Accelerating GPT-5.6 Sol Ultrafast(cerebras.ai)
707 points by pr337h4m 8 days ago | 276 comments
tl;dr: Cerebras and OpenAI have launched "Ultrafast Mode" for GPT-5.6 Sol, delivering up to 750 output tokens/second—reportedly 11x faster than Fable 5 and completing Humanity's Last Exam in 11 hours versus 78 for Claude Fable 5. The speedup is enabled by Cerebras' Wafer-Scale Engine, which packs 44GB of SRAM per chip to keep model weights on-chip and eliminate the memory-bandwidth bottleneck that slows GPU inference. It's currently available as a limited preview to select OpenAI API customers.
HN Discussion:
  • Excitement about speed improvements and their importance for iteration and quality of thought
  • Skepticism that Ultrafast mode maintains identical quality to regular GPT-5.6 Sol due to vague messaging
  • The comparison omits competing fast models like Mimo v2.5-Pro Ultraspeed, weakening the claims
  • ~Faster token throughput doesn't eliminate other bottlenecks like tests, typechecks, and grep
  • Anticipation for specialized ASIC hardware enabling local, offline, ultra-fast inference