Operating System powered by Qwen 3.8 27B at 1950 tokens/sec!
here is what 1,950 tokens/second @Alibaba_Qwen's 3.8 27b actually looks like on @cerebras:
i wrote a minimal python web server that turns cerebras inference into a live operating system.
zero apps on disk.
when you
New episode of Uncapped with the founder of Cerebras @andrewdfeldman and my partner @ericvishria.
We talked about the company's winding road to success, the market around chips, the current state future of the AI supply chain, the role of investors and Eric's relationship with
Connecting the dies and working around defects were fundamental challenges for wafer-scale computing. Making it work also came down to power, cooling, and reliability.
@seanlie explains how those lessons are shaping @cerebras’ approach to stacking DRAM.
From @seanlie’s
Qwen3.8-27B is now live at Cerebras speed.
The dense, open-weight model from @Alibaba_Qwen scores 34 on the Artificial Analysis Intelligence Index—making it comparable to models such as GPT-5.6 Luna, DeepseekV4 Pro, and Claude Sonnet 4.6.