GLM 4.6 isn't a "fast" model. It does well in benchmarks vs Sonnet 4.5.
Cerebras makes a giant chip that runs inference at unreal speeds. I suspect they run their cloud service more as an advertising mechanism for their core business: hardware. You can hear the founder describing their journey:
Cerebras makes a giant chip that runs inference at unreal speeds. I suspect they run their cloud service more as an advertising mechanism for their core business: hardware. You can hear the founder describing their journey:
https://podcasts.apple.com/us/podcast/launching-the-fastest-...