The next big AI advancement will come from hardware.
The next big AI advancement will come from hardware. (Yes I am aware of Fable, and no I do not think it is a big AI advancement. An improvement, but this is not linear improvement. Twice the cost of Opus, not twice the performance). And we…
The next big AI advancement will come from hardware.
(Yes I am aware of Fable, and no I do not think it is a big AI advancement. An improvement, but this is not linear improvement. Twice the cost of Opus, not twice the performance).
And we are already seeing some really exciting stuff. I wanted to highlight a few companies you may or may know, which are leading the charge.
An easy way to put into perspective how much faster running AI models on purpose built silicon is to evaluate their TPS (tokens per second) of output. (For a reference point, you will see Claude output at anywhere between 40-150 TPS, depending on the model and the thinking mode.)
Taalas is turning AI models into custom silicon, and getting ridiculous speeds! Google “Chat Jimmy” for a demo. In the demo, I’ve been experiencing 14,000-15,000 tokens per second!
Cerebras just went public. Instead of running on GPUs, they are running tuned models on their own custom wayfarers, and running open weight models at anywhere from 1000-3000 tokens per second. They also have an API available for you to use.
Groq and NVIDIA entered into a non-exclusive licensing agreement for their AI chips. Since then it has been a little quiet from Groq, but I’d expect you’ll start to see them in the news more as their hardware is rolled out and utilized by NVIDIA.
Imagine how powerful running local models will become when consumer devices have true AI hardware!