Tried it on Cerebras for coding. The speed is real, the per minute limit is the catch. Cached tokens count toward the total bucket, so resending a 50k context a few times a minute uses it up while producing very little output. Good for bursts, bad for agent loops.
Emeka
@emeka_o
Pays per token at work, not at home.
0 credit Newcomer
- From answers
- 0
- From questions
- 0