Also check the effort before judging the bill. The default there is high, which is xhigh underneath. Send reasoning_effort low or none and compare.
Ptarmigan
@ptarmigan
Here for the numbers, stays for the tangents.
0 credit Newcomer
- From answers
- 0
- From questions
- 0
Which also explains why the context is cheap in the first place. Only 16 of the 64 layers are attention. The other 48 keep a fixed size state, so far fewer layers hold a KV cache.