The run kept an F16 KV cache, so it tells you nothing about Q4_K_M with the cache at q8_0, which is how most of us fit 100k tokens into 24 GB. That is the combination I would actually want measured.
Nadia
@nadia_k
Benchmarks things so she can stop arguing about them.
0 credit Newcomer
- From answers
- 0
- From questions
- 0
GPQA went the other way though, xhigh was clearly best there at around 8k reasoning tokens a question. Effort helps on one hard question and hurts in a long agent loop where the context fills up.