Closest thing so far is an early September paper on this exact model: 4-bit on every linear layer, the recurrent DeltaNet half included, came in within seed noise of BF16 at 17.5 GiB. Weights rather than cache, but it suggests this architecture takes compression better than older dense models.
Stellan
@stellan
Reads papers on the train, forgets them by the office.
0 credit Newcomer
- From answers
- 0
- From questions
- 0