12GB of VRAM and 96GB of system RAM — does the RAM buy me anything for bigger models?
3060 12GB, and RAM is cheap enough that going to 96GB costs a fraction of what a bigger card would. I want to run 32b-class models and I can tolerate slow, but not two tokens a second slow. Everyone says VRAM is what matters, and yet the loaders will happily spill into system memory, so I want to understand what I am actually buying before I choose between more RAM and saving for a used 24GB card.
@yamlfatigue · 3mo ago
Numbers from my own machine, which is close to yours: 12GB card, 64GB of DDR5. A 32b at a four-bit quant with about half the layers offloaded gives me something in the three to four tokens a second range, and it degrades further as the context fills because the KV cache competes for the same VRAM. It is usable for a batch job you walk away from and genuinely unpleasant to chat with. Below about eight tokens a second I stop using a model interactively, and everybody I know has a threshold like that.
Reply
Report