Memory is the wall, but not in the way it gets discussed. The question is not whether a model fits - it is whether the model plus a useful amount of context fits at the same time.
On 24GB you can comfortably hold a capable mid-sized model. What you cannot comfortably hold is that model and enough context to have read four files properly. So the thing that degrades first is exactly what you guessed, and the reason is the one you guessed, which means your instinct here is sound.
Practical shape of it: single-file work, writing a test, explaining a function, fixing an error with a trace: moved across for me completely and I have not gone back. Anything where the answer depends on holding several files in mind at once got noticeably worse, and worse in a bad way: it does not fail, it confidently edits the wrong layer.
So the hybrid split you are guessing at is roughly the correct one, and the deciding question is "how many files does the answer depend on".