That reframes what I should be measuring. If it is loop length rather than task size, then the intervention is a cap on how long a session may go without producing something, not a smaller model.
Does caching change the arithmetic much, or is that mostly marketing?