The costs of local that do not appear in the pitch, and which decide it in practice:
Memory, not compute. A model has to fit in RAM alongside the operating system and everything else. This is what actually limits what you can run on an ordinary laptop, and it is why the demo machine always has a lot of it.
Thermals and battery. Sustained inference makes a laptop hot and drains it fast. Fine for a task, unpleasant as a background service.
Distribution size. Shipping model weights means a large download and an update problem. Users notice.
Support surface. Every machine is different. Cloud has one deployment; on-device has as many as you have customers, and the failure reports are hardware-specific.
That last one is why so many products announce on-device and quietly ship a hybrid.