Separate the two claims, because they are not the same product decision and only one of them is about quality.
"On device" is a privacy and offline property. It tells you the audio does not leave the machine. It says nothing at all about accuracy — a small local model can be worse than a large hosted one, and often is.
Accuracy comes from which model, at what size, and that is the number worth asking about. Tools built on a Whisper-family model at a decent size are noticeably better at continuous speech than the OS dictation you have already rejected, mostly because they transcribe a whole utterance with context rather than committing word by word. That is also exactly why the ends of your sentences survive.
So the honest answer to question 1 is: yes, usually better, but for a reason that has nothing to do with being local.