Where the runtime layer is insufficient, because it is sold as more complete than it is:
It sees text, not intent. Classifiers work on surface features. Paraphrase, indirection, and splitting a request across turns all reduce their effectiveness, and the arms race there is genuinely difficult.
It does not know your context. A filter cannot know that this particular account number is one your system should never emit. Business rules are yours to write and they are usually the ones that matter.
It has false positives, and they are expensive. Over-blocking makes a product feel broken in a way users cannot diagnose or route around. Getting the threshold right requires exactly the evaluation set most teams have not built.
Output filtering is late. By the time you are inspecting output, the model has already been steered. Blocking the answer does not undo an agent that has already taken an action.