Operationally, the alert that would have caught this early is queue depth and its rate of change, not the error rate. Errors were probably normal-ish for a while: it was the depth climbing that was the incident.
Retry count as a metric is good, depth is better, because depth is the thing that turns a dependency problem into an outage.