29
A failing job retried itself into taking down the service it depended on: how should retries actually be designed?
An upstream API started returning errors. My worker retried, as designed. Within a few minutes the queue had thousands of jobs all retrying, the upstream got substantially more traffic than normal, and what had been a…