27
How do you tell whether an agent got better or worse after you changed something?
I keep changing prompts, tools and models, and my evidence for whether it helped is that recent runs felt better. That is obviously not evidence. A batch of observability products has appeared aimed at this — tracing…