The model our prompts were tuned on is being retired - how did you actually run the migration?
We got the deprecation notice for the version four of our features are built on, with a suggested replacement and a date. Spot checks look fine, and then it does something subtly different with a long input and our extraction quietly loses a field rather than failing. I have about three days of engineering time for this. What did you do first - rebuild the evals, rewrite the prompts, or run both in parallel and diff the outputs?
@salted_hash_h · 3mo ago
Diff first, always. You cannot rewrite prompts sensibly until you know which of them actually changed behaviour, and in my experience it is a minority - typically the long ones and the ones leaning on a specific output habit. Spend day one building the harness, day two on the handful that moved, day three on monitoring you should have had anyway.
Reply
Report