How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise
DEV Community
How to Build a Post-Launch Eval Canary That Tells a Real LLM Regression From Sampling Noise
Is the model actually getting worse, or did I just get unlucky on a handful of prompts?
0 comments
No comments yet.