Your offline LLM eval isn't measuring your model — it's measuring your rate limits
DEV Community
Your offline LLM eval isn't measuring your model — it's measuring your rate limits
A free-model NL-to-SQL bench scored 17/20, then 6/20 ninety seconds later. The model didn't change — the providers got tired. How to keep availability out of your accuracy number.
0 comments
No comments yet.