the.bay.news

Teaching a local coding agent from its own mistakes: DPO on a 30B model

DEV Community
Teaching a local coding agent from its own mistakes: DPO on a 30B model
Our VS Code assistant was passing every test on its curriculum — which meant the curriculum had stopped measuring anything. Here's how we built honest eval sets, found two silent contaminations in our test bench, and used Direct Preference Optimization on the assistant's own redirect pairs to teach a 30B model to pick the right tool on the first try. Five OutOfMemory crashes, one counterintuitive fix, a clean 4-hour training run — and a pre-registered eval gate whose verdict we report as measured, including the part that failed.

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.