The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?
DEV Community
The Model Scored 30%. The Harness Scored 100%. Which One Did You Benchmark?
Four harnesses took the same public ARC-AGI-3 set from 13% to 100% without touching a single weight. Then Microsoft put the harness inside the training loop.
0 comments
No comments yet.