ARC-AGI-3 news: In the past two weeks multiple groups have scored at least 99.9% on the public challenges at relatively low costs.
lemmy.world
ARC-AGI-3 news: In the past two weeks multiple groups have scored at least 99.9% on the public challenges at relatively low costs.
The top three solutions come from independent researchers. The best solution was built by a group of PhDs and professors, who released a corresponding paper [https://arxiv.org/abs/2607.28287]. They all make use of some form of world-model. I’ve generally been a skeptic, and I still am, but this news surprised me because I expected ARC-AGI-3 to remain difficult for a long while. Note that the scores are self-reported and need to be independently verified. The solutions have not been tested against the larger private test set. Primer on ARC-AGI-3: > ARC-AGI-3 is an interactive reasoning benchmark which challenges AI agents to explore novel environments, acquire goals on the fly, build adaptable world models, and learn continuously. > > A 100% score means AI agents can beat every game as efficiently as humans. > > Instead of solving static puzzles, agents must learn from experience inside each environment—perceiving what matters, selecting actions, and adapting their strategy without relying on natural-language instructions.
0 comments
No comments yet.