the.bay.news

A LLM benchmark that gave a hard programming tests to gpt 5.6, but for much more languages than common benchmarks

programming.dev
A LLM benchmark that gave a hard programming tests to gpt 5.6, but for much more languages than common benchmarks
He’s right that python/js tend to score better on most tasks. OpenAI is very bad at less popular languages. Especially J, so that is a very bad skew in his results.

0 comments

Sign in to join the discussion — your thebay.events account works here.

No comments yet.