16 points | by wertyk an hour ago
6 comments
Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.
[0]: https://x.com/EpochAIResearch/status/2095602754282783108
Title: "major gains"
First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)
Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.
I think they mean cost per task, where Astra is now on the Pareto frontier.
more like 5.7 not 6
5.6.1
Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.
[0]: https://x.com/EpochAIResearch/status/2095602754282783108
Title: "major gains"
First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)
Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.
I think they mean cost per task, where Astra is now on the Pareto frontier.
more like 5.7 not 6
5.6.1