6 comments

  • ekojs 38 minutes ago

    Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.

    [0]: https://x.com/EpochAIResearch/status/2095602754282783108

  • NiekvdMaas an hour ago

    Title: "major gains"

    First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)

      flyaway123 31 minutes ago

      Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.

      dist-epoch 13 minutes ago

      I think they mean cost per task, where Astra is now on the Pareto frontier.

  • Readerium an hour ago

    more like 5.7 not 6