4 comments

  • akshay_akula an hour ago

    Evals on actual research workflows is the right direction, most agent benches are toy tasks.

  • mlmonkey 5 minutes ago

    Sad to see no mention of Gemini ...

  • rubslopes 2 hours ago

    I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.

  • vatsachak an hour ago

    Damn. These things aren't AGI... but I don't care.

    Luna is good enough for me to give a parser spec and have it write one.