Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.
More discussion here
https://news.ycombinator.com/item?id=49038433
System Card https://www-cdn.anthropic.com/c5fbac3f0b1280a933ebd26d3cb8bb...
Very impressive headline benchmark numbers. I expected a step change, but not past Fable. That said - it all depends on whether the classifiers make the model unusable...
Very interesting to see such a focus on cost for performance here
> arc-agi-3 30.2%
wow
here we go
The benchmark table is manipulative, borderline lying through statistics. In every line the top performing cell is marked red. Except the line where Sol leads, there it is marked in gray.