2 points | by ryanmerket 7 hours ago
1 comments
Subhead: “This matchup wasn’t close. GPT-5.6 Sol dominated the practical details that decide real-world usefulness: tighter instruction-following, cleaner formatting, and fewer correctness slips.”
Score: 113.0 vs 93.5
Subhead: “This matchup wasn’t close. GPT-5.6 Sol dominated the practical details that decide real-world usefulness: tighter instruction-following, cleaner formatting, and fewer correctness slips.”
Score: 113.0 vs 93.5