There is probably far, far more people trusting benchmarks and "I remade x thing in 1 prompt with y model" claims than you'd imagine, just ignore all of it. You know your workflow and what better should be and could be, measure on what works for you. You should note (and I may be wrong, not active in these part) a lot of those demos are very toy, recreating a known game, known app, known workflow, etc. very interesting but again a very toy example that should be taken with that in mind.
Benchmarks can still be useful, even for benchmaxxing labs. For example, the recent Beam model benchmarked much worse than DS4.1 Flash and GLM 5.3, both of which perform extremely well outside of benchmarking. Regardless of whether or not Beam was benchmaxxed, it's performance suggested that it wasn't capable of solving problems that other models in it's weight class could do easily.
The best-case scenario is that Reflection was being honest about their model's disappointing performance. The worst-case scenario is that they benchmaxxed, and it still managed to underperform compared to it's peers.
There is probably far, far more people trusting benchmarks and "I remade x thing in 1 prompt with y model" claims than you'd imagine, just ignore all of it. You know your workflow and what better should be and could be, measure on what works for you. You should note (and I may be wrong, not active in these part) a lot of those demos are very toy, recreating a known game, known app, known workflow, etc. very interesting but again a very toy example that should be taken with that in mind.
Benchmarks can still be useful, even for benchmaxxing labs. For example, the recent Beam model benchmarked much worse than DS4.1 Flash and GLM 5.3, both of which perform extremely well outside of benchmarking. Regardless of whether or not Beam was benchmaxxed, it's performance suggested that it wasn't capable of solving problems that other models in it's weight class could do easily.
The best-case scenario is that Reflection was being honest about their model's disappointing performance. The worst-case scenario is that they benchmaxxed, and it still managed to underperform compared to it's peers.