The post does is totally missing cost efficiency. They have to show cost per ELO points and then linearize by inverse logistic curve.
With this rating method you get wiiildly different results.
If you needed high-end inference you can always go a bit down from top efficiency and search for the sweet spot - which is probably ChatGPT 5.6 Sol.
It's expensive now, I expect once it is with inference providers it will be really dirt cheap. Then it would be truly be a "bicycle for the mind", which Fable promised to be except it proved to be too capricious for that.
When DeepSeek was released, it had an immidiate and significant impact on the US stock-market. Now when its becoming common knowledge that China is almost at parity with US SOTA models with good momentum, why is there no sentiment change on the market?
I’d imagine partly a belief, right or wrong, that protectionism and regulatory capture will reduce China’s models impacts on the US market. Just like automobiles: the big 3 should be terrified of the likes of BYD, but aren’t.
Back in 2025 it was common to test models in a different harnesses.
I remember watching a guy on youtube, who was testing every new model in opencode, cline, codex, claude, etc.
Why did it come out of fashion ?
EDIT: ah, yeah. the point was that a harness would often affect results (task completion rate, I think) for more than 10%
Specifically they use this harness: https://github.com/ArtificialAnalysis/Stirrup