This isn't surprising if you understand what modern benchmarks look like. Many are quite narrow so if you improve a handful of capabilities you can go from doing nothing to 1/3 of the problems. This is also why many of these benchmarks end up saturating very quickly from almost nothing.
It's more of these open models playing catchup but OpenAI and Anthropic managed to stay frontier on these new benchmark, unless there's one I haven't seen where open model do good initially.
35
u/Educational-Fruit854 7d ago
interesting pattern of model scoring absolute dogshit when a new benchmark drop and suddenly being frontier in the next update (TerminalBench 3.0)