Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yes, this is exactly why I said yesterday that OpenAI does benchmaxxing, seemingly quite a bit [1]. I got a flurry of downvotes for it at first, but I think people came around to it once they tried the model like you did.

Ultimately I think the issue is that OpenAI is under tremendous pressure to perform, but GPT-6 is not ready yet, so they had to push GPT-5 to its limits, and the only way they could do it was with really heavy RLHF, which has its shortcomings. Like, it is super obvious that Sol, Terra and Luna are all heavily biased towards working on a problem relentlessly because that's what their reward functions emphasized. That pushes up their scores in some benchmarks but does not translate to actual intelligence and capability.

[1]https://news.ycombinator.com/item?id=48849454



For me the weirdest is Luna. It costs the same as GLM to solve a task with it, but it just calls the same failing tool over and over again until we cut it out.

Now if you look at where GLM stands, or even DeepSeek v4 Flash, things get really interesting for what they provide.

US labs completely miss a cheap model that can solve problems for 95% of the people. Gemini 3.5 Flash could've been it if it didn't burn so many tokens.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: