Sure that can be called an attack, but then we must also concede these labs essentially massively attacked everyone else in existence to get the data, and continue attacking as we speak.
In a way you could see this as a case of Robin Hood. The US companies exfiltrated all the data on the planet just to hoard it for themselves now and accuse anyone who tries to get a piece of that back from them, and the Chinese labs are distilling it to offer it for cheap.
Obviously a bit more complicated than that but it still holds pretty well.
To be fair in every harness I've ever used that has LSP support, they never actually utilize the LSP for more deterministic refactoring tools. And even then when I do enable the LSP in many harnesses oftentimes it doesn't even use it at all.
Maybe they haven't been taught to do so or it's not integrated into the system prompt or the tools but all of them only ever use the LSP to read files/symbols.
Every harness I've used will happily just call the edit tool over and over or do a find and replace via sed or programatically call a python/perl script rather than rename a symbol via other means.
This is partially a "how much does the model follow instructions" thing. I use Pi with a vibed LSP extension and Claude (4.6 or thereabouts) almost never followed instructions to use LSP renaming tools - despite it being strongly emphasized in system prompt and agents.md. However I found Codex 5.3 would use them sometimes, and GPT 5.4/5.5 would prefer them.
If what ends up happening is that every listing has misleading AI photos but they have to disclose it, then also what ends up happening is nobody trusts them anymore. Consumers will know by default to not trust the photos.
In my eyes, thats a win since that's a better outcome than them secretly using AI photos.
Of course in my ideal world it would be outlawed altogether, But even if they were still allowed to use AI photos but were forced to disclose it, that's still a good first step.
What I read on social media about people and these resets gives off literal worst kind of addiction vibes. I've literally seen people talking about "Oh I had an existential crisis without Fable/GPT-5.6"
These people legitimately need help, or alternatively a social life.
Maybe its different on my end because I just use a sub outside of work for fun stuff. At work its not my money so I don't really care. I go to work, maybe use these subs at home every once in a while for a fun personal project and if I hit the limits (I rarely even do) I play video games or hang out with my wife/family.
People are borderline tying their identities to these models it seems, and yet most people aren't even building anything interesting.
Yes, 100%. The author seems to be describing, quite lucidly, how he is getting sucked into a gambling-like addiction to these LLMs… although he seems unaware of the implications of that.
I have a $20 Claude sub, and the closest I've come to hitting a limit was when I was using it to fight back against some scammers spreading a stealer on discord. I was having Claude decode their webhook strings, and the first time I blew them up with pings to get a job, then just started deleting the hooks from under them when they published new versions.
But for standard productivity, I've never come remotely close to hitting a limit.
Same here, and - while not trying to troll - the people I know who loudly proclaim how they keep hitting their limits are the ones who also seem to be blindly trying things through trial and error. A lot of wheel spinning, etc.
More like trying to prevent people from moving off of them.
The only people that these companies are really trying to limit are the tokenmaxxers on flat rate plans and enterprises on flat rate plans that are cheating and should be on the API plan.
For everybody else, it's way better to reset their quotas if you have the capacity. If you cut someone off, they're likely to either associate your brand with being unreliable or go try a different tool. Both outcomes are going to cost you money.
> It's Thariq from the Claude Code team here. This was my change! I made the AskUserQuestion tool so am generally in charge of maintaining it.
In a sense yes, I think it is actually reasonable to complain that the answer is too human/individualized here because it likely wasn't this individual human who made this decision, but he's making it seem like it is so that we are less likely to blame the company as a whole.
It's counterintuitive but when one singular person owns up to the problems that, at the root, are actually systemic to the decision making of the whole company it plays on the psychology of us as humans.
"I'm in charge of maintaining it" - This is not the same as "I'm in charge of all of the decision-making behind the implementation of how this tool works for users".
I actually agree exactly with your last point that one single person taking blame is counter-intuitive/non-productive here, but it actually seems like what these large companies desire is to have one person be the fall guy to play on people's sympathies.
If this were some small startup it would make sense but this is not that case.
It’s also an unforced error in my eyes for a multi-billion dollar company with competitors in the US and outside. Mistakes happen but no reason to turn a mistake into a revelation about processes that leaves no ambiguity, one that leaves the mythos (pun) of your company less than before.
Unfortunately this is a sign that the systems by which PRs are shipped need serious inspection.
For example in incident management, the industry settled on blameless postmortems. Blame is pointed at the systems which allowed incidents to occur, never the human(s) who triggered the incident. "Just don't make that mistake again" simply doesn't work. Humans make mistakes!
There are too many stakeholders in a product this large, even if every one of them wishes it were not so. The systems by which this PR shipped need serious inspection.
Habibi, I don't think anyone is really blaming you personally. We are piling on to the fact that this piece of critical infrastructure many of us depend on day in day out is being built in such a way that a single developer can wake up one morning, ship a change they thought might be nice, and then walk it back the next day.
I don't really care, in a most positive way towards you.
I don't hold you accountable even if you twirled your mustache like a villain. You are not responsible for Claude Code. Why did the leadership set up such a weak, error-prone process?
Again, not asking you directly, because it doesn't matter to me what you say, respectfully.
Seems strange for a company of your size to have one person push changes that should have easily caught this edge case then. Seems like a change even a small handful of people could have reasonably thought up this side effect of.
This is the kind of weakness I was alluding to… internal information that shows dysfunction is now shown to the public and there is no ambiguity. People pay attention and there is a low tolerance for unforced errors.
There has to be a second PR approver. There's no way a company as large this can allow an individual the ability to push to production without a secondary person involved.
When there's at least two people involved you can then start to look at more systematic factors that go beyond any single human mistake.
When both the submitter and approver miss that there's no changelog entry for a PR does that mean a checklist is needed? Should the changelog be automated with the commit message which can be improved for external consumption?
I'd find it very hard to believe that this high profile change just happened to be the only change missing from the changelog.
Bugs are always going to happen but housekeeping like writing something to a changelog when a commit is merged can be quality assured.
I don't know, this is a company that is pushing the narrative of software engineers becoming obsolete. Why do you think they would have respect for Software Engineering processes. As few as they are concerned, a few coders working in isolation with Claude Code is all that you need. (pun unintended)
If there really is no PR reviewer and no separate QA what is Anthropic even doing. I have respect for someone owning their mistake but as others have mentioned it reveals what may be a systemic issue.
I really wish there was just an easy guide on when to use Sol vs Terra vs Luna, and it just moves further into confusing territory when it comes to naming.
The naming convention is especially difficult to decipher depending on what your native language is. Of course a latin language speaker might be able to easily determine oh yeah each one is slightly bigger than the other but I still think it borderlines too confusing.
That aside all the numbers look amazing, and I'll be happy to probably main this alongside grok-4.5 for a while comparing the two on price and efficiency.
I vastly prefer the direction that OpenAI seems to be going with token efficiency and performance compared to Anthropic who seems to be moving towards a world where you just token-max as much as possible ignoring any and all costs.
I use the strongest model (5.5 now 5.6 sol) on the highest reasoning effort with /fast for everything. With a $200 pro sub I can't even use my weekly limit. And it's faster than using a weaker model that makes more mistakes which I have to waste time fixing.
Same, I am actually able to reach my weekly limit, but when I start going below 30%, I switch to normal speed, and that usually gets me through the week.
I absolutely save money and time by constantly using everything at the highest reasoning. I guess my use case and needs are different from others, but I really don’t understand how it can be true when people say they don’t need the highest reasoning and best model. Every time I drop down, things are missed, code gets unnecessarily bloated, more mistakes, and more iterations to solve the same problem. I think it might be because I’m spending a lot of time in a legacy system that I’m trying to clean up, and given the messiness, one needs all the reasoning available to decode what the hell is going on in there.
I have found that higher reasoning doesn't always produce better results. Models often tend to overthink and get themselves in a bind. And tasks just take longer for no reason.
I used to have the same experience until 5.6 sol xhigh. I have instructions in my code review skill and agents.md to encourage parallelism including multiple agents as long as quality isn’t impacted. I additionally instruct codex to not use less capable agents because at least with 5.5 this would seriously increase slop. Maybe sol is smarter about delegation. Hopefully because I’ll have to slow down or hopefully get approval for extra use credits. Now’s a great time for a limit reset if anyone from open ai is reading :).
My guess is that it's the same for Haiku/Sonnet/Opus: Biggest model for architecture and high level planning and technically challenging problems, medium model for simple implementation tasks, small model is for nothing
> I really wish there was just an easy guide on when to use Sol vs Terra vs Luna
Their dev guide has the following:
> Use gpt-5.6-sol for frontier capability, gpt-5.6-terra for a balance of intelligence and cost, or gpt-5.6-luna for efficient, high-volume workloads. The gpt-5.6 alias routes requests to gpt-5.6-sol
It's just how LLMs work - these are three completely separate models trained in parallel, with different numbers of parameters and using different amounts of compute.
They could hide this behind a harness that picks the correct model for you, but devs don't seem to like that.
There's also the 'effort' slider, which I guess how many experts in the MoE are evaluated and how long reasoning chains are allowed to go on, which is the 'smooth' scaling you are thinking of.
Use Luna. It's more performant than 5.5 and it's cheap. Hopefully it's cheap because it's more environmentally friendly than the bigger models. So you're doing a good thing. If it's a smaller model it may even be faster, but I haven't looked into it yet.
Previously it was much more obvious which model to reach for depending on your use case because they had the mini and nano naming conventions.
Getting rid of that seems like a step back. Just a personal nit though.
I've seen buzz about this elsewhere as well but to me effort levels seem more like spend limits disguised with another word. I don't think they should even exist.
The naming convention is bizarre and doesn't really mean anything to normies. Trying to pick between "Sol" and "Terra" is like asking the average person if they want the Max or the Ultra chip.
"i really wish this thing in my non native language was easier to decipher"
huh? if you dont know the words then read them in your native language. Sol/Terra/Luna are immediately unambiguous to an english speaker with any sense.
That isn't what "genuinely asking" looks like, you're criticizing using "questions" as cover. It isn't subtle, nor is it constructive.
I agree with them, Sol, Terra, and Luna are confusing names. They mean the same thing as GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast but require base knowledge for an analogy.
It feels like it was adding by the marketing department.
I'd agree it is similar to Anthropic's naming scheme, which I'd argue shares the same problems as this. It improves marketability/googlability, but decreases actual comprehension.
You don't actually explain why or how these names are "easy to understand" just state that they simply are. That's great; to me, they aren't obvious or intuitive at all. May have well just start randomly pointing at dictionary words.
I collect roman coins, with latin legends, so the sun/earth/moon references jumped out at me, and partly based on the opus/sonnet/haiku precedent I assumed that these names were referring to different model sizes/prices in a way that mapped to the names (Sun > Earth > Moon).
I'll admit though that until recently I never really thought about Anthropic's naming scheme as having meaning (an Opus being longer than a Sonnet, being longer than a Haiku).
This demonstrates the problem with an education that has no emphasis on the liberal arts, such as critical thinking. No justification is needed for why and how these names are easy to understand.
>They mean the same thing as GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast but require base knowledge for an analogy.
But do they though? When do you use GPT-5.6-Max-Low vs. GPT-5.6-Plus High? Or GPT-5.6-Fast-Xhigh? What's the Pareto optimal choice (outcome and price)? According to the benches it seems to bop around and the even if the benches are accurate the best choice isn't always consistent.
> When do you use GPT-5.6-Max-Low vs. GPT-5.6-Plus High?
You don't, because that isn't something I proposed using for model naming.
I called them GPT-5.6-Max, GPT-5.6-Plus, and GPT-5.6-Fast. Reasoning levels are distinct from the model design itself, and the UI makes that clear.
Plus, using that same flawed argument this would be called GPT-5.6-Sol-Low or GPT-5.6-Luna-High which also makes no sense/is confusing. So that argument applies (or more accurately doesn't), no matter the model names.
Your names are not good because “Fast” is not a descriptor of model size and overlaps with fast/ultrafast inference. And “Plus” collides with the ChatGPT subscription plan. Point being, naming is hard.
I do know what Sol/Terra/Luna mean, but was also confused for a second on the hierarchy. After doing a bit or research it dawned on me that they are arranged in the order of the sizes of the celestial objects but it somehow wasn't immediately obvious to me from the start.
Anthropic ships models with a helpful one-liner tag that makes the model hierarchy obvious. I think it wouldn't hurt if OpenAI did the same.
Did you not read the second sentence? Obviously I know what sol is given my first language being Spanish. I'm just speaking in a general sense that it can be confusing for others.
I already know plenty who had no clue what the difference between Terra and Luna would be.
My first instinct was Sol > Luna > Terra, since Sol is the farthest away, then Luna, and Terra is the closest. Size was not my first instinct. Or should Terra be the best model because its closest to people, then Luna because there have been people on it, then Sol be the worst because no human has been there?
Competition for cheaper and efficient models is a good thing, regardless of if you don't like SpaceX, Meta, etc. Especially from US based labs
I for one am really glad to get competitive models that will push the major labs to bring prices down. While Chinese open source labs are also great, unfortunately when it comes to US/Western political pressure it won't often have as much of a bearing on labs bringing prices down, especially for enterprises.
Also if these numbers are true, this is truly breaking ground finally for Meta.
I'd probably say there instead of a Challenge Mode and a Relax Mode like you said, it could just be a combined mode where there is a timer but after it goes out it simply continues the game on Relax Mode.
Or alternatively every word still has the timer and then at the end if you finish, it tells you how many words you completed under the timer and gives you a score based on that.
And then maybe an option for those who don't want the timer to show at all, since maybe it adds a bit of pressure. You can have just a simple option that removes the timer entirely from view
Another idea: maybe time how long you take for each word, and for the competitive among us, show stats on how long you took compared to everyone else, and a leaderboard for who took the least total time.
I mean its the same thing as the data center investments.
Think whatever you want about them, whether they're good or bad when it comes to environment, public health etc.
But one thing cannot be ignored - that they are not built to employ some large swath of people. They can be run with very lean teams, much leaner than the average person thinks for something so large. Any claim that they are employing some measurable amount of people is a sham they try to push onto the public.
If we actually punished corrupt officials and we had some kind of truth serum that forced people to admit yes/no as to whether they are corrupt, I would not be surprised if the majority of officials in the federal government would be culled. Its practically a breeding ground for corruption.
In a way you could see this as a case of Robin Hood. The US companies exfiltrated all the data on the planet just to hoard it for themselves now and accuse anyone who tries to get a piece of that back from them, and the Chinese labs are distilling it to offer it for cheap.
Obviously a bit more complicated than that but it still holds pretty well.
reply