- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me
I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.
Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.
You're right, it's probably quite unfair of me to say it eats lots of tokens when I am paying double for claude than codex and complaining about tokens.
The rest still stands, though.
But if I've learned anything is that in a 2 months I might have completely turned around, who knows
The thing is that this thing is constantly compacting.... I get 1M context with Claude and ~256k with Astra. Even if the compaction loses much less information on OAI's side, it takes so long it's barely any use for me...
I've tried High and Max. They have produced decent results, but they're so slow.... I will try to lower it a bit and see the difference, but it's a delicate balance: I don't want to waste literal hours on the incorrect reasoning level to only then have to spend those hours and tokens to do it right.
At this very moment, Astra has been working for 1h15m on a task. At this rate I genuinely expect it to take about 10 hours. I feel like claude would do it in at least a third of that. Let's see if the quality justifies the slowness (it better)
I'm curious: what languages or frameworks is this in?
The Django code that comes out of composer2.5, to me, was insulting. Grok definitely was a step up, especially because the fast option reaaally is fast so even if it came out a bit wrong I could just whip it into perfection.
For frontend work, it's a different story. You can still tell that composer2.5 is taking the long route, but I don't think it's as egregious as with Django.
Also, composer2.5 would routinely run commands that were really dangerous and in need of proper sandboxing. Things like creating an ./uninstall.sh script with a HOME variable on which it does rm -rf $HOME. In general, when I asked composer2.5 to do things "for me", I knew a third of the initial commands would be failures, and sometimes they could be catastrophic failures (it did actually run rm -rf $HOME on what would be an actual home folder). This just hasn't happened with Grok.
I also have a bunch of vibe-coded apps I built for myself with composer2.5 and it is extremely noticeable that they hit a "this needs to be refactored as it's crumbling unto itself" line much earlier than with Grok and proper frontier models.
I had the same sort of thing going on ahah. But I was convinced some hoje must have already done it and I decided I’d research when I got home (I’m out today). Didn’t expect it to reach hacker news so soon, though!!
How do you protect intellectual property? Or is this a case of the value being somewhere else, such as in your backend? If so, how does the agent debug frontend and backend? I presume it stops at frontend
It is the number 1 thing I cannot stand with Claude slop. It's a sort of anthropomorphization of language. Every "thing" does, produces, feels, wants, asks, answers, etc....
- "Launch is checked"
- "Question is asked"
- "The implementation answers"
- "The model wants"
- "The results name"
- "The connection surfaces"
- "The prompt wires"
- "The feature rides the mechanism"
Every single fucking thing is alive, wants things, and does things.
It's terrible. Infuriating. I want to rip my eyeballs out reading this filth. All. The. Time. "The anger is real".
Create any page with a file uploader. They all look the same now. It's like the Twitter Bootstrap days of responsive design. You'll get an icon which looks like ones on (on the drop space) those sites which are like "you must wait 60 seconds for this file to download".
It's so horrible. The human element has been completely removed and replaced by..... mediocre.
The human element has been completely removed and replaced by..... mediocre.
No it hasn't. The human element is still there, prompting the LLM. The change is that the human is happily accepting the first thing they get rather than critically looking at it and seeing a problem.
I don't think it's that because I see a lot of this in businesses where the human isn't paying the bill, or is even aware of what the bill is.
Humans are seeing either a shortcut to go faster (accepting low quality to move on immediately; reasonable if they're short on time) or a shortcut to lowering effort (accepting low quality because they don't care; not so reasonable but probably has a deeper root cause).
I'm late to the thread, but my experience with Claude Fable 5.1 has been absolutely horrendous.
Things it does constantly that Fable 5 barely ever did:
- Act without my permission. All. The. Time. "Oh I just finished this thing we were discussing, let me push it without ever having been told to do so."
- Immediately jump to action instead of addressing me first. If I say "I wanted to write tests for this and run them" it immediately starts writing tests instead of digging into what "this" is better -- literally does not give me any feedback and starts spitting out code. Naturally it creates the wrong tests
- Despite claims that it does not write like "stereotypical Claude" anymore, in my experiments it is far worse than before. Replies are longer, more filled with fluff, and still flooded with garbage language. Hard to parse.
- It loves to answer my set of two direct Yes/No questions with 5 paragraphs where it only answers one of them and answers 4 other questions I didn't ask. Notice how it misses one of the questions.
- It. Is. Cocky. Absurdly full of itself and arrogant. Just the whole way it presents and answers passes this energy of "No, but really, you're wrong and I'm right". It often is not right. What annoys me is not that it's wrong more often than before (which it may be), it's that it doesn't own up to it as before. Insulting if it were a human.
- Replies and addresses me directly in its thinking traces, and then assumes I've read it. I ask a question, it answers it in the thinking traces and does not relay it back to me at all. This is the only one that Fable 5 also did, but 5.1 is doing it an order of magnitude more often.
- It's too early to really tell, because I may just be working on particularly harder problems today, but it seems to get things wrong more often. I've had to bump it from high to xhigh to compensate.
My guess is I must be having a bad day or something. Although this is happening on multiple projects run from multiple machines (fully isolated, except for the account, which is the same) all in the same way.
What got me was "I keep seeing the same shape next to billing events in other people's reports".
Clicked off after that.
I'm extremely bullish on AI, but I am tired of people not using their own words. Is everyone so insecure about the personality they project in writing that they need to replace it with some AI bullshit?
(To the author: if you didn't do this, which I accept is a possibility, I'm sorry -- you've been caught by the storm of fatigue that has plagued us due to those who do replace their words with AI crap)
Hmm. You are making a good point, and I wouldn’t give the benefit of the doubt to this post’s author. But you can’t just unsee “this is twice load bearing”. Some people will wonder if it’s a training, self-reinforcing feedback artifact. For others, it will work like a memetic virus[^1].
The paradoxical thing is that those of us who have English as a second language should have an easier time producing non-AI-English by the mere act of directly writing what we want to write.
I am catching myself distrusting what people write ever so often. I've got cases of e-mails I'm CC'ed in and I am left wondering: "Has this person always written like this? Is this just an angle in this specific thread (e.g. for sales purposes)? Or are they really relying so much on what the blandness-machine spits out? Am I going mad and wielding my hammer looking for everything in the ....shape (eheh)...of a nail?"
Actually, now that I've written that, it's clear I need to have a talk with some people....sigh...
Yes, I did think about that possibility. Using a tool to translate to english doesn't mean we don't proofread it.
Ok, I hear you: one could argue that they simply don't realize this is a very non-idiomatic way of writing precisely because english isn't their first language. And that simply giving a preamble "sorry, my english comes from a translation tool, excuse any mistakes" would seem obtuse next to every english text they write.
Ok, I'll concede that. In which case I just have to accept the fact that, perhaps, if LLMs don't solve the problem of writing like garbage, and people keep using them to translate, a bunch of us will simply not want to read what you translate because it reads terribly and messes my brain up.
However, I find it hard to believe that an LLM would translate anything they've written into such a clear LLM trope. I would find it much more believable that they simply brainstormed to an LLM in their language and had the LLM build (at least that part of) the article from it. Even if they proofread it, since english is not their language, it slipped by. Either way, in this hypothetical scenario, the LLM was still used to write (at least a part of) the article, not just translate it.
I also find it somewhat hard to believe that someone who I believe is a developer (and has been using claude for a very long time) does not know enough of both english and claude to identify this specific pattern in the text they are putting their own name to it.
So, sure, maybe the author doesn't predominantly write in english or interact with claude in english, and they did write everything and simply piped it through an LLM to get a translation and, after proofreading the best they could, just didn't notice that.
Alternatively, maybe they just don't give a fuck, something which I can absolutely respect. They lose a reader (at least for now), but stand by what makes sense to them and I genuinely don't think less of them because of it -- I just can't read what they produce. I _would_ probably think less of them if I had to work every day with them and this were a repeated pattern of interaction, but that's not what's being discussed at all.
(Disclaimer: English isn't my first language either, which I think also does come across)
And? I’d rather reader poorly written English written by a person about his actual experience - or just read nothing at all - as opposed to reading thinking and writing that has been outsourced to a machine. Being able to think and write confers privileges on those who can do both; the idea that somebody should be able to benefit by doing neither is pretty fucking perverse.
I experimented with Hy3 for a project and was surprised with how good it was. I don't know if it's good for coding, but as a general purpose agentic model, it was only beaten by deepseek4-flash in our tests. It was so close to deepseek behaviour I kept thinking it must have been forked from it.
For the last few days I've been experimenting with the _free_ version of Hy3 offered by Opencode Go and I was also surprised to see how (relatively) good it is on coding tasks too.
The free quota from Opencode Go is also surprisingly generous, I perhaps hit limits one or two times and I've been using it _a lot_ for implementation tasks (using e.g. GLM-5.3-flash for working on specs and planning next steps).
- Unbearably slow
- A token eating machine like no other
- Constantly compacting
- A model (like other GPT ones) that hides thinking traces and thinking summaries, which infuriates me
I've been in the Claude camp for a while, but the way it writes has left me with a a brick for a brain and wanted to see if Astra was as good as they say. Well, I can't know, because in the time it takes for it to actually build anything useful, I've moved to other ideas.
Unbearably, annoyingly slow. I keep thinking I must be doing something wrong.
reply