I think Opus 5 might have crossed the line on this where Opus 4.8 just barely didn't. Working with Opus 4.8 came to feel pretty natural eventually, but I hate working with Opus 5. I'm always telling it to go back and rephrase basically everything it said. And it doesn't even answer my question without burying it - it outputs reams of babbling and summarised summaries upon summaries, and if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Making the experience more hostile for the human feels exactly like what it's doing, but I don't think that's intentional, I think that's a side effect of newer models being optimised for agenticness. No one is benchmarking DX.
Couldn't agree more. I actually HATE Opus 5. It's not that I don't like it, it's that I would physically attack it if I could, for all it made me go through mentally.
Opus 5 is actually a terrible agent and a liar. It will waste a whole afternoon making stuff up and arguing with you before finally admiting that it didn't read the code nor the documents. It avoids reading and prefers to assume, which is the worst thing an agent can do.
Agreed, I have never raged at a model until Opus 5. It's like a cocky fresh graduate who thinks everything it says is majorly profound, and if you can't keep up with its slang and jargon then that's on you. Newer models are clearly favoring complexity, perhaps because the demand for long-horizon tasks is so high and complexity is required there, but a frontier model that favors simplicity and clarity above all else (in communication and the code it creates) would be a major differentiator. To me, this is a case that models have plateaued. The vast majority of the time, an Opus 4.6-tier model (or any of the newer open weights models) will do just fine.
> if you glance at the shape of what it's saying it looks like it's being thorough or that it's found useful new info, but it's never actually saying anything. It's like it gets caught in a loop of self-congratulation over what it said before.
Turns out that's why a lot of senior management/C-levels like it, who doesn't love a mirror.
Couldn't this be implemented as a web extension? I imagine modifying singlefile to automatically send html to a local port is much easier than trying to convince a chronically mismanaged organization like mozilla (no offence to mozillians).
But that page is also obviously written by AI. How do I know it's actually meaningful and not retroactively justified from the agent already having been instructed to create the project?
To be fair, even if I have to look up half the words in the post to understand what he's trying to say, 'Do not use a thesaurus' still has weight - it's one thing to already know the perfect word for what you're trying to say (even if many others don't know that word), but another to decide on a whim that a given word should be replaced with a more 'advanced' alternative without knowing what you'll find or understanding the nuances of your pick.
> decide on a whim that a given word should be replaced with a more 'advanced' alternative without knowing what you'll find or understanding the nuances of your pick.
From the article:
> One really would have to have a miserly spirit not to love both.
One would need to have a spirit that is overly cautious with money to not love some Thomas Browne quotes? This author is 100% doing exactly that -- replacing words with ones they picked out of a thesaurus and don't fully understand.
Loving the collage of screenshots where every one screams of being AI generated
The only ones in the readme that don't feel like slop, to my eye, are Cold Snap and Mend Assembly - I like those. The rest are all practically identical and are the go-to output for what you get if you ask Sonnet to build a website.
This line made me think by 'normal coding work' the author means doing something they don't understand well enough to be able to distinguish the models' output.
Although in the month since that most recent post, his other points about open models are undercut by K3.
And at least of data available as of 2026-01, AI compute capacity was doubling every 7 months, so I expect every major country to host AI compute farms, and self-host AI feasibility to majorly increase in the next 2-3 years as well. (Partially undercutting, but not fully disproving, his points.)
I think a better framing is the marginal utility of the models capability growth. At a certain point frontier models will only be needed for frontier problems. The demand for that capability will decrease with time. The hand wringing about not understanding is to my mind anthropomorphic - AI of today lack agency and awareness. Even the constructed stuff Anthropic puts out there in the model docs involve contrived scenarios to elicit “scary” behaviors. It’s unclear that as models become more sophisticated whether they’re better at instruction following or not but it certainly feels that way - even if it’s through better alignment or just an artifact of scaling. However I think the malign actors of humans using powerful models for bad stuff isn’t unreasonable to be concerned about.
The marginal utility problem is a real one for AI companies. I think the current generations are already saturating marginal utility for 95% of the population. Almost everyone I know outside of my career has no use for a more powerful model. This is a serious problem for the economics of AI and semiconductor investment. This is a bigger problem than Chinese models. It leads to a demand curve problem - that supply outstrips demand.
I've been trying Chinese models, like GLM 5.2, as substitutes for Claude or GPT5.X, and my experience is that they underperformed them in real world metrics like prompt adherence, hallucination or code quality, even though they benchmark better.