I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it.
Do you have the same experiences?
I don't know what I could be doing differently to you but I found Opus 5 to be more reliable than even myself at times. Maybe your stack is unusual or you have conflicting commands in your prompts vs CLAUDE.md (that really confuses it)? It could be anything but this huge error bar in delivered quality is one of the biggest issues with LLMs.
Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it.
And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.
this would be _Great_ advice if you owned your own LLM and your knowledge was trapped in Amber because you were satisfied.
It's horrible advice given what we've seen consistent: changing alignments, changing guardrails, changing system prompts, changing inference priorities, etc.
Anyone who relies on these for their work product is chaining themselves to a matrix multiple of indetermintism.
I detect the regression already in planning with Opus 5, so I do not let Opus 5 implement anything. But it is a waste of time and tokens! Does planning with Opus 5 works out for you?
I would say "Elected errors _in_ Claude Opus 5" wouldn't be incorrect either.. Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it. Do you have the same experiences?
I don't know what I could be doing differently to you but I found Opus 5 to be more reliable than even myself at times. Maybe your stack is unusual or you have conflicting commands in your prompts vs CLAUDE.md (that really confuses it)? It could be anything but this huge error bar in delivered quality is one of the biggest issues with LLMs.
I do not experience any regressions, I don't really notice much difference either.
Opus 5 isn't very reliable for coding and introduces a lot of regressions every single time I use it.
And GPT 5.6 Sol over engineers just about everything. No LLM is perfect, its about learning the issues with each LLM and figuring out if you can live with it. Knowledge means that you can anticipate if it tries to pull something funny, and harness it against that behavior.
Sure, but I am a long time Opus user 4.5,4.6,4.7,4.8 and I wonder what's wrong with 5?
this would be _Great_ advice if you owned your own LLM and your knowledge was trapped in Amber because you were satisfied.
It's horrible advice given what we've seen consistent: changing alignments, changing guardrails, changing system prompts, changing inference priorities, etc.
Anyone who relies on these for their work product is chaining themselves to a matrix multiple of indetermintism.
Maybe it's my harness but I haven't seen it introducing regressions.
Do you not have unit tests, or how does it introduce regressions? You can tell it how to run the test suite in CLAUDE.md
It is a well known fact that projects with unit tests never have regressions.
I detect the regression already in planning with Opus 5, so I do not let Opus 5 implement anything. But it is a waste of time and tokens! Does planning with Opus 5 works out for you?
Opus 5 tries to modify the unit tests as a cover to its own regressions - thinking its own logic is correct and the test must be wrongly specified
Error message: API Error: 529 Overloaded. This is a server-side issue
Related: https://news.ycombinator.com/item?id=49066591 https://news.ycombinator.com/item?id=49056194 https://news.ycombinator.com/item?id=49067964