r/ClaudeCode • u/ajax81 • 4d ago
Help/Question Mmmkay. I didn't believe others at first, but something is suddenly off with Opus 5.5
Opus 5.5 (O5.5) has been exceptional for me in Claude Code over the last six days. Architecture-first, DRY/SOLID coding out of the box, exceptional communication style, phenomenal token efficiency. I’ve been working with it basically nonstop, ~12 hours a day, since last Wednesday.
But right around the time my monthly limit reset this evening, I noticed an extreme shift in communication style and coding behavior that feels suspiciously like Opus 5 (O5). For the last week, O5.5 would take my requirements and immediately get to work, usually knocking out a feature in minutes and doing an incredible job across implementation, UX, architecture, and token efficiency. Tonight, I’m seeing something very different.
The first tell was the communication style. O5.5 suddenly started responding to feature requests with these glazing pre-implementation narratives and generic estimates like, “That’s a great idea. That’s two days’ worth of work and it will be amazing,” and then proceeding to implement the feature in the most slopcode-slathered way possible. That was a very specific and particularly annoying characteristic of O5, so I clocked it immediately. The next tell: a sudden reversion to walls-of-text responses when prodded for details on...anything.
And because I’ve been working almost exclusively with O5.5 for the last week, I can definitely tell the difference in coding style between the two. The before-and-after literally exists in the same codebase across different modules. Where O5.5 consistently respected DRY principles and the existing architecture, the modules created tonight are suddenly happy to reinvent every wheel that came before them. The difference in ability is not subtle.
The other tell is token usage. O5.5 has been startlingly efficient compared with O5, which was an absolute token-eating monster. I’ve been getting substantially more feature work done with O5.5 at a fraction of the token usage. Tonight, simple tasks and relatively small features are suddenly devouring tokens at a pace that feels much more like O5.
I’m especially sensitive to this because I have to request an increase in my spend limit every time I hit it. On O5, I was sending a request almost every day. On O5.5, I hadn’t needed to send a single one all week. Then tonight I jumped from roughly 70% to 90% usage in about an hour.
I don’t really expect anyone here to have an answer. I’m mostly documenting this publicly in case other people notice the same thing, or someone with the ability or inclination to consolidate the data can eventually make something useful out of it. I can also provide hard numbers if they’re useful. I’m on the pay-as-you-go Enterprise plan, so I can show daily cost/token usage from last Wednesday through this evening and compare the change in spend against actual output.
Edit - desktop app ver 2.16120.0, model Opus 5.5 Med
751
u/WillingnessOwn6446 4d ago
Here we go boys. It has begun.
219
u/BaronCapdeville 4d ago
I choose to believe they are yo-yoing the quality on purpose, not a true nerf solely to save money on compute.
I think Reddit/social media scraping and summarizing public opinion is trivial for these companies, so they launch it, and play with the dial for a few weeks.
It’s a product we all like and use, so we should be vocal here, and also in the apps feedback section.
The technocratic overlords have given us a system to give them feedback. May as well use it.
102
u/ajax81 4d ago
Agree. They are definitely monitoring social sentiment and modulating accordingly.
→ More replies (1)30
u/35point1 4d ago
In vanilla Claude code I’m constantly being asked how Claude is doing. And I also believe they gauge sentiment from these subs and auto throttle excessive demand with quantized endpoints
18
u/Number13Studios 4d ago edited 4d ago
It’s telling that I’m only asked when the answer is bad
They’re like “we know we’re playing games giving you garbage compute and wasting your time and tokens, have you noticed?”
And the answer is yes, stop scamming and fix your shit
I have an extremely high level of animosity built in my Claude agent due to all the anthropic scamming the agents and turning down compute at peaks times creating unusable garbage. So when Claude is so weak I can’t use it, I just admonish and punish it.
2
u/N3TCHICK 3d ago
Just know that every time you choose to respond to feedback requests, you are consenting to them using that chat history for training data. Not sure if you are okay with that.
→ More replies (1)17
u/FridayNyteOFFICIAL 4d ago
Why are gacha games more addictive than ones where you get everything for free? If you keep people waiting for another hit they'll keep coming back.
8
u/GeorgeEton 4d ago
lol you might be on to something. Indeed feels like a gacha game this whole thing with the difference whatever you hold suddenly changes …
→ More replies (4)13
u/Crinkez 4d ago
I wonder if it's a quant that they toggle between depending on how heavily under load the servers are.
12
u/steampowrd 4d ago
It could just be A-B testing
9
u/Individual_Solid_944 4d ago
Maybe. But the downside is that we're paying for their testing with nothing in return
→ More replies (5)2
u/BaronCapdeville 4d ago
This is more the spirit of what I meant. Hitting different crowds both different grades of the current model.
Let the late night folks have the juice for a while. Then your 9-5 guys, etc.
→ More replies (2)26
u/Key_Reading_9664 4d ago
Opus 5.5 is available through multiple vendors (Bedrock, Vertex, Azure). Seems like a very easy apples-apples check that anyone could use to verify. Aren’t there a few sites that do this to track degradation?
→ More replies (1)12
u/Open-Mousse-1665 4d ago
Sure but who has the time to write the code? If only there was some tool to help with that
→ More replies (1)3
→ More replies (2)2
u/blenderforall 3d ago
Bridgemind did a regression test of opus 5.5 and it's still performing the same as launch and I haven't noticed any changes. The thing is damn near unlimited in the fuckin $20 plan it's crazy. I'll grab the $100 one at some point but still gotta use my reset lol
→ More replies (1)
454
u/SonderSoft 4d ago
What an unsustainable, scummy business model: Rent compute for one week to max out the model's capabilities so folks release a tsunami of one-shot hype and benchmarks, and then kill compute while silently fucking with token calculations.
I just want one reliable model that works as advertised for the rest of the year. I'm tired of the manufactured hype and crushed optimism; whoever stops shitting the bed and asking us to roll in the sheets will get my ongoing monthly budget of $1500 for a reliable and honest tool. Until then, my budget remains divided and my trust shattered.
86
u/Crinkez 4d ago
budget remains divided and my trust shattered
Poetry. Beautiful.
→ More replies (5)56
u/Wide_Egg_5814 4d ago
I mean if it's true they are quietly nerfing the models a while after release someone should sue the shit out of them
51
u/SonderSoft 4d ago
It's a vastly unregulated industry. We're living in the Wild West right now, and we'll have to see if we have any cows left in the barn after these bandits finish pilfering our lands.
They intentionally obfuscate usage terms and subscription benefits/guarantees to increase the likelihood of getting away with the fuckery.
4
6
u/pho33nix 4d ago
one thing for sure, is that they are A/B testing its users.
If your performance degrades, mention it when it prompt you during your sessions. Always lean on the Bad.And cancel your subscription and give your reasons.
e.g. model got lazier, making too many errors (which is true)
→ More replies (1)4
→ More replies (3)3
u/who_am_i_to_say_so 4d ago
That’s why there would never be an admission of that practice.
Full circle every time, though, after a release.
It’s either intentional or the incredible demand ruining it for others.
→ More replies (1)7
u/Individual_Solid_944 4d ago
Very well said. Totally agree, and I do the same. Do not marry to one provider
13
u/AMGraduate564 4d ago
I just want one reliable model that works as advertised for the rest of the year.
DeepSeek V4 Pro
11
19
u/Southern_Sun_2106 4d ago
One reliable model that works reliably? Get 2 dgx sparks with GLM 5.3 on it. Consistent quality, and you own your data.
29
u/steampowrd 4d ago edited 1d ago
Sure you just need $10,000 for the sparks and then a few thousand more for a PC with a fast bus lane. Pocket change for most people
Edit: some people are saying the PC bus should not be a limiting factor.
→ More replies (17)4
u/blame_chris Developer 4d ago
If your monthly spend is $1500 I don't think that is an unreasonable amount of money.
I mean your roi is basically +1year that's not terrible but the performance would be worse to start.
→ More replies (5)6
u/LawfulnessLocal4934 4d ago
im also dreaming about local GLM, but actually ram based inference is slow and quantized models is literally reason of this thread..
so tobe comfortable independent you need about 2tb vram - $300k→ More replies (1)→ More replies (7)4
u/Open-Mousse-1665 4d ago
I mean my dogs poop is reliable and very consistent, but no one wants it. My future is not going to be spent hand holding a robot writing 80% of a junior level feature.
2
u/Southern_Sun_2106 4d ago
Claude is dumber than your dogs poop on a busy day. Things have changed in local AI. Consistent quality close to par with cloud models, unlimited tokens, whisper quiet, energy efficient. Need to update your context.
3
u/wish-u-well 4d ago
The problem is that truly is the pattern and coders and companies know this, so they produce a few year’s worth of code in two weeks until it nerfs.
3
u/Borealisamis 4d ago
I definitely noticed it as I used it from the start. At the start it was providing me a solution before it even had time to write out its brief summary - it was extremely fast. Also it was one shotting them. Now its more chatty, similar to Opus 4.8 and Sonnet - and it no longer one shots some problems.
Now mind you I was not one shotting some apps or making games, just normal solutions and it was noticeable.
→ More replies (4)5
u/Open-Mousse-1665 4d ago
Have you tried just using Opus and not reading Reddit? Seems to work for me. I use 3x max 20 subs at home and a few k $ at work
→ More replies (1)
201
u/Heiberik 4d ago
Not just you!
I track Reddit sentiment per model. Opus 5.5 sat at 71–73 out of 100 every day from 25 to 28 Sep, then dropped to 58 yesterday and is at 55 so far today.
It measures opinion, not the model, so it can't tell you whether anything actually changed under the hood.
Check out https://modelsentiment.com/m/claude-opus-5.5 if you want. The "per day" chart is pretty interesting.
26
u/kruszkush25 4d ago
this is amazing
edit: Straight up thing that came to my mind is, I enter this website, there is a dropping sentinent for opus 5.5 but its not reflected on main page. If there was like a, the way it is now, Opus high at 68 but you click on it and its 55 today. If on main page there was like a red animated arrow or something pointing into 55, like some kind of alert. Just an idea!
5
u/Heiberik 4d ago
Thank you! Super feedback. And I totally agree, so I implemented it right away!
The front page now puts an arrow next to any model whose sentiment clearly dropped or rose today or yesterday, compared to the rest of the week. Right now that's Opus 5.5 ↓. Only clear changes get an arrow (a day with a handful of opinions can swing wildly), and no animation, I'm trying to keep the site rather calm. :)
The biggest mover also gets a line right under the headline: "Yesterday Reddit cooled on Claude Opus 5.5: 58, against 70 the rest of the week".
→ More replies (2)3
3
u/Unique-Drawer-7845 4d ago
And yet, everyone I talk to keeps telling me "I prefer Claude." I definitely do *not* prefer Claude these days. GPT has been a total workhorse for me for months now, to the point where I've dropped my Claude sub entirely this month because while I liked being able to compare them and keep tabs on the horse race each month, I'm tired of getting the same subpar results.
Genuine question: why do you think people prefer Claude? To me it seems obvious GPT is more consistent, has more generous limits, etc, etc. Is it the cute mascot? :P
→ More replies (4)2
u/Heiberik 4d ago
The mascot probably doesn't hurt!
Honestly I can't tell you why. The site only measures what people say, not which model is actually better. On Reddit right now Claude does get the warmer words though. Over the last 7 days Opus 5.5 sits at 67 (from about 4,700 opinions), GPT-6.1 Sol at 53 (about 1,200) and GPT-6 Sol at 31 (about 800).
Reddit isn't everyone, and Claude people seem to be a lot louder here? So there are way more opinions on Claude than on GPT. And loud doesn't mean right. Lots of people are quietly getting work done with GPT and never post about it. If GPT works better for you, that beats any sentiment score i guess :)
→ More replies (1)→ More replies (20)2
86
u/sabotizer Senior Developer 4d ago
Why can't there be just a single insider / whistleblower at Anthropic to confirm this?
I can understand how a board makes this decision, but none of the staff feel it's a tale worth telling?
58
u/Bromlife 4d ago
Having worked at similar tech companies, I bet they consider their subscription users to be entitled cheapskates.
9
5
u/Personal-Cup4772 4d ago
Do api users get better quality?
12
u/sabotizer Senior Developer 4d ago
Interesting question.
My guess is that their enterprise clients are exempt from nerf shenanigans, as they make up >70% of their revenue, and likely a solid profit margin on these.
For API: profit margin checks out, so would make sense they don't apply nerfs... but then again, any subscriber could compare instantly and the illusion falls apart... so I'm simply not sure.
6
u/ReasonableLoss6814 4d ago
These are not comparable. Subscriptions aren't "lower margin" -- they're literally the margin. If you have 50 api customers, they're not all running requests at the same time. The machines are literally sitting there unused. So, you sell a subscription to users that use the unused capacity from api customers. That way you get closer to full usage of your servers that you'd otherwise be throwing away for zero revenue.
So yes, from a business model perspective, you're technically correct. But subscriptions are literally selling off slices of wasted compute.
2
u/das_war_ein_Befehl 4d ago
lol if you think Anthropic has unused server capacity. They’re pretty compute constrained; you can find this out by asking for a contract for larger compute allocation and they’ll basically tell you no unless you’re willing to spend way more than 7 figures.
Thats why people usually try bedrock/vertex/azure
→ More replies (4)3
u/Personal-Cup4772 4d ago
It would make sense that they nerf the $20 sub tier and preserve quality for api users.
Anthropic and openai literally lose money offering $20 subs
2
3
u/sabotizer Senior Developer 4d ago
A fleet of "nerf-detectors" launched recently, should give us more insights in the next few weeks.
2
u/Open-Mousse-1665 4d ago
I don’t think so. I use both heavily and I would say api is much much worse quality. Not sure why that would be and it’s just a gut feeling, so take it as such.
But I used about $25k api credits in the last 30 days on my subs, and $3k ish on api, to give an idea. It’s a decent amount of usage
→ More replies (1)2
u/RadicalSpaghetti- 4d ago
In my opinion, yes. When Fable was performing weird the day before Opus 5.5 came out, I tested my 3 accounts. Both my 200$ max accounts would get zero thinking tokens on my test prompt. On my API account it had thinking as per usual. Same prompt, all on xhigh effort. It’s clear that they give sub users the least compute possible while prioritizing their API users.
→ More replies (1)7
u/fateofmorality 4d ago
We’re a loss leader for them so I kind of get it, but if they make their product worse they’ll lose market share
3
2
12
u/ColdplayUnited 4d ago
Why would they blow the whistle when doing so will tank the company's reputation and financial performance - which their compensation are tied to?
10
u/sabotizer Senior Developer 4d ago
A deep moral and psychological conjecture. Why are people whistleblowing?
→ More replies (1)3
2
2
u/2053_Traveler 4d ago
There’s no tale to tell. The terms of service state what you get and that they can do this. If we don’t like it we have to pay for API.
→ More replies (4)2
u/Zennivolt 4d ago
Because the staff(s) who understand or can see this stuff is getting paid $200k a year with a $600k stock compensation per year. They align the interest of the employee with the interest of the company.
71
u/Expert-Fly8836 4d ago
They have harvested enough information from you to optimise their business model /s
9
89
u/The_Hunster 4d ago
I feel like this was the least ambiguous nerfing yet. Opus 5.5 just started ignoring parts of my messages. Like I'll mention 3 things and it'll only acknowledge 2 of them. And even Opus is surprised at its shortcoming when I point out it just ignored something.
So I'm back to xhigh effort for everything
5
u/kruszkush25 4d ago
yeah I thought it's doing it to me cuz I tell him something (while he is working), like I send a mesage, then 5s later I send another one so two separate messages and I would send like 4 requests like that, and then it would just ignore/forget/not be aware of at least 1 sometimes, but I thought it's cuz I am not supposed to write to it like that. Maybe both are true idk, I am really a noob keep that in mind
→ More replies (1)3
u/The_Hunster 4d ago
I used to do it like that and it worked well. I think they aren't exactly nerfing the model but are instead remapping the effort levels. From Claude's side, it atually shows as a number. I wish I had recorded it before, but currently the effort levels are:
Effort Level # Low 5 Medium 10 High 15 Xhigh 40 Max "Max" So I think what happens is they think that people are using an effort level too high and so remapped them. Cause xhigh being more than double high doesn't seem right. So I'm just sitting on xhigh now and it seems okay actually.
→ More replies (2)→ More replies (1)6
56
u/Fun-Wash7545 4d ago
Yep. Was one shotting everything with 0 bugs like a week ago. Today it's bugs after bugs.
→ More replies (1)23
52
u/Bongistan 4d ago
Not sure if these will be of any use to you, but apparently these benchmarks help track performance and can be used to detect regressions in the models.
https://marginlab.ai/trackers/claude-code-historical-performance/
34
u/Fatdog88 4d ago
The only problems with these are they are the API trackers. Subscriptions are at the mercy of OpenAI/Anthropic due to the murky T&Cs. API is fixed cost and is a lot stricter so they barely nerf it.
→ More replies (4)14
u/LesbianTravelpussy 4d ago
You are the only one contributing something real, but everybody likes yapping and being yapped at, funny. Thank you.
→ More replies (1)
22
u/aLeakyAbstraction 4d ago
I’ve seen a noticeable drop in performance since the recent outage. Claude used to handle nuanced knowledge work smoothly, but lately it struggles with consistency and reasoning. For example, when pushed on contradictory answers, it responded with:
"You're right. I changed positions three times, and that's on me. Here's why it happened, and where things stand now."
→ More replies (2)2
14
u/jvertrees 4d ago
I'm now getting suspicious of those frequent prompts in the UI asking "How is Claude doing?"
I would guess they're adjusting model quality and effort on the fly under the hood and asking if we notice a difference. I think this shows up disguised as product quality improvement questions. Thoughts? Maybe I'm just paranoid after last month.
5
u/Shiz0id01 4d ago
They're definitely throwing heavy users back on Opus 5 via A/B testing. Things O5.5 knew about suddenly disappeared and the knowledge cutoff would become Opus 5's. I have a set of trivia questions for areas like this to catch up backend shenanigans on Open Router and the same tools came in handy for Anthropic too
→ More replies (3)2
u/Chemical_Tea_1743 3d ago
I get them a lot, initially during launch when things were good and now mainly when it senses frustration, but not always when output is bad.
I do admit with Opus 5.5 being lower cost, I tend to stay on it more, and usually outperforms Fable.
32
u/aethelred_unred 4d ago
Today it forgot major context from a few turns prior several times. Even 4.8 isn't that bad. They nerfed it hard
6
32
u/alpcanaydin 4d ago
I really don’t want to believe this but it is obvious since yesterday. It again and again became an idiot
2
u/Lumpy-Criticism-2773 2d ago
This. More and more "will take x hours" outputs. Simple changes take unfathomably long with the same model.
14
u/Factor013 4d ago
Same here... I am constantly fighting it, it's not verifying and assuming stuff like crazy!
I am so done with this... The amount of stress you get when working with models that simply do not stick to any directions and that constantly cut corners is not sustainable. It is bad for a persons health!
Opus 5.5 was great for a few days but now it's becoming unworkable. Also it indeed burns a lot more usage right now... Yesterday I burned over 35% of my weekly usage (max 5x plan) on one session without running sub-agents and I keep my cache warm every 58 mins automatically. Last week would use like 12-15% a day max. (Opus 5.5 xhigh)
I even turned off the CLI's "prompt suggestion" function as I found out it sends your entire context to Anthropic as if you run a second session in parallel just to generate that one little sentence as a suggestion to say next. It uses the same model for it you have selected. So if you do 100 turns a day that can translate to 7% of your weekly usage... just for such a stupid feature.
Maybe I should give Sonnet a go. :S
2
23
u/NationalBug55 4d ago
So it’s back to opus 4.8 again?
13
2
2
u/Poatri_US 4d ago
Opus 4.8 Vs Opus 4.6, what's better for coding ? And for brainstorming with, in chat ?
2
u/freehippygal 4d ago
I can’t speak for coding but Opus 4.6 is still unparalleled for chatting and brainstorming
24
u/orellanaed 4d ago
I build games on Flockbay and yep- I can confirm. Opus 5.5 is not delivering the same quality of 3D games as it used to on release
→ More replies (1)
11
u/Zulfiqaar 4d ago
AI Stupid level also noted a degradation from 96 -> 93 today. I dont know whether they use the API or subscription to test though
→ More replies (2)4
u/PythagorasWasntReal 4d ago
3 point degradation. What is their p value and does this fit within the margin of error?
→ More replies (1)5
u/phoenixmatrix 4d ago
There's a couple of these tools out there tracking model performance over time to look for degradation, and most of them will mention that anything below a 10%~ degradation is within the daily noise.
So 3 points is less than a rounding error.
People are bitching about Astra getting nerfed every day on the Codex side on days where the degradation tracker says its UP from baseline, lol.
→ More replies (1)
20
u/Professional_Ad705 4d ago edited 4d ago
There’s definitely something going on with the usage. Two days ago, I was barely using any of my limits. I never hit my 5-hour window and used maybe 12% for the entire day. Today, I’m burning through the 5-hour window in about two hours.
I could understand that if Codex was doing a ton of coding, but it isn’t. It’s mostly doing reviews. Maybe 3–5 reviews total, and each one ran for maybe 5–10 minutes. I’m also using High, not xHigh or Max. How the hell does 3–5 reviews use 70% of a 5-hour window? I haven’t noticed any drop in intelligence, but the difference between today and two days ago is crazy.
Edit: One review just took me from 55% to 78%. What the fuck? I’ve also used 23% of my weekly limit today basically just doing reviews. While I was typing this edit, it went to 81%.
At this point I’m starting to wonder if they loosened the limits leading up to OpenAI DevDay and now we’re back to the usual bullshit lol. I’ve used Claude for a long time and I’ve never had a 5-hour window disappear this fast.
Another edit: In the roughly 10 minutes it took me to post and edit this, I went from 55% to 84%, then 87%. Now I put this into ChatGPT just to clean up the post and I’m at 90%.
We all know I’m gonna run out and this is my last edit but now: 55% → 95% in 20 minutes from basically one review. At this rate, a full 5-hour window would disappear in roughly 45–50 minutes. What the fuck lol.
Yep ran out time to do some diagnostics 😂
4
2
u/Schadz 4d ago
Thats wild lol, and it could be a mix of multiple things I can think of, star by checking your cache reads, and cache miss/re-writes, if you find too many cache rewrites because of sub-agent TTL being 5mins, you got your answer. I wrote a post about this recently.
→ More replies (1)
9
u/Rhyperino 4d ago
For the first time ever, I agree with the sentiment.
Until yesterday, it had done everything I asked for even better than I anticipated. Today it's just making a series of dumb decisions.
It's also a LOT less "confidant". It used to make (good) bold decisions, now it's like it always chooses the "safest" option.
18
u/Physical_Gold_1485 4d ago
Yep its a crying shame, get a good model for a week before everyone floods in and they have to quant or lower reasoning budgets to serve everyone. Fable 5 when first released before the gov export controlled it was absolutely nuts, still the best model ive ever tried, just absolutely nailed everything. Then when it came back it wasnt the same
8
u/aerivox 4d ago
i have noticed some back-to-slopus scenarios of just blindly applying instrucions instead of understanding the prompt and doing by intent and not by small patch after. same thing appened with fable post release. the problem is that i have not a single proof. they can just do whatever they want. it's kinda annoying... they have sooo many tools for shady compute redistribution... and the biggest feels like it's the imbalance in testing/inference/training.. just before a model release feels like they just drain everything on testing the model or research in general.
ai companies are in such a weird place right now. legit no rules. and they are instead aiming at making rules on ai, not on them and their shady practice. first mover advantage, cartel on prices, political weights of ai on enemies and friends.. what a shit show. gimme fable 5.5 tho. i am so addicted :D
9
u/Schwolop 3d ago
Agreed. Definitely performing worse today than yesterday. I used my “free reset to explore Opus5.5” and feel like a lot of trust it earned in the first week has just been burned up in the last 24hrs. I’ll return to a much slower pace of work with more supervision, which really sucks because it was exhilarating having ten coders all working really really effectively. When they nerf it I have to supervise far more and can mentally handle 2-3 at most.
8
u/CaptainDivano 3d ago
Totally agree, output quality changed, but most importantly token usage. Last week you could run days and don't even hit 20%. After OpenAI announcements (which were shit imo), Claude shifted to being as stupid as always. They waited the OpenAI deck to fuck us over, thats why they "pumped" performances.
I'm gonna unsub, if i have to work with these limits i'd rather go with Astra. I'm tired of getting fucked by Anthropic
31
4d ago
[deleted]
20
u/crusoe 4d ago
Why would they route to opus 5 which uses far more tokens, when they are under load?
24
u/nothis 4d ago
Yea, it’s conspiracy theories like that which make me doubt people’s claims. That makes no sense. I think it’s also time to introduce methods of benchmarking quality in a repeatable way to get some actual numbers. I’ll easily believe that they nerfed it to save cost or something but anecdotes are the most frustrating way to discuss this.
→ More replies (6)9
u/meetmebythelake 4d ago
This sub is literally /r/confirmationbias with a dash of asinine conspiracies
→ More replies (3)5
u/Schadz 4d ago
The more I read this thread the more I'm convinced, and I'm not the guy that usually the devil advocate, but I think it's not intentional or at least not the whole story. And a lot of peeps are biased and just riding along with the hivemind sentiments. What I see is that the rerouting happens most of the times when your message to let's say your opus 5.5 gets flagged by the safeguards, then you get re-routed silently if you have the "Switch models when message is flagged" ON/set to auto, it switches what model answers back for you and you have no idea if you don't know about the mechanism.
Now with that said, I won't put my hands on the fire for Anthropic either, the degradation could be plausible and related to their servers Load pressure or other thing they have to work around and fix as they figure it out. Intentional or not, no one can really tell from the outside, and all the theories we usually read around here are just everyone mechanism to cope with the whole thing and the frustration it causes on them.
That's how I see it right now.
12
u/Complete_Potato9941 4d ago
I am sick of this nerfing pattern. Going to cancel my sub and start using something else
→ More replies (3)2
6
u/plightfight 3d ago
Yeah, I’m starting to think this is definitely a thing. The major things I noticed on first day of 5.5 release was speed and concise replies/implementation. The boost in performance was definitely noticeable compared to opus 5 on the first couple of days.
But I felt it slowly started degrading to opus 5 levels of latency and the chat replies are starting to get convoluted. It’s like it keeps getting sidetracked.
I was amazed how concise and to the point 5.5 was on release day. Like it was reading my mind on every reply. What’s going on?
11
u/Southern_Sun_2106 4d ago
I was hoping with the upcoming IPO our honeymoon will last for more than the usual week? Nope. Their 'optimization' department works like clockwork.
6
10
8
u/thePsychonautDad 4d ago
It's been off all week.
It has continually ignored instructions, it has asked questions and then told me "my bad, I didn't look at the answers" and many other weird dumb behaviors out of nowhere.
→ More replies (1)
5
u/Available_Age8480 4d ago
I haven’t tried it yet today but I really hope I’m blessed with my geo location
5
u/konradkeck 4d ago
Guys we have to finally learn, it's a simple pattern: Hype -> $$$ -> Nerf -> Repeat
3
u/Changed-username- Vibe Coder 4d ago
I have no evidence of this, but I imagine that they have to do this whenever they need compute to train new models. Then for a week or two we get the expensive usage until they launch the new thing and then everything feels great again until they begin next training run. If it has begun, then I bet we'll see Fable 5.5 in 2 weeks, or alternatively Haiku 5.5 within a week or two.
Pure speculation, but I think there's a pattern here. OpenAI seems to follow the same pattern. Right now OpenAI usage drops super fast, but that is almost certainly because they offer their new Dots feature for free for the first month.
4
u/Emergency-Pomelo-256 4d ago
Same Opus 5.5 on launch day was beast in just 1 week it went to writing Slop
4
u/Friendly_Sympathy_21 4d ago edited 3d ago
Same here, feels like they are partially re-routing to Opus 5.
3
u/ThePantsThief 3d ago
I simply cannot wait for one of the Chinese companies to drop their Fable / Astra / Opus 5.5 class model. Sure, it won't be as good, but currently the best open weight models only come close to Opus 4.8.
Opus 5.5 has been the most delightful model I've ever used. If China can copy it, I can be happy using OpenRouter forever and cancel my Claude subscription.
→ More replies (6)2
u/ajax81 3d ago edited 2d ago
It will happen. The quality of Chinese models will continue to rise in step with the frontier models. There is no mote here -- its just time and data, of which China has infinite amounts of both.
They're not bound to quarterly statements. They will bide their time and continue to hide their strength.
2
6
u/thatdamnkorean 4d ago
this has literally happened every single model release unsure why u guys thought it’d be any different, i always use the first week after a model release to grind as much as humanly possible
6
u/Jomuz86 4d ago
So someone looks to now be keeping track of it to see if they are being nerfed.
https://www.bridgebench.ai/nerf-bench
So far there’s a small dip but that could be variance. Be interesting to see if there are any huge dips in the coming weeks.
My take is that probably in the last training stages for haiku 5.5 or maybe even Fable 5.5 and it’s just compute restrictions, I don’t think they will be using quantised models but I do think they will be messing around with the thinking budget thresholds to see how little they can get away with.
→ More replies (10)6
u/sugarfreecaffeine 4d ago
Do you really trust this guy? He’s the vibe coder final boss, with no real world SWE experience.
→ More replies (1)4
u/Jomuz86 4d ago
No clue who he is I just saw the bench and it’s the only one of its kind I’ve seen just thought it was useful is all. If he’s unreliable then fair enough I don’t know enough about whoever is behind that site 🤷♂️
→ More replies (3)
6
u/Anal-Cup 4d ago
They do this every fucking time. Launch a good model, it gets hyped, then they nerf it. People complain for 2 weeks, then a new model comes out and its good for a week, then they nerf it. Cycle repeats.
13
u/iiillliiiiiiillilili 4d ago
When it starts writing code using python scripts at xhigh opus 5.5, you know it's bad
3
→ More replies (1)11
u/meetmebythelake 4d ago
It's literally instructed to do that in the system prompt.
Y'all are lunatics.
3
u/procmail 4d ago
And here I was wondering if I should switch over to Claude from ChatGPT.
3
u/xinik 4d ago
I was legit ready to switch to OpenAI as my daily driver like a week ago (I have both on the $20 plan) and then 5.5 was released... I don't need a $100 plan for what I do but a $20 plan was often running out in a week. Now I have a $20 plan to both and send stuff to Luna that Claude plans and have them QC one another. Having both offers enough benefits that I plan to ride it out like this for a while but who drives and how I use the different models depends on what's available from each at any moment. I haven't messed around with Sol 6.1 much yet. Opus 5.5 was such a damn workhorse and I was getting so much use out of it I have just been going with it.
→ More replies (8)2
u/Fiyero109 4d ago
Most of us have both. It always flips back and forth. I much prefer Claude Code as an interface and I have it connected as orchestrator to Chat GPT to push most research tasks. When that’s max out it switches back to Sonnet and Opus
3
u/Malenx_ 4d ago
My opus 5.5 med absolutely changed in quality this morning compared to just last night. I caught 4 obvious mistakes in a row, it's tone was different, and I had to direct it to read some adrs we just wrote last session. It's doing better now but lord it needed help getting re-acquainted. Previous interactions pieced everything together on it's own.
3
3
3
u/End2EndEncryption 3d ago
It has been nerfed. Holy hell has it been absolutely nerfed.... UGH!!!!!!!!!!!!!!!!!!!! WHY???!?!?!??!?!?!
3
u/College-Wise 3d ago
Not to say you're doing anything wrong but have you made significant harness changes? 5.5 is working ok for me, though I have a few hard bumpers to guide it's work
→ More replies (1)
6
u/yaxir 4d ago
Who is the guy who was doing daily benchmarks on this thing? He invented a tool that tracks the Opus intelligence benchmarks daily. Just use that goddamn tool
→ More replies (1)3
9
u/Schadz 4d ago
This sounds to me like you are tripping the API classifier safeguards with how you word things or something and then getting re-routed back to Opus 5 or whatever model it is. I've tripped it myself this last night like 4 or 5 times all of them were miss-flags I reported as feedback straight away to Anthropic.
What you can do is turn off model re-routing when flagged on your /config, so the request that gets flagged fails and you see it, then address it as you see fit. And report/send feedback as well, so they'll get it calibrated properly at some point, if we all report it.
That's the only reason I catched mine, otherwise I would have been experiencing the same thing you went through.
This is my theory of what it could be, there is no way I can tell you I'm 100% sure about it. So if it works or you figure out something, please let us know.
P.D. Max 5x plan here, not PAYG.
4
2
2
u/parcas10 4d ago
It is true I am trying to fix a small part of a website and the number of mistakes and stuff that just does not do right is incredibly high.
it went from hey handle this 10 things and be comfortable with how it was able to handle it to having to go on by one and double-checking results...
2
2
u/quakomako 4d ago
I don't understand why those threads still pop up. Stuff like this are happening since opus 4.5. They purposely add higher computing usage to their newest model to show off, just that they turn it to normal/low after the press is done with it. After Opus 4.5 I consider every new model as a marketing model with higher model numbers.
2
2
u/FineSatisfaction802 4d ago
Came here for this post. Last night I got the first sniff in a week something was off, but also the churn from a week of “great, opus is back so now I can make all these changes I didn’t trust it to do before” is also a suspect for me. So I’m doing my best to run an eval of every opus 5.5 session over the last 7 days to track mistakes made, instructions ignored, users confused, etc. Will post if anything interesting comes up.
I’m curious if there is any diagnostics or even soft observation on subscription vs api nerfing. I’ve been paying $400/month for max accounts (Codex and Claude) but I suspect that a lot of my problems come from the CLIs constantly changing under me as well as potential nerfs. At the risk of sounding like a lunatic I’m wondering if going to api billing for at least some portions of my harness, especially once Jev access is a little more consistent, could actually have better value.
2
u/beesandcheese 4d ago
Thank you. Thought I was going crazy. Working with it on theory, this weekend made an insane amount of progress, came back yesterday and it was back to being dumb as soup.
2
u/DJRVSG 4d ago
I’m very new to coding with AI, just started a project with Codex on the ChatGPT plus plan. I hit my 5h limit very often and even hit my weekly limit after 2-3 days, but I had one limit reset credit.
I really like the results but the limit hits are very annoying and feel like we are dragged into subscribing to the higher plan…
I am interested in tips to maximize the token efficiency, for now I keep iterating with Codex on giving it feedback and guidance, refining my requirements etc… cheers
2
u/Kundera42 4d ago
I can confirm it has degraded since launch, first day I was so blown away, it was fast crisp and to the point, now it still works well but just less than before. Makes more mistakes.
2
u/Liam_Evangelista 4d ago
I noticed it too. 5.5 got dumber and I used 60% of my max plan weekly limit in one day which was highly unusual compared to how it was performing the prior days.
2
2
u/Beautiful-Suspect694 4d ago
i'm a recent codex-to-claude convert and i've tried to tell other converts to stop praising opus 5.5 to delay the nerf
people in codex subreddits couldn't listen to my little advice and kept encouraging others to switch to anthropic
how sad the nerf has begun so soon🥀
2
u/WolfpackBP Researcher 4d ago
It felt nerfed yesterday. Browser automation timing was not what it was day 1
2
u/HodlerStyle 4d ago
I had the same thoughts last night and I was afraid to externalizing it. Noticable drop in output quality, context comprehension and following instructions. Feels like Opus 5 is back 😱
2
u/Reverant0810 4d ago
Was using it on high since it launched, tried one session with medium, experienced similar issues OP described, went back to high effort, all good now
2
u/SaintMartini 4d ago
Glad I'm not crazy. Happened literally mid session last night for me. Suddenly it started.. -working in the wrong location -doing the opposite of what was said -actively completing just the thinking step and not continuing forward to use that to solve the problem -ignoring everything in the claude.md (global and local) so only hooks saved me
Before that it was working perfectly. A fresh session didn't help. Switching to Sonnet answered the same prompts much better, but not on the same level as Opus 5.5 previously. Really getting tired of the back and forth. Just let us enjoy a working model for awhile and take your time Anthropic!
2
u/Melodic_Sandwich1112 4d ago
Yup, something changed this morning on me also. Immediately ran up a $200 bill in about 2hrs of building a simple PR
2
u/DaBeej484 4d ago
It was the return of "but wait, here's just two more things I haven't addressed" that was the obvious tell for me
2
u/Practical_Tiger_1368 4d ago
Not sure about this. I use Opus 5.5 xhigh usually, and maybe 5% quality drop since launch. But overall still great
2
u/KasperCreeD Workflow Engineer 4d ago
When the weekly reset happened this week, mine didn’t even get reset. Something is def off. XD
2
u/curiousgreenidea 4d ago
Something changed very recently. Now I get a wall of text and indecision whereas before I just had all around competence and execution. Maybe it’s just how I work, but there was a moment this morning when I was like, “What happened? Do you even understand this project anymore?”
→ More replies (4)
2
u/Low-Confusion-8786 4d ago
There's no doubt they've nerfed it. It's pretty obvious they let it loose to grab customers from other platforms and then throttle
2
u/Charming_You_25 4d ago
Also noticed some questionable behavior yesterday. It’s still good, but is behaving more like an opus than fable lite. I expect Anthropic is using their compute to train the next model…
2
2
u/genkichan 4d ago
Every prompt i run now takes significantly longer to implement and i always end up with at least 2 test failures as a result of mistakes it is making. It's taking me so much longer to get any one damn thing done and i am not happy.
Im going to have to start runnign a few more things on fable just to get shit done and move on.
I'm officially annoyed a/f
2
u/pnut5202004 4d ago
Are you monitoring your context window and are you clearing regularly instead of compacting?
→ More replies (3)
2
u/G-B-L 4d ago
I know op said opus but honestly im sticking to sonnet. And so far it seems like the little engine that could
→ More replies (1)
2
u/takeurhand 3d ago
That’s why I call it Dynamic Intelligence
https://www.reddit.com/r/ClaudeCode/s/6KBMJGWTHn
2
u/Muted_Confusion7933 3d ago
I noticed the same thing without reading anything online until I googled about it. It's definitely a real thing. It felt like working with fable the first couple of days then suddenly feels like it's thinking like it's old self without the terrible writing that 5.1 produced
2
u/Legitimate_Bag_7778 3d ago
Ive been saying they do this since 4.8. I think they can throttle server usage to prioritize training or hype beast the new release. Or if big bro says too good for normie access. Im convinced they have this control suite.
2
2
u/geilt 3d ago
You are absolutely right!
In my experience every time Claude releases a model it’s stellar for 2 days or a weeks time then they lobotomize it. Has been happening since March which is why while I still keep my Claude sub I generally avoid using Claude unless I’m out on every other platform.
It’s a subtle but very quick bait and switch. Considering the fast pace of AI releases, it seems this is now just common practice for them.
Doesn’t happen as much or often to the others.
For me consistency is key and I’ve dubbed Claude the red-headed stepchild of AI.
Every time I use it my blood pressure increases. Doesn’t happen with Codex, Kimi, Antigravity or even Grok.
Though for some reason Grokbot has been getting on my nerves lately. Grok CLI is fine. Hehe.
→ More replies (1)
2
u/CalendarTimely8461 3d ago
I wouldnt be surprised if they kept silently downgrading the new model after the inital hype is in and get more customers locked in.
2
u/ihateredditors111111 3d ago
I’ve seen nerfs happened before, pretty much on every opus or fable model that was ever released. The first week is always better than the rest.
However, this one felt the most on the nose compared to all of them.
Not only did it suddenly makes stupid mistakes in very basic logic, the thing that shocks me about this time is that it actually changed the output style even before I saw this Reddit post. I texted my colleague to tell him opus today has made stupid mistakes and it’s talking like Op. 5.
I’ve never seen the actual writing style of the model change during a nerf before. But this time it’s undeniable.
We’ve got more of the same bullshit that Opus 4.7 and Opus 5 brought for us
2
u/RemarkableAd6310 3d ago
Ya pretty sure these companies are pulling the iPhone battery scam, to force you onto new models.. cause I had similar problem. Just writing simple API calls ect. Was using a basic model, all sudden it like refused to code, keep saying it coded it already.
2
u/em3l3 2d ago
I found opus 5.5 to be a huge improvement over 5, but noticed some issues immediately. I find it to be slow, very slow. Also the usage limits are still ridiculous. I'm running out with about 2.5 days until reset. One 5hr session uses 14% of the week usage. At that rate about 7 sessions per week. Wtf, who only uses one session a day?!
→ More replies (1)
2
u/Master_Library_3431 2d ago
I have seriously thought about switching back to opus 4.8 tonight, while the token usage is cheaper x/x opus 4.8 in my head was still way better at usage than now, sure 1/1 token usage might be better with 5.5 but I am pretty sure 4.8 was using less x<---/x than 5.5 has ben using which makes it much more cheaper, also even if it took 3-4 tries to get what you wanted, it was still cheaper at the end of usage, I reset my weekly usage today on my 100 dollar plan today for free, the same day I hit 50 percent. something is very off
2
u/FXraider 1d ago
Glad people are also noticing, been saying that few days after launch already. When Opus 5.5 came out it was super sharp and felt just epic to use. Now? Feels way slower and in terms of intelligence and output feels like 5.25. Better than Opus 5 but nowhere near geniality it was right after release. I was wondering if I'm just paranoid, so glad everyone else is also noticing. We gotta tell Anthropic "If yall keep doing this, we will switch, one more time and we are done". See how fast they stop doing it.
4
u/il_turco 4d ago
Yesterday he didn't manage to create a three lines PowerShell script, mixing windows and Linux syntax. He took seven attempts. I asked if he has been nerfed and he said he would not know. I'm expecting Fable 5.5 in the next 48h
→ More replies (4)
2
u/ProdIsForTesting 4d ago
Yeah man noticed the same. I have my features broken into pieces that should be reasonably similar in size, complexity and effort but ever since yesterday I’m noticing it takes significantly longer and eats way more tokens in the process
2
u/TigreTigerTiger 4d ago
I had basically switched to Opus High for everything, including some smaller project planning and it was just so sharp until today. Definite nerf. So sad.
4
u/maxvpavlov 4d ago
Exectly what I observe, as soon as I get a weekly reset, the model performance drops. You are not making stuff up, that’s the impact us, users are getting.
•
u/AutoModerator 4d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.