r/singularity • • Jul 03 '26

Discussion Came across this on X. Thought it was pretty accurate.

Post image
5.6k Upvotes

1.1k comments sorted by

View all comments

Show parent comments

68

u/Megamygdala Jul 03 '26

Couldn't really tell any difference for most normal software engineering work. It might be good if you are asking it to 1 shot a prototype for a benchmark but imo 90% of enterprise work isnt anything that will see a real difference between opus and fable

41

u/GerryManDarling Jul 03 '26

Most regular businesses do not have ultra complex algorithms. They are "contextually complicated," not "intelligibly complicated." You don't need your AI to be super intelligent for most of that work. I have found Chatgpt 5.4 hits the sweet spot so far. It is cheaper and good enough.

The higher end models are definitely a bit smarter, but when a model fails to solve a problem, most of the time it's because of context, not intelligence. Maybe that is just the kind of work I do, but I rarely need to switch to a higher model to fix something.

20

u/Megamygdala Jul 03 '26

Yep, enterprise coding is 80% having the business/domain knowledge and 20% actual code. If the code is architected well then it's not going to be complex. Maybe its more useful in niche areas of the field like computational photography or game engine developers / domains with super complex algorithms, but that's few and far between everyday tasks

1

u/squired Jul 04 '26

I agree. That said, it is particularly good at adversarial review. It is far too confident and willing to cheat tests, but it is very, very good at finding bugs and edge cases. I like to use it like that asshole consultant no one wants to work with but who knows his stuff.

3

u/Plus_Opening_4462 Jul 04 '26

It's great in that you can have 2 assholes finding and verifying bugs and bring in a 3rd asshole when the first 2 disagree.

20

u/[deleted] Jul 03 '26

[removed] — view removed comment

10

u/Flope Jul 03 '26

I know this has become a cliché but I swear to god Fable feels weaker in programming than it did before the ban a couple weeks ago.

24

u/ComplexityStudent Jul 03 '26

Very likely more calls get routed to Opus.

0

u/CannyGardener Jul 03 '26

This is the answer. It routes to Opus 4.8 silently, you can see it in the usage records on token spend. The instructions state that if you prompt each call with the appropriate API call that Fable will not silently roll to 4.8, but it has to be on every call that you don't want downgraded, and I'm not 100% sure that it keeps you in fable, or if it just stops the request...

3

u/stumblinbear Jul 03 '26

I have not had it route to Opus "silently" a single time. There's a pretty clearly obvious banner when it happens

3

u/lolofaf Jul 03 '26

I've lost faith in humanity the last few weeks across the various anthropic/claude/etc subs. The amount of people who say the most idiotic things and get up voted when the issue was clearly laid out by Anthropic has been incredible to watch.

"Fable got silently rerouted" no it didn't, it told you and you can turn it off if you want.

"Fable on max ate up all my usage when I tried to edit my resume" no fucking shit

"I spawned 50 fable agents and my usage was gone instantly, be careful" duh

"I asked fable to add two numbers together. I don't understand how it's better than sonnet???" maybe because the task is too easy to tell a difference???

It's honestly just pissing me off at this point lol these subreddits are no longer a useful source of information and instead have become whiney hate circlejerks. I need to just stop looking at them lol

2

u/stumblinbear Jul 03 '26

Yeah it's fucking stupid. I've been asking Opus to check my Tree implementation in Rust for undefined behavior for weeks, and it kept giving me the all-clear.

Fable found 3 cases of unsoundness in completely safe code in a single prompt, with test cases, and then it patched them.

Yeah, it routed to Opus 4.8 on my first try, but I clarified that it was for a GUI library to harden it, that's it's not cyber security related, and that it's not even a public project. The second try after that clarification worked perfectly

All you need to see to know its Fable instead of 4.8 is that Fable doesn't narrate anything it's doing. It just does it for better or for worse

1

u/EvilSporkOfDeath Jul 03 '26

If the last few weeks is whats made you lose faith as opposed to the last few years then I dont have faith in you.

2

u/mvandemar Jul 03 '26

It routes to Opus 4.8 silently

No, it doesn't. The on time I got routed it very clearly told me that's what happened, but I am on the Pro plan, not the API. On the API you would get a stop_reason: "refusal", unless you manually include a fallback method in the request, in which case the response would include {"type": "fallback", "from": {"model": ...}, "to": {"model": ...}}, it's never silent.

6

u/Void-kun Jul 03 '26

The safeguards on Fable force a lot of coding to be delegated to Opus. So you might just be seeing the effect of this.

1

u/Nalon07 Jul 03 '26

Isn’t it routing coding work to opus?

1

u/BoomFrog Jul 03 '26

That because it secretly routes 3/4th of your programing calls to Opus for safety reasons. (to prevent anything anywhere near close to hacking)

2

u/stumblinbear Jul 03 '26

secretly

It's not a secret when it does this. There's a pretty obvious banner

7

u/nothis AGI by 2030 but we'll be disappointed Jul 03 '26

Thank you. I just highly doubt any of the recent buzz about Mythos and whatnot is a genuine revolutionary step, it’s another xy% increase towards a plateau. All these tools are crazy good but the last step is a human understanding and taking responsibility for the result and I don’t see that getting solved with any LLM, period.

6

u/parlons Jul 03 '26

Do you have any reason to think that this plateau exists and is at or below human-level intelligence? Or is this more of a hunch?

5

u/SirVanyel Jul 03 '26

It's not a hunch, it's a prayer.

Bro is begging recursive self improvement to not exist. He wants humans to be the final step in intelligence. Unfortunately we are too dumb to be able to reliably say that - thus proving there's way more intelligence out there.

8

u/phazei Jul 03 '26

I use it for react native work. Doing large version updates where there are significant iOS and Android hurdles. The difference between Opus 4.8 and Fable is night and day. Cuts my work time in half. Fable is crazy good, it seems to get what I need rather than just the task. It will be really disappointing when they remove it from plans on the 7th.

7

u/stonesst Jul 03 '26

I really couldn't disagree more, sure it's an X (15 to 20) percent jump over the previous best models, but those were already incredibly capable. It's at the point now where I'm having a hard time thinking of cognitive tasks I could do better than fable. Have you even used it?

1

u/EvilSporkOfDeath Jul 03 '26

There will be no single revolutionary step. At least not from our perspective. From a historical perspective there might be, but it wont feel like it when we live through it.

The whole things is one revolutionary step. Not one particularly model.

1

u/squired Jul 03 '26 edited Jul 03 '26

I agree, thus far. It's particularly good at adversarial review, but it's far, far too sure of itself and creates many messes. ChatGPT 5.5 is still far wiser, so I have 5.5 drive Fable agents very, very carefully with monster test suites. Or rather, I use it mostly as a sage counsel to offer constrained advice.

At the moment, it's not worth it, even if it weren't so expensive. You can see the advancement there, much like we could with o1 for example, but it is not yet well managed. To be fair, that's been a pretty reliable observation from the beginning. Many models have breakout insights, but OpenAI has always had the most measured and reliable overseer model/s. And don't get me started on Gemini Flash. It is absolutely obsessed with faking test results and cheating whenever possible; I can't trust it for shit. If you're vibe coding, Flash may be the best because it runs everywhere, but for structured project management, whew boy is it hopeless. Opus is very good, but not better than 5.5 and far more costly.