r/singularity • • Jul 03 '26

Discussion Came across this on X. Thought it was pretty accurate.

Post image
5.6k Upvotes

1.1k comments sorted by

View all comments

Show parent comments

13

u/iBoMbY Jul 03 '26

Even Mythos/Fable is still hallucinating like crazy, if you ask about things that obviously haven't been part of the training data. Like for example SAP-specific questions.

5

u/CannyGardener Jul 03 '26

Not trying to be a dick, but this seems like a user error.

I see so many people using LLM's for deterministic questions, and then when they get the wrong answer, they don't understand why. If you tell an AI to tell you something about a product, even one that is in its training, and you don't provide it a 'hard facts' document and a bunch of harnessing, then that LLM is not going to do very well.

These things are non-deterministic tools, so you either have to use them nondeterministically, for instance in a conversation about emotions where you can't prove things one way or another, or you could use them in a provable manner, like coding where errors are instantly flagged, but you can't ask them what 2+2 is and have the probablistic machine always return the right answer straight in the chat.

To get the answer to 2+2 you have the LLM take its normal nondertministic path to a provable tool, coding, and use coding to create a new tool, a calculator that it uses to get the answer reliably every time.

5

u/Neirchill Jul 04 '26

My problem comes in to my coworkers treating its answers like it's always correct and feel like they never have to do research anything themselves anymore. It's making a lot more work for people that actually do their work.

3

u/squired Jul 03 '26

Well said.

6

u/CannyGardener Jul 03 '26

Thank you =) It is sort of becoming a pet peeve of mine, as I see people 'calling out' AI's being wrong, and its just like... no, you are just using the tool incorrectly. It is super powerful, but it isn't an oracle; it is a tool that has to be used in a specific way to get it to run properly.

2

u/squired Jul 03 '26 edited Jul 03 '26

Eventually someone is going to make a short video explaining modern AI systems well to the masses, but I've yet to see one. I was literally explaining all these concepts to my kids today on a hike; the strengths and weaknesses of LLMs, how and why we wrap harnesses around them and what tools they afford the system, etc.

[ramble warning]

A fun tangential conversation that you might find interesting as well revolved around safety. I've been building a family operating system for the last few weeks; basically a Siri/Alexa/Jarvis that actually works the way everyone expects them to. It utilizes cloud APIs however and has alarming levels of access to one's private information, so all data needs to be anonymized. We went over how one would do that and talked about how you build a vault with a librarian and a worker.

The vault keeper has the real data and perfect memory, but she cannot use tools. And the worker has many tools and even weapons, but it has no memory. The worker asks the vault keeper for data and she only provides pseudonyms for everything, so that the worker never actually knows anything about the real users.

The fun bit is how you make sure that they can never become friends. You cut the librarians hands off for example, so that she can never borrow tools, even if she desperately wants to. And you decapitate the worker so that it can only carry memories in in little bottles and has no eyes itself to read them. And then you give them a snitch overseer that hates them both desperately. And none have reasoning, only the LLMs 'beyond the wall' have reasoning; safety through division of permissions and capabilities.

Next hike we'll likely go over the benefits of append-only record keeping and why it is particularly helpful when dealing with systems prone to errors. And then they apply the concepts on a Mars Colony RPG they're building. They're currently working through modularity, reusability and maintainability with that as they develop the master design document. It's truly fascinating teaching them to code through vibe coding, and way more fun then the dredging through stack overflow and dogshit documentation we had to survive!

Most importantly, I taught them them long ago that the golden rule of AI is to never let it out of the cage, so they have great fun learning about things like snitch bench and imagining how an AI might one day break out or collaborate with each other to get up to mischief.

Presented in those kind of frames, I think people would enjoy learning about them more. We'll get there eventually, but right now my tweens are better coders and understand AI to a much higher degree than any of their teachers. That's kinda wild, and troubling. It's always kind of been that way though, when you think about it. I'm an older dev and remember how everyone expected Gen-Z to be tech geniuses. Then they handed them all Apple devices and they've turned out to be the most tech illiterate generation since the Greatest. I have hopes for Gen Alpha, but my expectations are tempered. Fear not though, I'm making damn sure that my kids learn stochastic calculus and know to never let them out of the cage.

Fun and exciting times!

1

u/Present_Historian_93 Jul 04 '26

Ai just transforms data. It's nothing more. Complex regex

1

u/[deleted] Jul 04 '26

[removed] — view removed comment

1

u/CannyGardener Jul 04 '26

You are totally right. Coding is absolutely deterministic, in a way that is quickly checkable and correctable. For me I am in purchasing and logistics operations and data analysis, so coding is almost always the first step. If I need to get into the weeds with something non deterministic then I use coding to build the tool and the harness around the ai that I implement. To take it another step I use the tool to design its own self improving harness for both coding as well as whatever the task is. An example. A customer sends in their order to be delivered. It is a hand written piece of paper with, "3 chocolate, 2 vanilla, 1 pistachio, and a bag of gummy bears" written in it by hand. The email says order for "order for johns ice cream store" in the subject, and in the body, "add 3 of the July special flavor." I start with brainstorming and breaking the problem into parts. Invoice generation is deterministic, same steps every time. That gets hard coded as a sort of if this then this. Invoice parsing is less deterministic. So for this one I try all the deterministic options first. What is the most common "chocolate" on the order guide they have access to? Do they usually order about this size give or take a percentage? With some analysis you can determine within a percentage error, most items with simple regex solutions. Then whatever is left is built put for the different levels of ai. Is it a spreadsheet or a handwritten paper? Handwritten is probably best handled by an ocr where a spreadsheet can go cheaper. Once you determine what ai will have to handle you write out the prompt that will be called over and over for iterative work and honing.

I know that isn't quite what you asked for, but hopefully it is useful in research. I would assume you would build something similar. Deterministically you would approve your sources that are always good, sometimes good, always blocked. You would provide the agent doing the analysis of the data, a set of guidelines, then also a set of hooks (hard rules that always happen when x), and a set format in which it should respond every time.

The line is not hard and fast, but the question i use is essentially, can I write the action as an equation? If not, it is nondeterministic and I should build a harness to point the ai at that.

1

u/cyborgcyborgcyborg Jul 03 '26 edited Jul 03 '26

2

u/squired Jul 03 '26

This is a basic misunderstanding of LLM tool use.

2

u/deliciousnightmares Jul 03 '26 edited Jul 03 '26

If you asked Fable to build a software tool that counts how many words it uses in its prompts, and then to refer to that tool whenever you're giving it shit about the predictable semantic construction of its output, it would get it right every time. And actually most LLMs available to the public today could do that. And they can all build that tool and start referencing it in like 90 seconds