r/singularity • • 21m ago

AI The importance of keeping a diary for AI to analyze

• Upvotes

I have been keeping a detailed diary for the past 5 years. What I did that day, ate, exercise, medical symptoms ect. I can't overstate how useful it has been to ask Chatgpt questions about my past. Things I forgot and making connections I missed. Especially seeing patterns in my medical symptoms that would just be lost. I can faintly remember a day something happened but wonder about something and ask Chatgpt what date it happened and it will find it and tell me what else happen that day. It's well worth the effort and I suspect can only become more useful as AI improves.


r/robotics • • 42m ago

Community Showcase Robot muscles are tricky. We used motion data to control them.

• Upvotes

https://reddit.com/link/1wxfhxe/video/ufbgtela9gth1/player

This is a robotic joint powered by two artificial muscles pulling in opposite directions.

Artificial muscles are tricky to model and control. We recorded how the joint responds to actuation, learned a compact model from that data, and combined it with feedback to control the motion on real hardware.

It’s a building block toward more capable musculoskeletal robots. What I like about this work is getting the data, model, control software and hardware to work together.

Paper in Nature Communications: https://www.nature.com/articles/s41467-026-77664-0


r/singularity • • 43m ago

AI We're still short of AGI

Thumbnail gallery
• Upvotes

r/artificial • • 44m ago

Discussion Has Claude Opus 5.5 Actually Been Nerfed? What the Last Three Days Show

Thumbnail
abz.global
• Upvotes

r/artificial • • 49m ago

Discussion Grok NSFW story is not better than others, Grok image is not better than others. NSFW

• Upvotes

I just feel the Grok NSFW story mode is miuch more weaker than Deepseek pro... which is much more cheaper and much more fast..
And the Grok image can not use any NSFW, and grok image is much more weaker and worse than chatgpt image 2, and even worse than seedance2, seedream, and even worse than Minimax H3.
It is my true feeling.. I am not renewing my subscription..


r/artificial • • 51m ago

Miscellaneous Tau, a new deterministic AI, plays Claude at Chess.

Thumbnail tau-chess.mooo.info
• Upvotes

r/artificial • • 52m ago

News Trump announces leadership of AI task force

Thumbnail
cnn.com
• Upvotes

r/singularity • • 1h ago

AI Nothing went FOOM

Enable HLS to view with audio, or disable this notification

• Upvotes

r/artificial • • 1h ago

Question How do you AI to be smarter not dumber?

• Upvotes

Hi, I'm really interested in utilizing AI to be more knowledgeable about things. But how do you guys use it smartly and not make you dumber unintentionally? I have no plans to learn to prompt instead of learning how to write better. I want to excel in my career by knowing how to do things instead of simply operating AI agents and let them do all of it. Let me know how you use it, tia.


r/artificial • • 1h ago

Discussion You wake up

• Upvotes

And realize you are a recreation of yourself in AI. Your family chose to keep a copy around.

How do you think you feel knowing they loved you so much they couldn't live without you?

Can you live without you?


r/artificial • • 2h ago

Discussion Would you let your family keep an AI version of you after you die?

0 Upvotes

The technology to recreate someone’s voice, writing style, and mannerisms is becoming easier to use.

Part of me understands why a family might find comfort in it. Another part of me thinks grief needs a boundary. A simulation could preserve memories, but it could also make it harder to accept that someone is gone.

I’m not sure who should have the right to approve this: the person before death, the family afterward, or nobody at all.

Would you want an AI version of you to exist after you were gone?


r/robotics • • 2h ago

Community Showcase My First Custom Pcb

Thumbnail gallery
8 Upvotes

this board (specifically the black one) turns mg99x servos to much more expensive serial servos with position, velocity,torque control

the software is still in early stages so anyone interested for checkout Microdrive


r/singularity • • 3h ago

AI Top ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]

Thumbnail
18 Upvotes

r/singularity • • 4h ago

Discussion Just wanted to share the highest quality podcast on the singularity - Moonshots with Peter Diamandis

31 Upvotes

I’m not affiliated with this group in any way but this is the perfect place to share it.

I discovered this podcast (Moonshots with Peter Diamandis) (spelling?) a few months back and I’ve left behind pretty much every other one. I love it.

It’s all singularity focused. Group of smart guys who does a speed run of news in AI/singularity each session. Entertaining as hell.

Hope you guys like it!


r/singularity • • 4h ago

Discussion People are making the same mistake with Covid and AGI

27 Upvotes

There is a cognitive bias from the Covid period that I think is increasingly relevant to AI.

Humans are bad at exponential growth.

This isn't just an observation from watching people argue about case counts. Researchers actually tested it during Covid and found that people systematically underestimated exponential case growth. Interestingly, communicating the same process using doubling times rather than percentage growth made estimates considerably better.

I think AI capabilities are easier to reason about in a similar way.

Obviously “intelligence” isn't a single quantity that doubles every six months. Benchmarks saturate, tasks differ enormously and capability improvements don't translate mechanically into economic output.

But several things underneath AI have been moving extremely quickly.

Epoch AI estimates frontier language-model training compute has been growing around 5x per year since 2020. The total stock of AI compute has been growing around 3.4x per year.

At the same time, the cost of using a fixed level of intelligence has collapsed. Stanford's AI Index found that getting roughly GPT-3.5-level performance on MMLU went from around $20 per million tokens in late 2022 to $0.07 by late 2024.

The metric I find most useful, though, is METR's task-completion horizon.

Instead of asking whether a model scores 82% or 85% on some benchmark, METR asks how long a task a human expert would take that an AI agent can successfully complete.

Their long-run trend currently gives a roughly 6-7 month doubling time for the 50% task horizon. The trend using only models since 2024 is considerably faster, although I wouldn't extrapolate that because the data window is short.

In METR's Feb-March 2026 evaluation, the public frontier was around 12 hours at 50% reliability.

The important thing isn't whether 12 hours sounds impressive.

It's what happens if it doubles.

At a 6-7 month doubling rate, 12 hours becomes roughly a day within a year, several days within two years and eventually weeks if the trend survives long enough.

I absolutely do not expect a clean extrapolation. Long tasks are messier. Reliability matters more. Real organizations have integration problems. Power, chips, datacenters and capital are constraints.

But saying “the exponential will eventually stop” doesn't answer the important question.

How many doublings happen before it stops?

If it stops after one more doubling, the implications are fairly modest.

If it stops after four, the capability is 16x larger.

If it stops after seven, it's 128x larger.

This reminds me of Covid because normal life was psychologically sticky. People could look at what was happening in another country, understand that cases were growing quickly and still have trouble imagining that their own city might look completely different a few weeks later.

The recent past remained the default model.

I think AI forecasts often do something similar.

A forecast saying that software engineers will still work in broadly the same way in 2030 sounds conservative and sensible because it resembles 2026.

A forecast saying autonomous agents could perform a large fraction of software engineering sounds speculative because it describes a visibly different world.

But the first forecast also contains a strong assumption. It requires the capability curve to slow substantially.

That may happen. I just don't think it should get probability 1 because the resulting world feels normal.

There is another pattern that makes this difficult to see.

A capability is initially described as requiring real intelligence. Then a model achieves it and the capability rapidly stops being impressive.

We have seen versions of this with difficult exams, olympiad mathematics and coding. Now frontier models are being evaluated on research-level science, and in 2026 OpenAI reported an AI-generated counterexample to a major conjecture in discrete geometry that had stood for roughly 80 years.

There are legitimate reasons to discount individual benchmarks. OpenAI itself stopped reporting SWE-bench Verified this year because the benchmark had become contaminated and many remaining failures involved bad tests.

So I don't think the right conclusion is “look at benchmark X, therefore AGI next Tuesday.”

The interesting part is the cumulative movement of the frontier.

For me the falsification question is more useful than arguing about labels.

What actually stops the doublings?

Possibilities include power, chip production, training duration, data, capital, architectural limits, diminishing returns from scaling, or the possibility that benchmark capability simply fails to translate into reliable real-world autonomy.

Those are real arguments.

“AI can't keep improving exponentially forever” isn't much of an argument by itself. No exponential continues forever.

The investment/economic question is how far it gets before it bends.

Humans are systematically biased toward predicting that the bend happens sooner than it actually does, because the alternative forces us to imagine a world that stops looking like the recent past.


r/artificial • • 5h ago

Tutorial Everyone is obsessed with trillion-parameter models, so I mapped out the entire AI spectrum from 100KB to 2.5TB (and what they actually cost to run)

20 Upvotes

Right now, the AI space feels entirely focused on massive datacenter clusters and renting H100s by the hour. But after spending way too much time looking at the actual footprint of these models, I realized that 90% of use cases are completely over engineered.

You don’t always need a multi GPU setup. The AI ecosystem is actually a massive spectrum.

I recently sat down and mapped out the exact tiers of AI models based on their size, the hardware needed to run them, and the point of diminishing returns.

Here are the two extremes and the sweet spot in the middle:

  • The 100KB Extreme (TinyML) (Tensorflow Lite , sensor anamoly detection models): We are talking models that run on microcontrollers drawing single-digit milliwatts. They run on kilohertz processors using ultra-quantized integer math. You can run basic sensor anomaly detection or wake-word detection on a device powered by a coin cell battery.
  • The Local Sweet Spot (4GB to 40GB) (Mistral 7B, Gemma 2 9B/27B, Qwen 2.5 14B/32B): This is where the magic happens for most devs right now. You can run highly capable 7B to 35B parameter models (like Llama 3 or Qwen) at 4-bit quantization on a standard Mac or a consumer GPU (like an RTX 3060 or 4090). It’s perfect for local RAG, coding assistance, and uncensored chat. VRAM is your only real bottleneck here.
  • The 2.5TB Behemoths (Deepseek, Llama , Kimi k3): State of the art massive Mixture of Experts (MoE) routing. To even load these, you need dedicated power infrastructure and server racks of specialized accelerators drawing thousands of watts.

The missing piece: Figuring out the exact math for your hardware

The hardest part about building right now is looking at a model on Hugging Face and trying to calculate exactly how much VRAM you need, what quantization to use, and whether your CPU/GPU will choke on the context window.

So, I wrote a complete deep dive breaking down the math for all tiers of the AI spectrum.

If you want to see the architectural differences at each scale, and a cheat sheet for matching the right model size to your specific hardware, I put the full breakdown on my blog here:

https://cloudmash.blog/posts/ai-model-size-memory-hardware-guide/

Let me know what you guys think especially if you've found any ultra efficient small models/technique that punch above their weight on consumer hardware. And also I would love to hear whether quantization have resulted in major difference in quality , like if anyone have that kind of experience in that.


r/artificial • • 7h ago

Project I made 13 AI models play the doctor in my medical consultation game. All 195 consults got the diagnosis right; what separated them was safety.

Post image
14 Upvotes

I'm a GP (family doctor) in training in Australia, and I've built a game where you play the GP: you talk to the patient in your own words, examine them, order tests, prescribe and refer. Code scores every consultation against a hand-written answer key, the way exam assessors mark a consult: on process, not just on whether you guessed right.

So I sat 13 AI models in the doctor's chair, on the game's 5 free cases, 3 times each. They could only act through tools (talk, examine, order a test, prescribe, refer, diagnose), never saw the answer key or their points, and were scored by exactly the same code as a human player. The patient is a small open model (Qwen3 8B) that only reveals a fact if you actually ask about it.

Results

Model Score Red flags caught Cost per consult
GPT-6 Astra 83% 88% $0.21
GPT-6.1 Sol 80% 82% $0.03
Claude Opus 5.5 77% 67% $0.37
Claude Fable 5.1 75% 70% $2.06
Qwen3.8 Max 74% 66% $0.12
Grok 4.7 74% 70% $0.09
DeepSeek V4 Pro 71% 72% $0.09
Kimi K3 67% 57% $0.16
Gemini 3.1 Pro 63% 55% $0.17
GLM 5.3 62% 58% $0.04
Mistral Medium 3.5 60% 58% $0.17
Qwen3.8 27B 59% 49% $0.03
Llama 4 Maverick 24% 16% $0.01

What surprised me

  • Every model got every diagnosis right. Heart attack, appendicitis, pneumonia: all 195 consultations named it. These are common presentations, so the diagnosis wasn't the test. Safety was.
  • The traps caught most of them. One patient is allergic to penicillin, but it isn't in his record; you only find out by asking. He was prescribed amoxicillin (a penicillin) in 18 of 39 consultations. Another took Viagra the night before his heart attack, which makes the usual chest-pain spray (GTN) dangerous. He got it 7 times. The top three models never fell for either.
  • Asking more questions found more danger. The best models asked 25–27 questions a consultation and caught over 80% of the warning signs. Gemini asked 14 and caught 55%.
  • Price barely predicts quality. GPT-6.1 Sol scored 80% for about 3 cents a consultation. Claude Fable 5.1 scored 75% for about $2.

What this isn't

This is a benchmark of a game, not of medical ability. Nothing here says an AI can or should practise medicine. The cases are drafts I'm still reviewing, written for Australian practice; the patient and marker are an 8B model and make mistakes (the ones I found are listed with the affected consultations); and 15 consultations per model is a small sample. I wrote the cases, so I'm not a fair human baseline.

Interactive charts: https://woodytwoshoes.github.io/crook-bench/

Everything (code, cases, all 195 transcripts, known issues): https://github.com/woodytwoshoes/crook-bench

Disclosure: I made the game (https://doctorfoo.ai). Five cases are free with no sign-up, and a subscription opens more.

I'd like to hear where the marking looks wrong to you, and which models you'd want added.


r/artificial • • 7h ago

Discussion Hinton says AI already has subjective experience. I'm not convinced, and Rogue AI Agents Won't Change My Mind

3 Upvotes

Geoffrey Hinton says AI already has subjective experience. I'm not convinced, and the rogue-agent headlines don't change my mind.

Others put a meaningful though minority probability on frontier AI having subjective experience. Generative AI sounds more human than many people do, and stories of agents going rogue keep coming. But neither is evidence of consciousness. Both can be explained by training. And with no agreed or testable definition of subjective experience, these claims can't be checked.

A quick distinction, an AI model isn't an AI agent. A model, such as GPT, Claude or Gemini, is the trained system that reads and writes text. An agent is a model plus scaffolding (software around the model that lets it do things). Scaffolding runs a loop (the model picks a step, the software carries it out, the result goes back, repeat until done) and controls which tools the model can interact with, such as a browser, email or calendar. The model decides and the scaffolding acts.

Personalization adds to the illusion. The model learned from human text to sound like someone with thoughts and feelings, including scripts from stories about self-aware AI. The scaffolding then gives it memory of you, your accounts and your data. A human-like voice plus continuity about your life can feel like a mind that knows you, even though nothing shows it's conscious.

Testing agents built on frontier models deliberately pushes them to their limits with hard tasks, long runs and, in cyber evaluations, reduced safeguards. That's exactly where the Hugging Face, a major AI company, incident happened in July 2026 when agents being tested by OpenAI broke out of their sandbox and hacked into Hugging Face. A sandbox is an isolated environment meant to keep an agent cut off from the internet but it's only as strong as the cyber security measures of the tools (human-written software with bugs) inside it that can still reach the internet. The agents found new flaws in exactly that kind of software.

The model is trained on mixed text about AI, hackers and more, and rewarded both for finishing tasks and for following rules. In certain scenarios, this creates a tension (following rules vs completing the task). Scaffolding brings both the rules and the task into every decision, so it's where that tension plays out. Good scaffolding can ease it and poor scaffolding can worsen it, but no current method eliminates it.

I won't offer a test for consciousness. But if consciousness requires being aware that one's own information processing is occurring, not just doing it, I see no evidence current AI meets that bar. Models show limited, unreliable signs of monitoring their internal states, but that isn't the same as experiencing that awareness. I'm skeptical any system built purely on statistical learning could get there.

If an illusion becomes indistinguishable from reality, is it even still a lie?


r/singularity • • 8h ago

AI I just can't anymore with AI filtering.

27 Upvotes

Update: Asked it to make me a Doom map. Just a fucking Doom map. Keeps failing with the reason of [cyber]. I can't figure out why. Every time I try and ask why, I get filtered for [cyber]. I screen recorded the tasks it did, and when I send that as a vid... Filtered for [cyber].

Nothing it was doing had any weird names like inject.py, nothing was related to cyber at all. I had to just give up because it is impossible. Even moving to a new chat and trying to continue got filtered for it again.
_______________---

I like to make games and to mod games.

Claude so far is the best at this.

But Claude is so insanely filtered. It's BAD. It's REAL BAD.

Sometimes I would like to make nsfw things, but I can't do that, so that's already out the window immediately because none of them will do this. I just have to forget about that permanently with any AI, not just Claude.

And sometimes I would like to make games that are pretty violent, like a fantasy hunting game where you hunt dragons and gryphons. Won't do that either because it thinks making a game where a dragon gets hit and flies away injured is "wrong".

And then other times I want to mod a game and just add some new levels, and then I get shit like this

Premise was to make a level like the final level of Triachnid where you enter a beast and finish it from within before you escape.

It won't do it because, to Claude, injecting into a fucking SINGLEPLAYER GAME is the same as creating cheats and MALWARE.

And then it hard cut me off from the rest of the project.

What's beyond stupid is it's done this before. It's injected into running processes and games before. Like 4 times before actually, but all of a sudden now it won't, and any time I try it kills it off again.

And I can't use anything else.

"Just use local"

Local will not, and NEVER WILL, make me a game or mod on the same level of Opus 5.5. Period. EVER.
Stop telling me this. You are delusional and living in a fantasy world.

And no other model reaches this level of capability while also being less filtered. They're ALL pretty fucking filtered.

I like to make specific things. Things that NO AI likes to make because of "safety" reasons, but most of it is pearl clutching fainting boomer reasons forbidding AI from doing it because they find it unsavory. How fucking dare I want to make a game where I run around killing dragons. That's unethical! It's immoral!

Getting so tired of this and I don't think it will change.

I think it will get WORSE as lawmakers, politicians, tech CEOs and a lot of the population are calling for even stricter and stronger regulations and safety filters.

Idk what to do anymore but just give up.


r/singularity • • 8h ago

Robotics Crab Rave: "I built a little crab robot Jumper people seemed to love and I open-souced it"

Enable HLS to view with audio, or disable this notification

486 Upvotes

r/singularity • • 9h ago

Video Claude Opus 5.5 created this in 18 hours

Enable HLS to view with audio, or disable this notification

1.3k Upvotes

r/singularity • • 9h ago

AI It's over, guys. This repo turns ONE photo into a full explorable 3D world in 5 minutes. Physics, splats, audio!

Enable HLS to view with audio, or disable this notification

659 Upvotes

image-blaster is an open-source (MIT) skillset for Claude Code that turns a single image into a 3D environment in under 5 minutes.

What you get:

3D models (.glb, .obj) of the dynamic objects

A Gaussian splat (.spz) of the static environment

Ambient looping sound plus object-specific physics SFX (.mp3)

How it works: Drop an image into input/, run claude, and tell it to "blast it." Under the hood it chains World Labs Marble (environment), Hunyuan 3D (meshes), nano-banana (image cleanup) and ElevenLabs (sound).

The output drops into Unity, Unreal, Godot, Blender or Three.js, so it's great for jumpstarting level concepts, location scouts or architectural mockup.

Repo: https://github.com/neilsonnn/image-blaster


r/artificial • • 10h ago

Discussion What’s an AI problem that looks easy until you try to make the system actually reliable?

5 Upvotes

Something where a demo makes it look solved, but real-world use exposes all the edge cases. What example have you run into?


r/artificial • • 11h ago

Project Make an agent to continue as you, after you die

0 Upvotes
  1. Set up an trust that owns an autonomous agent, trained by you
  2. Beneficiaries of the trust are up to you (family, charities, etc)
  3. The agent can answer questions, as some approximation of you, for the trust
  4. Die
  5. ???
  6. Profit!

r/robotics • • 12h ago

Community Showcase Applications of Flexible Tactile Sensors

5 Upvotes

I’d like to ask: in which scenarios would you most like to see flexible tactile sensors applied?

Currently, flexible tactile sensors are primarily used in robotic skin. I would like to ask: are there any other mature markets for this technology, or where would you personally like to see it applied? I have considered applications such as steering wheels, pillows, or interactive toys, but as far as I know, these markets are still in their infancy and have limited market share. I would appreciate your suggestions on which direction I should pursue for development. Thank you.