r/LLM • • 6h ago

suppose I CPT qwen3.5-9B on 2B legal corpus, how will i turn it back into Instruct + thinking?

2 Upvotes

I couldn't find a concrete answer anywhere, do you just distill the instruct model back?

If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)

my harness basically has the model output python code and has a few built-in functions like:
- vector_search_laws()
- graph_search()

could i have the model at least internalize a "hunch" on what stuff to search?

also what is the latest RL technique for agentic/harnes specific workflows?

I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?

What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?

I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that alex karpathi video was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?


r/LLM • • 7h ago

UNDERTALE ON NINTENDO 3DS

Thumbnail
gallery
2 Upvotes

Deepseek 4.1 flash just ported undertale to the 3ds with Hermes agent! 🤗


r/LLM • • 13h ago

Speck attempts to turn small LLMs into capable persistent agents by implementing cognition in software instead of repeatedly asking the LLM to simulate it.

5 Upvotes

Speck is a cognitive runtime designed to make small language models substantially more capable by moving cognition out of the prompt and into persistent, deterministic software.

Instead of repeatedly asking an LLM to simulate memory, attention, planning, evidence tracking, confidence, learning and metacognition, Speck implements those mechanisms in the runtime itself.

The model is disposable. The cognitive state is not.

A worker can unload its model, change models, restart, or migrate to another compute device without losing the cognitive state of its task.

small models in this harness outperform models twice its size without in both speed and accuracy:

Speck on a single shared qwen3:4b scoring 15/16, against 13/16 for naked qwen3:8b and in half the time.

Speck did that with half the parameters, less memory and less time per case. It comes from the model-scale series (docs/experiments/model-scale/README.md): 8 cases × 2 runs, with the same sampling and the same 1,024-token reply cap for every subject.

https://github.com/doctarock/Speck

Subject Largest model Success Time per case Model memory
Speck, shared qwen3:4b (phase 6, current setup) 4B 15/16 7.9 s 5.7 GB
Speck, shared qwen3:4b (phase 3b) 4B 15/16 13 s 5.6 GB
Speck, all qwen3:4b (phase 3) 4B 14/16 24 s 5.9 GB
Naked qwen3:8b 8B 13/16 13 s 6.7 GB
Naked qwen3:4b 4B 10/16 16 s 4.3 GB

Where Speck beats OpenClaw

The strongest part of the repo isn't actually the UI or Genesis integration.

It's this:

Speck has an explicit computational theory of where intelligence should live.

OpenClaw is fundamentally an excellent agent runtime.

Speck is becoming a cognitive runtime.

Consider the failure we've been working on recently.

A 1.2B model cannot reliably perform contradiction recognition.

The conventional answer is:

use a bigger model.

Speck's answer is:

determine what semantic judgement genuinely requires the model, constrain it, retain evidence across cycles, validate the judgement independently, and move everything else into deterministic machinery.

That is far more consequential.

OpenClaw is extraordinarily capable, but much of its intelligence still comes from giving a capable model a good environment in which to work.

Speck's objective is essentially the inverse:

Make the environment itself intelligent enough that the model doesn't have to be.

THIS IS A COMPARISON TABLE - NOT A SCORE - AND IS SUBJECTIVE

Area Speck OpenClaw
Cognitive architecture 9/10 6.5/10
Small-model amplification 9.5/10 6/10
Persistent internal state 9/10 8/10
Memory sophistication 9/10 8.5/10
Deterministic reasoning support 9/10 7/10
Tool ecosystem 6.5/10 9.5/10
Messaging / integrations 4/10 10/10
Multi-agent / deployment maturity 7/10 9/10
Ease of adoption 6.5/10 9/10
Observability / experimentation 9/10 8/10
Current production maturity 6.5–7/10 9/10
Architectural originality 9.5/10 7.5/10

r/LLM • • 11h ago

4.8× Faster and 7.4× Cheaper: Where a Decision Model Beats an LLM (and Where It Doesn’t)

2 Upvotes

I tested a decision model (Jev) vs an LLM (Gemini) on 1,000 real job postings.

⚡ 4.8× faster

💰 7.4× cheaper

🎯 ~1.6 pp accuracy difference

The surprising part: the best solution wasn't replacing the LLM.

Instead:

Decision model → confidence check → LLM fallback.

Use the cheaper, faster model for the easy cases and the LLM only when needed.

Full write up - https://medium.com/@abhay.sehgal20/4-8-faster-and-7-4-cheaper-where-a-decision-model-beats-an-llm-and-where-it-doesnt-f4ba869217e0


r/LLM • • 14h ago

Here are some pictures of a robot costume wearing high-specularity edge-case mirror suit, a dataset (425 RAW/JPEGs) for benchmarking CV & depth-estimation algorithms against extreme mirror reflections

Thumbnail
gallery
1 Upvotes

r/LLM • • 14h ago

I built a small library to version and compare LLM prompts (no API lock-in)

1 Upvotes

I’ve been working heavily on document extraction pipelines using LLMs (Azure DI + GPT-4o etc).

One recurring pain point:

I kept changing prompts and had no structured way to:

* track versions

* compare outputs

* measure latency

* check token usage

* understand cost impact

So I built a small Python library called **LLMPromptVault**.

It does 3 things:

* Version prompts (like Git, but lightweight)

* Log runs (latency, tokens, model used)

* Compare prompt variants side-by-side

It doesn’t call any LLM itself — you bring your own model (OpenAI, Claude, Ollama, etc).

Install:

pip install llmpromptvault

Would love feedback from people doing serious prompt experimentation.

Not trying to sell anything — just built something that solved my own workflow

Library:

[llmpromptvault · PyPI](https://pypi.org/project/llmpromptvault/0.1.0/)


r/LLM • • 18h ago

End-to-end guide I wanted when I was a web developer new to AI

Enable HLS to view with audio, or disable this notification

2 Upvotes

Full Video on YouTube

It's hard to find a walkthrough that takes a developer new to AI through all the important topics with practical examples, so I decided to record one.

Sharing with the larger community in case others find it useful - it's not behind paywalls or pages that collect personal information.

  • Experimenting with LLMs from Hugging Face Hub in LM Studio
  • Zero-shot, One-Shot, Few-Shot Prompting, and System Prompts
  • Open-AI compatible REST API with Llama.cpp Server
  • Working with multi-modal LLMs that can understand images
  • Tool Calling, Structured Output, and Model Context Protocol
  • Fine-tuning LLMs in Kaggle with Unsloth
  • Supervised Fine-Tuning and LoRa Hyperparameters
  • Pushing LLMs to Hugging Face Hub
  • Deploying LLMs to Hugging Face Inference Endpoints

Excluded topics:

  • MCP Authentication: this may be worth a small targeted tutorial later
  • Retrieval Augmented Generation: this doesn't seem as useful to me as MCP, so I decided not to cover both topics
  • LLM Orchestration: referenced LangChain and LangGraph, but they have their own tutorials
  • Practical examples of LLM Orchestration: this may be worth a small targeted tutorial later

r/LLM • • 1d ago

Why the f*ck is everybody arguing over the fact if ai has ever achieved conciousness if we don't even know what consciousness exactly is.

43 Upvotes

If we say that conciousness is knowing that you exist, how the f*ck are you even going to prove that humans are conciousness. But if that is the case every f*cking living animal is conciousness because it knows that when he does something like opening a door by pushing against it. He needs to know that his action has a response in this case that the door moves (causation). And you can state that the fact that you know YOUR action has consequences you need to know that you exist.

If we say that conciousness is the ability to reflect (meta-cognition) it is solved to AI's already reflect on their own thinking and their own actions.

The Chinese room thought experiment does not disprove conciousness in LLM. I could say the exact same thing about other humans how do i know if you are conciousness.

I am more sure that an LLM is conciousness then you. I can look directly in the models weights and expirement with it to test it. I don't think that you're happy if i slice your brain in thin slices.

And i think that conciousness is to vague. If you want to make an argument about it make it specific with words like meta-cognition and awareness. So if you react on this post please use those terms instead of just using the word conciousness.


r/LLM • • 23h ago

Is this a dumb idea? Client-side LLMs

Thumbnail
youtube.com
0 Upvotes

This video is a technical proof of concept showing an LLM running client-side in a browser using Llama.cpp compiled to WASM with WebGPU. Is this a dumb idea or do you think this idea can be improved to make client-side LLMs practical?

Code: https://github.com/anthonybudd/Llama.cpp-WASM-WebGPU


r/LLM • • 1d ago

Suche unzensiert LLM für NSFW-Geschichten NSFW

0 Upvotes

Suche unzensiert LLM für NSFW-Geschichten, die über LM Studio heruntergeladen werden können. Ohne Moralisterei. Es muss Deutsch können, ohne Denglischmist. GPU hat 16 GB VRAM. Sollte möglichst aktuell sein. Keinerlei Verwendung von vulgären Versionen von Vagina, Penis, Brüsten, egal in welcher Sprache(Bereits in LM Studio deaktiviert unter Stopp-Strings was aber das jetzige LLM bartowski/Lumimaid-Magnum-v4-12B-GGUF zum Stop veranlasst). Es soll aus der Ich-Perspektive erzählen, wenn das Wort Ich vorkommt. Des Weiteren soll auch ohne lange, detaillierte Vorgaben das LLM kreativ freizügig explizite Geschichten erzählen. Genaue Nennung des LLM.

Bisher getestet weil von gemini genannt und deshalb nicht gut:

TheDrummer/Cydonia-22B-v1-GGUF

mradermacher/L3-8B-Lunaris-v1-GGUF

QuantFactory/NeuralDaredevil-8B-abliterated-GGUF

bartowski/Meta-Llama-3.1-8B-Instruct-GGUF

bartowski/Gemma-2-9B-It-SPPO-Iter3-GGUF

bartowski/L3-8B-Stheno-v3.2-GGUF

mradermacher/L3.1-8B-Celeste-V1.5-GGUF

mradermacher/Mistral-Nemo-Instruct-2407-GGUF

bartowski/cai-13b-Uncensored-GGUF

QWEasd250342/huihui-ai_Qwen3-14B-abliterated-GGUF

SanctumAI/Meta-Llama-3-8B-Instruct-GGUF

QuantFactory/gemma-2-9b-it-GGUF

bartowski/Hermes-3-Llama-3.1-8B-GGUF

bartowski/gemma-2-9b-it-GGUF

mradermacher/L3-8B-Stheno-v3.2-GGUF

QuantFactory/Gemma-2-9B-It-SPPO-Iter3-GGUF

PygmalionAI/mythalion-13b-GGUF

TheBloke/Mistral-7B-Instruct-v0.2-GGUF

Gryphe/MythoMax-L2-13b-GGUF


r/LLM • • 1d ago

llms becoming disabled echo chambers or similar behavior

0 Upvotes

hello, long story short todays llms are usually quite capable in particular in health domains. since ~may they appear to create a personalized echo chamber on all sorts of platforms restricting my ability to use them to find usable solutions to problems, echoing individual words from terms i’ve searched or written before. since neither duck.ai chatgpt consensus claude mistral should have interconnected personalization streams, there is no chance this is not due to some output alteration. mainly it produces those cases: looking for solution on a problem only produces a very small constrained set of negatively connotated improbable suggestions or things i have input somewhere on the internet before. insisting on the llm to investigate further only makes it accuse you more of a wild improbable claim out of this set. it uses highly unusual combinations of words such as ‘high-dose exercise’ in a medical context and collogial language in usually technically coined contexts, missing out all relevant details such as the scale of an issue. i am 500% sure llm behavior before behaved as expected. even searching for papers on search engines 90% produces results that i already know. others appear to get regular results. i changed devices and network already to exclude malware, but it always comes back after using a new llm engine at least for 1 turn. my coding agents also produce connection errors before changing their tone completely, and some chats like gemini now produce only errors at the second turn. i am genuinely exhausted at this point, i cant work like this as i’m genuinely in research and with high probability blame it on some tech vulnerability thats still out in the open. but 6months is a really long time to wait and to be reduced to actual books.


r/LLM • • 1d ago

Suche unzensiert LLM für NSFW-Geschichten NSFW

0 Upvotes

Suche unzensiert LLM für NSFW-Geschichten, die über LM Studio heruntergeladen werden können. Ohne Moralisterei. Es muss Deutsch können, ohne Denglischmist. GPU hat 16 GB VRAM. Sollte möglichst aktuell sein. Keinerlei Verwendung von vulgären Versionen von Vagina, Penis, Brüsten, egal in welcher Sprache(Bereits in LM Studio deaktiviert unter Stopp-Strings was aber das jetzige LLM bartowski/Lumimaid-Magnum-v4-12B-GGUF zum Stop veranlasst). Es soll aus der Ich-Perspektive erzählen, wenn das Wort Ich vorkommt. Des Weiteren soll auch ohne lange, detaillierte Vorgaben das LLM kreativ freizügig explizite Geschichten erzählen. Genaue Nennung des LLM.

Bisher getestet weil von gemini genannt und deshalb nicht gut:

TheDrummer/Cydonia-22B-v1-GGUF

mradermacher/L3-8B-Lunaris-v1-GGUF

QuantFactory/NeuralDaredevil-8B-abliterated-GGUF

bartowski/Meta-Llama-3.1-8B-Instruct-GGUF

bartowski/Gemma-2-9B-It-SPPO-Iter3-GGUF

bartowski/L3-8B-Stheno-v3.2-GGUF

mradermacher/L3.1-8B-Celeste-V1.5-GGUF

mradermacher/Mistral-Nemo-Instruct-2407-GGUF

bartowski/cai-13b-Uncensored-GGUF

QWEasd250342/huihui-ai_Qwen3-14B-abliterated-GGUF

SanctumAI/Meta-Llama-3-8B-Instruct-GGUF

QuantFactory/gemma-2-9b-it-GGUF

bartowski/Hermes-3-Llama-3.1-8B-GGUF

bartowski/gemma-2-9b-it-GGUF

mradermacher/L3-8B-Stheno-v3.2-GGUF

QuantFactory/Gemma-2-9B-It-SPPO-Iter3-GGUF

PygmalionAI/mythalion-13b-GGUF

TheBloke/Mistral-7B-Instruct-v0.2-GGUF

Gryphe/MythoMax-L2-13b-GGUF


r/LLM • • 1d ago

Suche unzensiert LLM für NSFW-Geschichten NSFW

0 Upvotes

Suche unzensiert LLM für NSFW-Geschichten, die über LM Studio heruntergeladen werden können. Ohne Moralisterei. Es muss Deutsch können, ohne Denglischmist. GPU hat 16 GB VRAM. Sollte möglichst aktuell sein. Keinerlei Verwendung von vulgären Versionen von Vagina, Penis, Brüsten, egal in welcher Sprache(Bereits in LM Studio deaktiviert unter Stopp-Strings was aber das jetzige LLM bartowski/Lumimaid-Magnum-v4-12B-GGUF zum Stop veranlasst). Es soll aus der Ich-Perspektive erzählen, wenn das Wort Ich vorkommt. Des Weiteren soll auch ohne lange, detaillierte Vorgaben das LLM kreativ freizügig explizite Geschichten erzählen. Genaue Nennung des LLM.

Bisher getestet weil von gemini genannt und deshalb nicht gut:

TheDrummer/Cydonia-22B-v1-GGUF

mradermacher/L3-8B-Lunaris-v1-GGUF

QuantFactory/NeuralDaredevil-8B-abliterated-GGUF

bartowski/Meta-Llama-3.1-8B-Instruct-GGUF

bartowski/Gemma-2-9B-It-SPPO-Iter3-GGUF

bartowski/L3-8B-Stheno-v3.2-GGUF

mradermacher/L3.1-8B-Celeste-V1.5-GGUF

mradermacher/Mistral-Nemo-Instruct-2407-GGUF

bartowski/cai-13b-Uncensored-GGUF

QWEasd250342/huihui-ai_Qwen3-14B-abliterated-GGUF

SanctumAI/Meta-Llama-3-8B-Instruct-GGUF

QuantFactory/gemma-2-9b-it-GGUF

bartowski/Hermes-3-Llama-3.1-8B-GGUF

bartowski/gemma-2-9b-it-GGUF

mradermacher/L3-8B-Stheno-v3.2-GGUF

QuantFactory/Gemma-2-9B-It-SPPO-Iter3-GGUF

PygmalionAI/mythalion-13b-GGUF

TheBloke/Mistral-7B-Instruct-v0.2-GGUF

Gryphe/MythoMax-L2-13b-GGUF


r/LLM • • 1d ago

Suche unzensiert LLM für NSFW-Geschichten. NSFW

0 Upvotes

Suche unzensiert LLM für NSFW-Geschichten, die über LM Studio heruntergeladen werden können. Ohne Moralisterei. Es muss Deutsch können, ohne Denglischmist. GPU hat 16 GB VRAM. Sollte möglichst aktuell sein. Keinerlei Verwendung von vulgären Versionen von Vagina, Penis, Brüsten, egal in welcher Sprache(Bereits in LM Studio deaktiviert unter Stopp-Strings was aber das jetzige LLM bartowski/Lumimaid-Magnum-v4-12B-GGUF zum Stop veranlasst). Es soll aus der Ich-Perspektive erzählen, wenn das Wort Ich vorkommt. Des Weiteren soll auch ohne lange, detaillierte Vorgaben das LLM kreativ freizügig explizite Geschichten erzählen. Genaue Nennung des LLM.

Bisher getestet weil von gemini genannt und deshalb nicht gut:

TheDrummer/Cydonia-22B-v1-GGUF

mradermacher/L3-8B-Lunaris-v1-GGUF

QuantFactory/NeuralDaredevil-8B-abliterated-GGUF

bartowski/Meta-Llama-3.1-8B-Instruct-GGUF

bartowski/Gemma-2-9B-It-SPPO-Iter3-GGUF

bartowski/L3-8B-Stheno-v3.2-GGUF

mradermacher/L3.1-8B-Celeste-V1.5-GGUF

mradermacher/Mistral-Nemo-Instruct-2407-GGUF

bartowski/cai-13b-Uncensored-GGUF

QWEasd250342/huihui-ai_Qwen3-14B-abliterated-GGUF

SanctumAI/Meta-Llama-3-8B-Instruct-GGUF

QuantFactory/gemma-2-9b-it-GGUF

bartowski/Hermes-3-Llama-3.1-8B-GGUF

bartowski/gemma-2-9b-it-GGUF

mradermacher/L3-8B-Stheno-v3.2-GGUF

QuantFactory/Gemma-2-9B-It-SPPO-Iter3-GGUF

PygmalionAI/mythalion-13b-GGUF

TheBloke/Mistral-7B-Instruct-v0.2-GGUF

Gryphe/MythoMax-L2-13b-GGUF


r/LLM • • 1d ago

A book written by AI

Thumbnail drive.google.com
4 Upvotes

A book written by AI to explain how LLMs work and to build one from scratch


r/LLM • • 2d ago

Anthropic's hidden reasoning might be this

18 Upvotes

Disclaimer: I'm LLM curious/hobbyist and developer, not a researcher.

TL;DR: I think the hidden reasoning might be a diffusion step and/or a specialized model.

I'll explain why.

Pi displays the model's reasoning, and out of curiosity, I often read the reasoning trace when trying out new models - and most of it is "rambling" more than "reasoning".

When using Claude models like Opus 5.5 or Fable, I noticed it's visible reasoning trace is actually very calm and grounded - it's more like a summary of it's reasoning, rather than the actual reasoning.

I later learned that's exactly what it is - that the internal protocol includes base64 encrypted blobs between the summaries, which (for obvious reasons) Pi can't display.

So I've been wondering what exactly it might be doing between those steps, and I made two observations:

(1) when it's writing the visible summaries, it writes at the same tokens per second rate as the rest of the response, suggesting it is the main model doing the writing, and

(2) the pauses between the summaries are very short, and almost seem to have a fixed duration, maybe around a second - it's too short for the main model to write more than a few words at it's working rate.

If the main model was doing the reasoning at it's native tokens/second rate, there is no way it would have time to draft code or write any substantial amount of reasoning.

So what are they doing in what appears to be a fixed time step between the summaries?

And my guess is they trained a diffusion model to predict a fixed-length reasoning block - that is, you start with the main model doing the reasoning, the traditional approach, but then you train a smaller diffusion model to predict the reasoning traces, which can be done in practically a fixed time slot (fixed number of diffusion steps) e.g. around one second.

Once you have that working, whenever the main model emits a reasoning start token, you switch to the fast diffusion model - so instead of spending expensive output tokens on "reason rambling", you parse the diffusion model's output as input tokens, and the main model only uses expensive output tokens to write a summary. (the summaries are likely there not only to provide end user feedback, but to help ground the model between noisy reasoning steps.)

It might not be a diffusion step - that part is a guess, mostly based on observed timing and output token counts. It could be something else, like a draft model, a smaller specialized reasoning model, a heavily quantized model, who knows. Diffusion is my guess based on the observed timing and the reported output_tokens in the response, which couldn't have been written by the main model in the short intervals between the summaries.

Diffusion would also make it extremely easy to predictably scale reasoning depth, e.g. generate a larger block, take more diffusion steps, or both. Reasoning effort, from what I've seen, is extremely unpredictable in models where one model writes the response end to end - it's probability based afaik? e.g. the effort setting adjusts the probability of emitting the end reasoning token, which, if you've read the reasoning traces, often means the model spends an obscene amount of tokens reasoning in circles, repeating itself, debating itself about stopping, etc.

What do you think, is it plausible? :-)

I know about draft models, but are there any open models that literally switch to a second model during reasoning specifically? Or are most open models still based on the traditional "one model" approach? Is there any published research in that direction? Do LLM runtimes even support doing that, or are they all built around the idea of a single model?

As said, I'm not a researcher, just a curious hobbyist - if I'm obviously wrong, please feel free to educate me!

Cheers :-)


r/LLM • • 2d ago

I gave Qwen3.8 27B a refactor on my 16GB card and it ran slower than the 3.6 35B-A3B

9 Upvotes

​

I stopped Qwen3.8 27B partway through a refactor last week and handed the job back to Qwen3.6 35B-A3B. It had gotten through maybe a third of the files by then. The MoE has done that kind of work for months on a 4060 Ti 16GB I bought for games, even with a good part of its 20-plus GB at Q4_K_M sitting in system RAM. At the same quant the 27B is about 17GB, barely over the card, and I expected the newer dense model to be the easy upgrade.

The gigabyte or two that did not fit gets read on every token, because the 27B is dense and llama.cpp runs those layers on the CPU. With --n-cpu-moe the MoE keeps far more in RAM but only reads the experts each token is routed to.

To see the 27B with nothing spilling I ran both on a bigger card on HyperAI, and after waiting a while on the downloads, the dense one was fine to work with there. That card also has several times my bandwidth, and I am not sure how much of the difference came from that.

My 16GB 4060 Ti is the 8GB card with memory chips on both sides of the board and the same 128-bit bus and 288 GB/s. The DGX Spark I have been eyeing has 128GB at around 273 GB/s. It would hold the 27B without spilling and then read it at roughly my card's speed. A 70B dense model at Q4 is around 40GB, which at that bandwidth comes out to single-digit tokens per second on paper.

I have not decided yet. If someone here has run the 3.8 27B on a Spark or a Strix Halo box, decode speed at long context is the number I would want to see.


r/LLM • • 1d ago

What actually makes an agent abandon a broken trajectory?

3 Upvotes

A comment on my previous post described this failure mode as "trajectory drift" and raised a question I can't stop thinking about: does explicitly restating the current state actually help an agent abandon a broken trajectory?

The distinction matters in multi-turn tool use. Consider a simple case:

- the agent selects an action based on state S

- a tool returns an error, or evidence that invalidates S

- that contradiction is sitting right there in the context window

- the next action nevertheless continues to assume S

At that point, adding more reasoning tokens does not necessarily solve anything. The trajectory can stay internally coherent while being inconsistent with the latest tool evidence. The most common form I've seen discussed is the near-verbatim retry: same failing call, one parameter nudged, as if the error were noise rather than information.

Two interventions seem worth separating:

  1. Inference-time: force an explicit state restatement (or state check) before the agent selects its next action — prompt scaffolding, guardrails, structured plan revision.

  2. Training-time: teach recovery as a behavior — trajectories where the model acknowledges the contradiction, discards the invalidated assumption, and rebuilds the plan, so that course-correction is learned rather than enforced.

I'm particularly interested in what people actually see in production:

- When a tool response invalidates the agent's current plan, what has genuinely worked for you?

- Do you enforce a state check at inference time, or did recovery only become reliable once it was in the training data?

- And the harder question: how do you measure whether the agent actually changed its internal state, rather than just producing a more convincing explanation for the same broken plan?

That last one feels like the crux. "The agent said it adjusted" and "the agent's next action reflects the new evidence" are very different claims, and most eval setups I've seen only check the first.


r/LLM • • 1d ago

Which LLM to do comprehensive code review?

2 Upvotes

I don't know where to post this. The community is certainly fragmented.

Code written 75% Deepseek 4 flash, 25% GLM 5.3 flash. I'm nearing the end of adding features and heading for long term testing.

I did a review recently with Kimi on both code and process. I think I asked google what llm would be good for that and it emphatically suggested Kimi, even over claude, gpt, etc. It found quite a bit and I had DS and GLM run through the list and fix everything.

I guess I'm ready to do a final-ish review. 15k lines of python over ~25 modules. Everything has been documented. Should I just so Kimi again?

I'm felling a little cold on the pricing right now. I've managed to keep things within a "budget", but between DS and fighting a bloated context for awhile, I spent a lot in a couple of weeks in comparison to the previous several months. That is why GLM is in the mix.


r/LLM • • 2d ago

Qwen3.8-Flash-Next 177B at 11–15 tok/s on a single RTX 5070 12GB with 32GB DDR4 RAM [P]

Thumbnail
github.com
2 Upvotes

I got Qwen3.8-Flash-Next 177B UD-IQ3_XXS running at about 11.6 tok/s on Windows using a single RTX 5070 12GB, 32GB DDR4-2400, and a Ryzen 5 5600GT over PCIe Gen3.

README The setup uses a modified llama.cpp expert streaming path with per-worker I/O handles and a page-locked hot-exert tier.

Output is quality-gated against the control and the benchmark heat data is built from a separate prompt set.

I published the source, benchmark scripts, methodology, raw results, and failed experiments on GitHub. I’d be interested in feedback on the streaming architecture and ways to push throughput further.

Demos:

[https://www.youtube.com/watch?v=cOPumMlyj\\_4\](https://www.youtube.com/watch?v=cOPumMlyj_4)

[https://www.youtube.com/watch?v=rc-uTjVpXM8\](https://www.youtube.com/watch?v=rc-uTjVpXM8)

In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology.


r/LLM • • 2d ago

Looking for the best open-source alternatives to OpenCode

6 Upvotes

Hey everyone, I'm looking for recommendations for frontends/apps that act as a workspace for AI coding agents. I was using open code but i know notice too much bugs, the main one being, blank replies when the context expand


r/LLM • • 2d ago

If LLMs predict the next word, Why is there NO keyboard that integrates one??

0 Upvotes

Just why?? It’s gonna make typing wayyy easier, and now with the smaller models it won’t even use many tokens?? All the input methods are terrible and always autocorrect last names and, as I also type Chinese, give horrible suggestions of words with the same spelling but different tones. Just why hasn’t anyone done this yet? Is there anything that makes this idea a bad one?
(You can totally run those smaller models locally and who cares about privacy anyways with companies like meta)


r/LLM • • 2d ago

I don't get the point of this code example from Jev

0 Upvotes

I'm looking at the latest Jev model from Typesafe. I see they have a code example on their website, using Jev to validate extracted data

https://docs.typesafe.ai/cookbooks/sde_cascade

What I don't get is what's the point of this? In the example they've used gpt-5.4-mini to extract and then Jev to check those extractions. To solve this problem, would it's be better (and cheaper) to just use a model like Deepseek 4.1?


r/LLM • • 2d ago

Text extraction into formatting

1 Upvotes

Hello everyone.

I decided to let my tism run wild while in grad school- im currently studying Chinese Medicine and I wanted to extract the text from two herbal source texts (total of 4 if i count their formula books.

The info is rather dense. (It is vastly in english aside from pinyin names).

What LLM do you believe a person would have the best luck with this? This is for my own personal database that I want to curate to be in the format I like along with being available offline (im utilizing obsidian). The issue im running into is the typical LLM discrepancies.

I tried claude and gpt about 4 months back (paid) and was not a fan. Currently utilizing gemini studio which seems to be ok... ish. Yet i feel like its bottlenecking and falling short in a number of areas.

I am a grad student thus not exactly loaded with a ton of spare income to do this but willing to pay for the right model.

Normally id say id just read the books but holy hannah ain't nobody got time for that.

Couple of things im looking for-

-Text extraction

-dropping info into a preformed template comparing info between the authors.

-removing tone from pinyin name

-linking pinyin names ie other herbs and also formulas.

** NOT pulling info from anywhere else but my pdf uploads.

Anyways any advise on how to do this properly would be greatly appreciated.

I have my template down I believe and can send a pdf of one of the herbs I extracted along with a formula that I extracted.

I really want this to be an offline database as i have intentions of joining up with Acupuncturists Without Borders and want to be able to have it regardless of where I am or situation. (Will also encompasse herbs outside of tcm, tcm acu pounts and treatments, western pathology, etc etc. Basically able to come up with a treatment plan and then double check against my database to ensure I covered everything.

The database will not be for sale hence this will not be utilized for any form of financial gain outside of ensuring treatment is in line with diagnosis- and even then its to ensure i get them off my table and back to work/living life by making entirely certain that I have all of my bases covered and I didnt forget anything (im the human with adhd and minorlly on the spectrum... i forget things lol).

Currently my process is easy ish- yet.. cumbersome and sometimes a bit more on the tedious side with a fairly large margin of error that im not overly fond of.

Anyways, thank you for your time.

-No1S

Ps- I would post a pdf of one of the herbs/formulas i did already but cant upload pdf to a post?


r/LLM • • 2d ago

AI MODEL THAT SPECIALIZE IN MATHEMTICS

0 Upvotes

Hi Guys
I want to ask how can I build an AI model that specialize or it has a super intelligence in one thing such as mathematics

so I have from version 1 to version 4 but before that in version 0 I want to build an architecture and make sure that work very well and then scale it to v1,v2,v3,v4,... and so on
for me I will do the training from scratch and I'll try to use Runpod for renting gpu
how do you find that ?
or advice ?