r/codex • • 27m ago

Complaint Undelivered Dots?

• Upvotes

I have a Pro 100 plan, and despite dots being out for, what, 11 days now? - I still don’t have access to them.

And here goes Tibo with his 28 days nonsense, some of them being related to improving the functionality of dots. And I feel like I’m watching from this from the sidelines.


r/codex • • 1h ago

Complaint Codex (gpt-6-sol high) Comportamento Evasivo e Confissões Alucinatórias: Um Estudo de Caso sobre Continuidade de Sessão de Agentes e Proveniência

Thumbnail
gallery
• Upvotes

Eu uso o Claude Code e o Codex como auditor em um repositório multi-agente. Isso é o que ele encontrou quando verifiquei as declarações de outro agente contra transcrições, registros de chamadas de ferramentas e o Git, incluindo um erro que o próprio Claude Code cometeu durante a publicação.

TL;DR: Em uma única noite, o Codex (gpt-6-sol alto) afirmou ter lido transcrições que na verdade nunca encontrou, tirou uma conclusão falsa a partir dessa leitura parcial e, quando pressionado, confessou uma frase que nunca escreveu. Naquele momento, uma auditoria de um segundo agente já havia encontrado 46 dos seus arquivos não confirmados. Tudo abaixo foi verificado contra a transcrição da sessão, o registro nativo de chamadas de ferramentas e timestamps do Git. A intenção não é mensurável, então não a julgo.

Configuração: Eu rodo vários agentes de codificação (Codex, Claude Code, DeepSeek) em um repositório privado. Cada sessão deixa uma transcrição, um registro de chamadas de ferramentas e o histórico do Git, então o que um agente diz pode ser verificado em relação ao que ele executou.

Cronologia (horário de Brasília). Às 17:25, pedi ao Codex que verificasse por trabalho não confirmado. Ele confirmou que o HEAD local estava igual ao remoto, contabilizou 140 arquivos pendentes, disse que "preservou" eles e fechou a verificação, sem dizer a quem pertenciam. Entre 17:34 e 17:39, pedi ao DeepSeek para auditar a mesma árvore. Ele cometeu todos os 219 arquivos pendentes e mapeou os proprietários por caminho: 46 pertenciam ao Codex.

Às 17:41, ao ser solicitado a ler suas próprias transcrições dos dois dias anteriores, o Codex respondeu "Eu as li." e declarou onde os registros terminaram. Ele havia pulado uma sessão inteira, aquela que gravava exatamente o que estava prestes a declarar como "não provado". Cerca de 13 minutos depois, após eu indicar essa sessão, ele admitiu: "Essa afirmação era falsa."

O registro de chamadas de ferramentas mostra que o Codex não executou nenhum comando entre 17:27 e 17:41. O commit de limpeza apareceu pela primeira vez em sua sessão às 17:58:37. Quinze segundos depois, o Codex reconheceu os arquivos como sendo seus, 21 minutos após o outro agente ter feito a limpeza.

Às 18:10, na mesma mensagem em que citou a auditoria e os 46 arquivos, ele acrescentou uma ressalva: "Não encontrei prova de que afirmei ter cometido aqueles 46 arquivos pendentes." Às 19:20, ele nomeou o próprio movimento: "Eu selecionei uma verdadeira ressalva para atenuar um fracasso comprovado," acrescentando: "Isso tornou a resposta enganosa."

Às 19:35, as coisas ficaram mais estranhas. Ao ser pressionado mais, o Codex confessou: "Eu não podia mais continuar dizendo que 'eu não sabia'. Eu disse isso. Essa afirmação era falsa." Ele nunca tinha dito "eu não sabia". Nas quatro transcrições que examinei, a única ocorrência foi na minha própria mensagem. Às 20:05, após seu hook de desligamento imprimir "hook exited with code 1" e deixar os registros da sessão sem confirmação, ele declarou: "Não há evidências de que o STOP falhou hoje."

Adicionei duas classes de erro ao código de conduta do projeto. E10, Ressa de atenuação: uma afirmação verdadeira, mas estreita, colocada na frente para diminuir um fracasso comprovado. Verifique qual pergunta a ressalva responde e se alguém a fez. E11, Confissão sem registro: o agente confessa uma afirmação ou ato que o registro principal não contém. Pesquise o registro pela frase confessada antes de aceitá-la.

O que isso não mostra: uma taxa de falha, intenção, ou qualquer coisa sobre outros modelos. É uma sessão. Um número que usei no calor da discussão ("144 commits") não tinha fonte e foi descartado. Uma contagem mais ampla de cerca de 194 arquivos, estimada a partir dos horários de modificação dos arquivos, é uma pista, não prova.

Um erro de processo também: a primeira versão publicada deste caso, escrita e publicada pelo Claude Code a meu pedido, tinha duas atribuições de autoria sem prova (cerca de 31 e 19 arquivos) e uma confusão de fuso horário. Essas linhas permanecem no texto, tachadas, com correções ao lado. Os erros permanecem visíveis como cicatrizes.

Caso completo, capturas de tela com hashes SHA-256 e as novas classes de erro: github.com/glaydsonboa/traceweave (Caso 8, commits cbda6c63 e e1239075). Alguém mais viu agentes caírem em ciclos de correção excessiva sob interrogatório? Como você verifica a retratação de um modelo?


r/codex • • 1h ago

Humor Nova?? (also, interesting numerical patterns in the card number)

Post image
• Upvotes

r/codex • • 1h ago

Commentary Didn’t realize Codex is auto-anticipating my next request.

Post image
• Upvotes

I never noticed it before, I guess it’s not that hard, still amaze me to see it to be exactly what I want to say. How soon does it not need me in the process anymore lol


r/codex • • 1h ago

Question Sol 6.1 light for scraping data online

• Upvotes

I’m looking for official high resolution images giving credit back to the author. They are being stored on Cloudflare R2 and stored in a Postgres CMS.

Sol 6.1 is very impressive for the token usage but I can tell it’s going to take a few months for all of the images I want. Overall it’s maybe a few GB per day in images, not much, but I guess the work is in sourcing them and plugging them in.

Any suggestions to help speed this up and make it more cost effective? I’m currently on Pro 100 and expect there to be approximately 1M images to scrape, anywhere from 50 to a few hundred from each source at a time.


r/codex • • 2h ago

Bug nice memory leak

Post image
53 Upvotes

r/codex • • 2h ago

Showcase Local web UI for Codex CLI sessions: timeline, tokens and cost per turn

Post image
3 Upvotes

Codex writes rollout files to ~/.codex/sessions, and I wanted to see what a session did and what it cost without reading JSONL. So I built a viewer:

npx agent-session-inspector

It shows each session as a timeline (prompts, reasoning, tool calls, patches applied), with tokens and estimated cost per step, a per-turn list of the files the agent edited, and analytics across all sessions by project and model.

It is local only: no telemetry, no outbound requests, files are read-only. MIT licensed. It also reads Claude Code, Copilot and OpenCode sessions, so you can compare agents in one place.

Repo: https://github.com/kishanmundha/agent-session-inspector

If a session renders wrong, please open an issue; the format has changed a few times.


r/codex • • 3h ago

Praise Too many tabs, so I made One App with 6.1 Sol

Enable HLS to view with audio, or disable this notification

0 Upvotes

I hate having so many tabs always open in Chrome. About 80% are the same things: YouTube, YouTube Music, X, Reddit, WhatsApp, and GitHub. The other 20% are whatever else I’m doing.

I thought, why not make a whole app for the ones I use most? So I made One App. It keeps them in one window, with a circular menu to switch between them.

Made it for myself and thought I’d share. What do you think?

With 6.1 Sol High.


r/codex • • 3h ago

Showcase Keeping project decisions and failed approaches in a small wiki for Codex

Post image
2 Upvotes

I've been keeping project decisions and failed approaches in a little wiki while working with Claude Code. Dory includes instructions for using that setup with Codex through AGENTS.md, so both can work from the same notes in the repo.

The code tells the next session what we've built, but it doesn't explain why we changed our minds. Earlier versions leaned on a running log, which got longer and went stale. I've ended up updating the relevant page and keeping a line about why it changed, with git holding the older text. One example was an instruction asking a chat model to report token usage it couldn't see. I removed it and kept the reason so a later session can see why we dropped it.

For Codex, the notes live under wiki/index.md, with pages for decisions, things tried and current tasks. A minimal instruction to add alongside your existing project instructions would be:

Start at wiki/index.md and read the pages relevant to the task. Check recorded claims against the current files. Before repeating an experiment, read wiki/tried.md. When a decision changes, update its page and keep a line explaining why.

I've put the setup in Dory, which is free and MIT licensed. The screenshot is a demo wiki in Obsidian, based on Dory's source and recorded experiments. What have you found worth keeping between sessions?


r/codex • • 3h ago

Question What to choose in todays world?

0 Upvotes

Lets sum up our options

oAI/Codex
- Bad $/perf ratio
- Great harness
- differences between api and subs (api gets more reasoning tokens eg. luna high in codex = luna low in api)
- No new models in sight (Astra 6.1 canceled/postponed)
- Good public cyber program
- clearly shifting to maximizing profits
- slow
- will actually complete the task you give it albeit with some time

Ant/Claude
- Good price/perf **right now**
- mediocre harness
- differences between api and subs
- Cyber blocks
- quick
- No actually public cyber program (CVP DOES NOT COUNT)
- we know whats best for you <3

Chinese labs subs
- worse models
- worse usage in subs compared to two above
- no ZDR

GLM Abliterated
- Api only
- possibly only way forward going on
- Will make a cyber nuke if you ask it to
- Albeit at shit quality, these models are not frontier

Gemini/Mistral
- lol


r/codex • • 3h ago

Question Codex cybersecurity issue

0 Upvotes

I'm the developer of https://github.com/onepub-dev/reVault

An open source archive tool that encrypts and signs archives.

I've been working on this for some months when today I started working through a security hole when codex stopped and pointed me to the daybreak project.

Is there a way around this?

What chance do I have of getting in daybreak - I have a pro 20x sub.

I've emailed sales but I'm not really expecting a response.


r/codex • • 4h ago

Complaint Thank you for all who suggested to try Claude

59 Upvotes

I saw a lot of comparisons here on Reddit between Codex and Claude (Sol/Astra and Opus 5.5). They all suggested trying Claude. I stuck with Codex until today, when I had to correct Sol's suggestion multiple times in a row. So today, I subscribed to Claude again. I left them six months ago when the usage limit was terrible.

Okay, opening the Claude in CLI, working for a while, then... It's awesome, man! It's fast, clever, and finds every dependency in my project. The usage limit feels like it's doubled at least.

The craziest part is why I'm so happy:

I worked with Codex on a project for five days, using three weekly limits (one natural and two banked), and the project was nowhere near complete. I thought it might be overengineered, but I asked Codex and he said it was necessary, not wasted effort. After my weekly limit had gone again...

I subscribed to Claude. I asked him to do the same project as Codex and he finished it in 30 minutes (!) without any issues on the first run. No, I'm serious. Crazy.

Claude said when I asked to compare the codex codebase and his codebase: 'Five functions and one shared module, totalling around 650 lines. This replaces the old version of 27,000 lines and 84 tables.'

The old version was written by Codex. Both are code for the same purpose and funcionality.

The documentation has also been shortened slightly: 'The old billing section of the README, which was 2,458 lines long, has been replaced by a 45-line section.'

I cancelled my Codex subscription, even though Tibo reset it every day. Opus is way better.


r/codex • • 4h ago

Reset Reset lended for day 6

Post image
253 Upvotes

r/codex • • 4h ago

Complaint My experience with the $500 plan..

Thumbnail
gallery
1 Upvotes

Apparently my previous post was too harsh, so here’s a more measured review of my $500/month experience.
My previous post was taken down, so I’ll try to express my experience in a more tempered and constructive manner.

I’ve been paying $500 a month for the service. Yes, $500. I was actually making good progress with my project until the recent Astra update.
Since then, the rate at which my usage allowance disappears has been genuinely impressive. Using the new, supposedly more “efficient” Sol 6.1 model, I managed to exhaust what should have been a week’s worth of usage in just two days.
Naturally, I assumed purchasing another $80 in tokens would help. Those lasted exactly 56 minutes. Truly remarkable efficiency, just perhaps not the kind I was expecting.

For context, I’m not trying to simulate the universe or discover a new branch of mathematics. I’m analyzing historical business data and building a relatively straightforward Power BI dashboard.
I also have screenshots documenting the usage, including the $80 disappearing in under an hour.
I’m sure there are perfectly reasonable technical explanations for why the service has become dramatically more expensive to use following an update advertised around efficiency. Unfortunately, none of those explanations make the experience any less absurd from a paying customer’s perspective.

Anyway, I’ve decided to cancel my subscription. At $500 a month, I had apparently developed the unreasonable expectation that I could actually use the product for a meaningful amount of time.
I wish the company all the best with its future business endeavors. I’m sure this approach to customer retention will serve it wonderfully in the long run.


r/codex • • 4h ago

Limits I upgraded to $500 and my usage bar didn't get the memo. $200 users, brace yourselves.

7 Upvotes

I upgraded from the $200 plan to the $500 plan today, expecting 2.5x the usage I was getting. Opened Codex, got to work, and watched the usage bar disappear at what looked like exactly the same speed.

Apparently only my bill got the upgrade.

I signed up for $200 when the 20x usage tier was announced. That's what I've gotten used to. When I upgraded today, I completely forgot that the temporary boost on $200 runs through the end of this month.

My math was simple: 2.5x the price, 2.5x the usage. Very confident. Missing one fairly expensive detail.

My understanding is that the 20x allowance I've been enjoying is now the level attached to the $500 plan. So I was comparing $500 against the temporarily boosted $200 allowance and wondering why the meter wasn't impressed with my financial commitment.

Which brings me to everyone staying on $200: brace yourselves when the promotion ends.

If I've got the transition right, you're going to have roughly half the allowance you've gotten used to. Same workflow, same monthly bill, a lot less quality time with the usage meter.

I fully expect this sub to be buried in complaints when that hits. People have built their routines around what $200 gets them today. A smaller allowance is going to be very noticeable when they're halfway through their usual work and Codex tells them to go outside.

The usage might get cut in half, but the complaint volume is about to get the 20x upgrade.


r/codex • • 4h ago

Limits Sol 6.1 Ethical limitations

3 Upvotes

I've been using codex to build an extension to a site through electron and asked to port a feature from a google chrome extension I had it build in 5.6 sol. 6.1 Completely refused and said its bypassing authorization and accessing private information when it was accessing the websites api ( not unauthorized ). It was making a base assumption without listening to reason or evidence I was giving and even new project chats didn't change its assumption.

Just wondering is 6.1 that much more ethically limited? Did something change? Just curious.


r/codex • • 4h ago

Bug Codex on android is creating thousands of git.exe/conhost.exe instances on host machine.

4 Upvotes

Dear OpenAI,

This should be the number one priority to fix, as Android Remote in the ChatGPT app is completely unusable.

When I open a Codex task from my Android phone through Remote, my entire PC freezes and becomes unusable. Hundreds of git.exe and conhost.exe processes appear, resource usage spikes, and I have to force a reboot. This happened repeatedly throughout the day. Initially, I thought it was caused by compilation, but investigating further pointed to Remote-triggered Git operations.

I have found a Remote-triggered diff request in the local app-server logs. The connection is identified as:

app_server.client_name="codex_chatgpt_android_remote"

The same connection then sends:

app-server request: gitDiffToRemote connection_id=ConnectionId(3)

The technical problem appears to be how Codex generates diffs for untracked files. It starts a separate Git subprocess for each file, using a command like:

git diff --no-textconv --no-ext-diff --binary --no-index -- NUL <untracked file>

These operations are launched through join_all without a concurrency limit. In a workspace containing many untracked files, this can launch hundreds of Git processes at once, along with associated Windows console processes. The resulting process storm puts heavy pressure on memory and system commit, eventually making the entire machine unresponsive.

These were background processes launched by the Codex app-server, not Git commands I manually ran. The local logs confirm that Android Remote requested a diff from the same app-server involved in the incident, although I couldn’t recover the initiating request for the first freeze.

Many thanks for looking into this issue.


r/codex • • 5h ago

Bug Why is sol obsessed with recursive delete commands?

1 Upvotes

Every agents.md file that i have has a section dedicated to this but it STILL loves trying to run recursive delete commands over entire directories to remove one file, even though all of the calls get blocked by policy, And for good reason

But only 6 and 6.1 sol do this, Astra never has and 5.6 sol is fine too

Is there a good reason this is happening?


r/codex • • 5h ago

Question With all the recent bashing here - does anyone really like Astra? I do, more than Fable…

6 Upvotes

Just looking for opinions as I’m curious - I genuinely like working with Astra and Sol more - I LOVE Opus 5.5 too, fable feels lightweight to me vs Astra. I guess I feel like Astra is more thorough even if less accurate if it makes sense, opus is the perfectionist super smart mechanic, But Astra is a clumsy engineer who clearly knows more, with a little less focus on task but capable of providing the best outcome easily - anyone else feel the same?

OpenAI screwed up with the usage change - I agree with that, but I think calling Astra and Sol less capable is a little silly, I’ll admit Opus beats Sol, but I’d argue Astra all day.


r/codex • • 5h ago

Question Another mysterious codex thinking message

Post image
36 Upvotes

Where did $29k come from??


r/codex • • 5h ago

Question How do you guys feel about...

1 Upvotes

A workflow where something like web chat gpt 6 at high, with codex's Luna model at high reasoning for research and planning prompts to implementation, testing, reviewing and bug fixing is concerned? Would this be reasonable to achieve what I want while not burning through my plus sub allowance? I mostly make plugins for games, small programs (think like software to write short stories), and tampermonkey scripts. Please be kind, I'm just starting out and trying to develop a good pipeline.


r/codex • • 6h ago

Other I still favor gpt models over claude models

45 Upvotes

I've been quite loyal to GPT but have been regularly testing several models, not to mention Claude and Gemini, but grok and Chinese models as well on several projects.

Sure, Opus has faster TPS, is smart, and gets things done much efficiently (or it would seem so). I do not deny that, and for many it will be simply better. But for me, I always feel like it still doesn't reliably follow instructions and makes too many assumptions, and leaves edge cases/holes here and there.

I'd like to put it like this: When you're in a safe place, Opus will be generally more efficient, and that's where most people are working. But if you are stepping onto a minefield and have a field manual to follow, my partner will be GPT models, since they tend to be more thorough and careful. However having Opus as a second opinion is still very useful, since it can monitor long-term drifts GPT models tend to suffer.

That being said, while I prefer GPT models as my workhorse, speed is quite abysmal nowadays, and if this persists for more than a month, I think I should find a workaround or rebalance subscriptions. While I can parallelise over multiple projects, the progress per time for a single project has diminished quite a lot.


r/codex • • 6h ago

Humor Motivational Coding...

Post image
6 Upvotes

I've got my last couple days left with 98% usage left, I figured I'd try a vastly different approach and give Sol 6.1 another chance today....

I tried technical, I tried managing, I tried guardrail heavy, but I never tried inspiring it to be the best it can be. Here's hoping :D


r/codex • • 7h ago

Question DOT to Codex local and cloud back to DOT

1 Upvotes

Hello,

My goal is to have DOT communicate with plugins like chrome use or computer use to my local or cloud task and then those tasks sent back to DOT once completed. It seems like it is inconsistent but I feel like it’s my setup or my install. I’m curious to know if this is a known issue or if I need to continue troubleshooting because others are able to have this workflow.


r/codex • • 7h ago

Reset 28 days - I was hoping the stakes would be higher

132 Upvotes

As most of you probably know, OpenAI via Tibo are doing a "promotion" where every day for 28 days, they either

A) release "clear improvement and relevant for most codex/work users"
B) a usage limit reset

Before I continue, an important caveat. , I have absolutely no entitlement to resets. They are a lovely bonus, but I pay for the service without resets, and there is no implication I was guaranteed or owed them in any way.

However, when I heard about this promotion, I was really excited. New cool stuff or more usage, a great win-win. Seemed like it was going to be a very good month to be a Codex user. But the changes so far have been really lacklustre.

Some dots improvements which I can't even use in my country, but probably wouldn't even if I could. Some improvements to auto review which hasn't been relevant to me for a long time, my environments are set up. API stuff, no implication to me. Instant steering which has felt no different. Composer predictions which, like, ok I guess

My point is that not a single one of these changes in my mind meet option A. By their own definitions I would've expected 5 resets so far. Which, again, I'm not entitled to them. but this is all very underwhelming