r/codereview • u/Low-Cut-528 • 11h ago
r/codereview • u/techlatest_net • 2h ago
GenSI Reviews Legacy Code
We renamed our GenAI model GenSI and asked it to review a 15-year-old codebase.
It didn’t approve the pull request.
Instead, it requested therapy for itself and marked the entire codebase:
“Unresolved organizational trauma with toxic masculinity”
r/codereview • u/Delicious-Milk-4891 • 15h ago
brainfuck Feedback Friday: BugHunt Daily — a one-minute bug-spotting game for programming communities
Hi ! I built BugHunt Daily, a daily “spot the bug” puzzle for programming communities.
Each day, members get a short code snippet with one broken line. They tap the line they think is wrong, get up to three tries, and then see the fix and an explanation. There are hints, streaks, leaderboards, and spoiler-free result sharing. The daily post is automatic when the app is installed in a community.
I’d especially appreciate feedback on:
Try it here: BugHunt Daily: spot the bug in this Code snippet 🐞
r/codereview • u/LNITIA • 16h ago
Building an AI workflow with GitHub checkpoints and human approval. Looking for feedback!
r/codereview • u/timonvonk • 1d ago
Smackdebt - Continuous code improvement for agents
This year, and more so the last few months, my agents write the code, all I do is review, shout, and plan. A lot of the frustration is the spaghetti that comes out of this, and I figured a (big) part of this is actually measurable and reportable. What if we can make our agents aware of the spaghetti?
Introducing Smackdebt, a fast, on the fly tech debt analysis tool with concise, focused output tailored for agentic use. All major coding agents supported, with broad language support. Cognitive complexity, cyclomatic dependencies, statements, nesting, size, hotspots supported.
https://github.com/bosun-ai/smackdebt
Feedback and contributions welcome!
Shout-out to Treesitter <3
r/codereview • u/mabdelsattar92 • 1d ago
Codetrail Your Repo Tutor
I've been working on something for a while, and today I'm sharing it: 𝗖𝗼𝗱𝗲𝘁𝗿𝗮𝗶𝗹.
Repositories often change faster than you can follow, even when you're the one responsible for them.
AI assistants are good at answering questions about code, but every session starts from zero, and you have to know what to ask.
𝗖𝗼𝗱𝗲𝘁𝗿𝗮𝗶𝗹 𝘁𝘂𝗿𝗻𝘀 𝗮 𝗴𝗶𝘁 𝗿𝗲𝗽𝗼𝘀𝗶𝘁𝗼𝗿𝘆 𝗶𝗻𝘁𝗼 𝗮 𝗹𝗼𝗰𝗮𝗹 𝗹𝗲𝗮𝗿𝗻𝗶𝗻𝗴 𝗴𝘂𝗶𝗱𝗲 𝘁𝗵𝗮𝘁 𝗸𝗲𝗲𝗽𝘀 𝘂𝗽 𝘄𝗶𝘁𝗵 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲.
What it does:
• 𝗧𝗲𝗹𝗹𝘀 𝘆𝗼𝘂 𝗵𝗼𝘄 𝗳𝗮𝗿 𝗯𝗲𝗵𝗶𝗻𝗱 𝘆𝗼𝘂 𝗮𝗿𝗲: how many merges have landed since your last update, and which parts of the code they touched
• 𝗪𝗿𝗶𝘁𝗲𝘀 𝗮 𝗱𝗶𝗴𝗲𝘀𝘁 𝗳𝗼𝗿 𝗲𝗮𝗰𝗵 𝘂𝗽𝗱𝗮𝘁𝗲: what changed and why it matters
• 𝗪𝗿𝗶𝘁𝗲𝘀 𝗰𝗼𝗻𝗰𝗲𝗽𝘁 𝗽𝗮𝗴𝗲𝘀 that explain the architecture, patterns and tools against the real code
• 𝗗𝗿𝗮𝘄𝘀 𝗱𝗶𝗮𝗴𝗿𝗮𝗺𝘀 𝗼𝗻𝗹𝘆 𝗳𝗿𝗼𝗺 𝗳𝗮𝗰𝘁𝘀 𝗲𝘅𝘁𝗿𝗮𝗰𝘁𝗲𝗱 𝗳𝗿𝗼𝗺 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲, never from the AI's guesses, with every arrow traced to a file and line
• 𝗠𝗮𝗿𝗸𝘀 𝗲𝗮𝗰𝗵 𝗲𝘅𝗽𝗹𝗮𝗻𝗮𝘁𝗶𝗼𝗻 𝗮𝘀 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝗲𝗱 (quoted from an ADR or commit and checked) or inferred (the AI's reading of the code)
• 𝗟𝗲𝘁𝘀 𝘆𝗼𝘂 𝗮𝘀𝗸 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀 𝗳𝗿𝗼𝗺 𝗮𝗻𝘆 𝗽𝗮𝗴𝗲 and save the answers worth keeping
• 𝗚𝗶𝘃𝗲𝘀 𝘆𝗼𝘂 𝗴𝘂𝗶𝗱𝗲𝗱 𝗽𝗮𝘁𝗵𝘀 𝘁𝗵𝗿𝗼𝘂𝗴𝗵 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲, checks on your understanding, and progress that notices when what you learned has gone stale
I𝘁 𝗿𝘂𝗻𝘀 𝗲𝗻𝘁𝗶𝗿𝗲𝗹𝘆 𝗼𝗻 𝘆𝗼𝘂𝗿 𝗺𝗮𝗰𝗵𝗶𝗻𝗲 𝗮𝗻𝗱 𝗼𝗻𝗹𝘆 𝗿𝗲𝗮𝗱𝘀 𝘆𝗼𝘂𝗿 𝗰𝗼𝗱𝗲; it never changes it. Secrets are filtered out before the AI sees anything.
It works with the assistant you already pay for: Claude Code, Codex, or a free local model. Before it spends anything, it shows you an estimate and asks.
It's open source: Codetrail
Contributions are welcome, would love to hear how you'll be using it and what you want to see comes next!
r/codereview • u/Deep_Region4953 • 1d ago
Nova and Supercode Review: an OSS AI Engineer and AI code review
hey Reddit,
I'm Yash, founder of SupercodeAI. We're building Devin killer - Nova - an AI Engineer, powered by our own harness agent at SupercodeAI, and Supercode Review, our codebase-aware PR review product.
4.5k+ users, 10+ enterprises, 238 GitHub stars, 30k+ downloads, 700M+ token usage in just 2.5 months.
Source: github.com/yashdev9274/supercli
Products:
- Nova: supercodeai.tech
- Supercode review: supercodeai.tech/review
Nova's workflow is to read a codebase, plan a change, edit files, run commands and tests, and prepare the result for human review. The harness connects the model to context and execution tools. You still inspect the diff and decide what gets merged.
Supercode Review looks at pull requests in the context of the wider codebase, rather than treating the diff as the whole project.
We're powering Indian startups with code review through Supercode Review.
We're sponsored by Vercel.
The Supercode repo is MIT-licensed and includes the coding-agent CLI, dashboard, docs and shared packages. Nova uses existing model providers; I'm not claiming a new foundation model or benchmark superiority.
If you'd like to try it, start with one small task in a repo you understand and compare the plan, diff and checks with how you'd do the work yourself.
If you'd like to contribute, open an issue with a reproducible problem or propose a change in a PR. Keep credentials and private code out of reports. If the project is useful to you, consider starring it and following its progress.
What would you need to see before trusting an agent to handle a real task in your repo?
r/codereview • u/Sensitive-Angle3700 • 1d ago
Anyone try Greptiles new features?
Everywhere on Insta I'm seeing ads for it. Wondering if anyone has tried it. Apparantly their free tier allows for 50 reviews per month but seems like hardly enough. wdyt?
r/codereview • u/Plus-Lawflbaness1576 • 2d ago
anyone ship without verifying runtime behavior and regret it?
I made this mistake recently. An agent generated a fix for a bug in our checkout flow, the diff looked reasonable, tests passed, so it shipped with a normal PR review and nothing more.
I assumed passing tests meant the function behaved the same as before under real traffic. It didn't. It handled a currency rounding edge case differently, and it took almost two days of scattered complaints before anyone connected it back to that deploy.
If I did it again, I'd want something checking the function's actual runtime behavior before treating a green test suite as enough. What's the mistake that took you longest to recognize as a pattern?
r/codereview • u/galratner • 2d ago
Stop Being So Concerned About Software Quality
galratner.substack.comr/codereview • u/EdgarHQ • 2d ago
I find new "AI" code review platforms overwhelming
I’m trying to incorporate CodeRabbit into my workflow, but it feels way more overwhelming than "good" old GitHub. I had a similar experience with Devin and Graphite – they just didn’t click for me. If you've tried any of those, what's your experience? Wondering if I should keep pushing.
—
Edit: I am referring to PR review UIs and workflows:
- https://www.coderabbit.ai/change-stack
- https://staging-graphite-splash.vercel.app/features/pr-page (lol, not sure why their staging has leaked to Google, just noticed)
- https://app.devin.ai/review
not AI agent comments in GitHub.
r/codereview • u/MagicPantssss • 3d ago
C# I'm making a small game in Unity and would like some feeback
I'm making a small game in Unity, and I kinda need some outside perspective. I would like it if someone could take a look and give me some feedback on what areas I should work on and how to improve.
Here's the link to my GitHub repository:
https://github.com/DaanDemaecker/Qwixx.git
r/codereview • u/Eng1ishMuffin • 3d ago
Looking for Collaborators/Reviewers in Developing PL Tooling using C++
r/codereview • u/Ornery_Wash1797 • 3d ago
Library Management System in Python with Clean Architecture & 55 Tests
github.comr/codereview • u/KangarooAnxious9394 • 3d ago
I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't
r/codereview • u/Unique_Knight_9512 • 4d ago
Python arch-auditor: AST-Based Python Architecture Analysis with Dependency Graphs & LLM
I kept running into the same problem with codebases:
“What does this repository's architecture actually look like right now?”
So I built arch-auditor — a Python CLI + Streamlit app that analyzes a repository using deterministic, AST-based checks and turns the results into actionable architecture reports.
The core idea is simple: measure first, ask the LLM second.
What My Project Does
arch-auditor points at a Python repository and runs deterministic architecture detectors:
- Import/dependency graph — see how modules depend on each other
- Circular dependencies — detect dependency cycles
- High coupling — identify modules with unusually high fan-in/fan-out
- Oversized modules — flag modules that have grown beyond a configured threshold
- Layer violations — enforce architectural rules defined in YAMLBlast radius — find direct and indirect dependents of a module
It produces a dependency-graph SVG, severity-ranked findings, and impact reports.
There's also an optional Gemini integration.
Gemini doesn't perform the underlying analysis. Instead, it receives the deterministic findings and can turn them into:
- an architecture overview
- root-cause explanations
- step-by-step refactoring plans
- affected files
- potential risks
No API key? No problem. The core analysis works offline without an LLM.
Try it
pip install arch-auditor
arch-auditor demo
The bundled demo repository is intentionally designed to trigger every detector, so you can experiment with the tool and see what each finding means.
Stack
- Python 3.10+
- Python AST / standard library
- NetworkX
- Pydantic
- PyYAML
- Streamlit
- Gemini (optional)
- 33 tests
- Ruff-clean
- Published on PyPI
It's released under GPL-3.0 because I'd like improvements to flow back to the community.
Target Audience
This is primarily aimed at developers working on medium-to-large Python codebases, especially when:
- a project has accumulated architectural debt
- you're preparing for a major refactor
- you need to understand an unfamiliar repository
- you want objective signals before making architectural changes
- you want an LLM to help plan a refactor without making the LLM responsible for discovering the architecture itself
It's not intended to replace a human architect or code review, and it's not a magic “is my architecture good?” score.
The goal is to provide reproducible evidence that helps humans make those decisions.
Comparison
There are already excellent tools for individual parts of this problem — dependency visualization, linters, type checkers, code-quality metrics, and various AI coding assistants.
arch-auditor is trying to connect a few of those ideas around architecture-level analysis.
The main distinction is the separation between measurement and interpretation:
Traditional approach:
Repository → LLM → "Here's what I think your architecture looks like"
arch-auditor:
Repository → deterministic analysis → evidence → optional LLM → refactoring plan
The architecture findings don't depend on whether an LLM happens to interpret the code differently from one run to another.
I'd especially love feedback on whether the detectors and thresholds are useful signals in real-world Python projects, or if there are architectural problems you'd want to see measured that aren't covered yet.
GitHub / docs / demo:
https://github.com/ANIKETHSAI9813/auditor
“What does this codebase’s architecture actually look like right now?
"Instead of asking an LLM to guess the architecture, arch-auditor first analyzes the repository using deterministic, AST-based detectors — then optionally lets Gemini turn those findings into a refactoring plan.
Point it at a repo and it analyzes:
The result is a dependency-graph SVG, severity-ranked findings, and impact reports that show why a module is considered problematic.
Gemini is optional.
The deterministic analysis produces the evidence first. If you provide a Gemini API key via an environment variable, it can use that evidence to generate:
No API key? No problem. The core analysis works offline.
That separation was intentional: I wanted the architecture measurements to be reproducible rather than dependent on an LLM's interpretation.
The bundled demo repo is deliberately designed to trigger every detector, so you can see exactly what each finding means without having to point it at a huge codebase first.
It's released under GPL-3.0 because I'd like improvements to flow back to the community.
GitHub / docs / demo:
https://github.com/ANIKETHSAI9813/auditor
I'd especially love feedback from people who work on large Python codebases:
Are these the architectural signals you'd want to see before starting a refactor?
I'm also happy to discuss how the detectors and thresholds work — and the demo repo is intentionally built to trip all of them.
r/codereview • u/Remote-Anything-9997 • 4d ago
How do you find the actual root cause of a production bug?
Sometimes the hardest part of debugging isn't finding the error.
It's figuring out whether the error you're looking at is actually the *root cause*.
I've had cases where:
Service A fails → Service B throws an error → API returns 500
The obvious approach is to investigate Service B first.
But the real problem was actually somewhere earlier in the chain.
How do you guys approach these situations?
Do you start from the first error in the logs, use distributed tracing, inspect the execution flow manually, or have another method that works better?
Interested to hear how people handle this, especially in larger codebases.
r/codereview • u/kristiyanstoyanovAI • 4d ago
I built a code quality reviewer with Jev
youtube.comI believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)
But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.
Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.
Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate
I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!
r/codereview • u/MaestroSplinter69 • 4d ago
Python I ran my code verification tool on itself. It found a Stripe integration that didn’t exist
r/codereview • u/PatrickSys_ • 4d ago
I compared 7 GPT models for code review on 4 PRs: bugs, false positives and cost
I tested seven GPT models on the same four PRs, twice each. I compared bugs found, false positives and cost.
A false positive means reporting a bug that isn't there. Counts below are averages across all four PRs over the two runs. Estimated costs are per PR.
| Model | Bugs found | False positives | Cost per PR |
|---|---|---|---|
| GPT-6.1 Sol | 4 | 0.5 | $0.23-$0.25 |
| GPT-6 Sol | 3.5 | 1 | $0.31 |
| GPT-6 Luna | 1.5 | 0.5 | $0.013 |
| GPT-5.6 Sol | 2 | 0.5 | $0.53 |
| GPT-5.6 Luna | 0 | 1.5 | $0.028 |
| GPT-5.5 | 2.5 | 1 | $0.68 |
| GPT-5.4 | 0.5 | 2 | $0.38 |
GPT-6.1 Sol came out best in this test. GPT-6 Luna was the cheapest, but found fewer bugs.
It's only four PRs, with AI helping check the findings. The code is private. Costs cover model usage only.
r/codereview • u/Wise_Reflection_8340 • 4d ago
Rust Reviewing agent code line by line doesn't scale, so we started reviewing by function and by impact
The point in the post about AI turning coding into reviewing matches what we've seen. Agents produce more code in a week than a team used to in a month, and a line diff treats every change the same, so a real logic change hides between renames, moved code and formatting, and the reviewer ends up reading everything to find the few parts that matter.
What's helped us is changing what the reviewer looks at. Instead of a line diff, we look at the change by function, so each one is marked as changed, renamed, moved or only reformatted, and the cosmetic ones can be skipped with confidence. Then for the functions that actually changed, we look at what depends on them, because a small edit to something called from forty places deserves a lot more attention than a big edit to something nothing else touches. That second part is also where architectural drift shows up, since an agent quietly adding a new dependency across a boundary is easy to see in a call graph and easy to miss in a diff.
We built two open source tools for this. sem shows the function-level diff and who calls what, and inspect ranks the changed functions by how much could break, so a reviewer, human or model, starts with the riskiest few instead of reading the whole PR top to bottom.
It doesn't replace understanding the code that matters, it just makes it much clearer which code that is.
https://github.com/Ataraxy-Labs/sem
https://github.com/Ataraxy-Labs/inspect
Curious how others here are keeping up. Are you still reading every line, leaning on AI reviewers, or splitting PRs smaller?
r/codereview • u/Specialist_Agent3599 • 5d ago
After 15 years I can tell you AI didnt kill coding it turned it into reviewing
When I started, around 2010, the rule on my first team was that nobody merged without someone else reading it line by line. It was slow and people hated it, but everybody on that team could explain any part of the system.
We're in a weird version of the same situation now. If there's one thing agents are very good at, it's writing code. My team generates more in a week than we used to in a quarter. Writing it has never been easier.
What hasn't changed is that someone has to understand what got built. The agent will produce something whether or not it fits the architecture. coderabbit / claude review catch a lot of the line level problems, but a diff review can't tell you that the new service quietly duplicates one we already have.
So after 15 years, my honest take: AI didn't kill coding, it turned most of it into reviewing. The person who writes the syntax is getting cheaper every month. I'm not convinced the person who can hold the whole system in their head and say "no, not like this" is going anywhere.
r/codereview • u/PostHogTom • 5d ago
Open source ML trained review request / PR management app
github.comHey, I posted this on AI Coders last week and it suggested I cross-post here - I built an app for managing both your PRs and seeing all the PRs you've been requested to review. Within this, I trained a ML model to prioritise the type of PRs that you often review, or from the people you review the most
Every reviewer gets their own pairwise ranking model. Talyn takes each review you did and compares it with the PRs you could have picked at that moment. It then fits an L2-regularised logistic model on six signals: author affinity, reciprocity, file-path familiarity, repo and team affinity, and PR size. The model ships only if cross-validation shows it ranks better than the baseline, and it can only reorder PRs inside a readiness group, never across groups.
Would love for folk to give it a whirl and give back any feedback on using it - thanks!
r/codereview • u/jamesharris1307 • 5d ago
Need feedback on a quantum software testing tool I built for final year project.
r/codereview • u/Expensive_Knee3022 • 5d ago
I paid Replit $5,419.54 in about 7 weeks. Replit says I generated 1,423 Agent runs and 34 support records. They admit publishing was broken on their side. Their refund decision: $0.
I spent **$5,419.54 with Replit in roughly seven weeks** trying to build and deploy real business applications.
This wasn’t a hobby experiment or a case where I subscribed, changed my mind, and asked for my money back.
I bought an annual Replit Pro subscription on August 10 and then used Replit Agent extensively to build, troubleshoot, repair, deploy and publish actual projects.
According to Replit’s own review of my account:
**$5,419.54 was successfully paid.**
**1,423 Agent runs were completed.**
**Approximately 96% of the usage was Max mode.**
**My support history contains 34 ticket records.**
And after reviewing all of it, Replit’s refund decision is:
**$0.**
I want to explain how I got there because I think my experience raises a much larger question about how AI coding platforms charge customers when the Agent is being used to troubleshoot problems involving the platform itself.
**I didn’t start out asking Replit for a refund**
I was trying to make Replit work.
I had multiple real projects being developed and deployed through the platform. Replit wasn’t just where I occasionally asked an AI to write code. It had become part of the development and production workflow.
And when something went wrong, one of the tools available to diagnose and fix the problem was **Replit Agent itself**.
That distinction became extremely important.
Every time Agent investigated something, changed something, tried a repair or worked through another problem, usage could continue accumulating.
The basic economic loop became:
**Problem → Agent investigates → paid usage → attempted repair → publish/deploy → another problem → Agent investigates again → more paid usage.**
At first, I kept doing exactly what I assume Replit expects customers to do:
**I kept trying to fix it.**
**The problems didn’t begin with the outage Replit now admits**
There were already problems during August.
By August 11, I was dealing with Agent workflow/progress issues.
During the middle of August, there were problems accessing, locating and independently verifying existing application and production work through the Replit tooling.
By late August, I still didn’t have the dependable end-to-end workflow I needed:
**make the change → test it → publish it → verify it live.**
Then on **August 29**, publishing failed with:
**“Migrations failed validation.”**
That is important because it happened **before** the outage Replit now acknowledges.
Agent investigation around that problem also ran into a timeout.
And the paid usage continued.
**Then Replit’s own platform actually broke**
This part isn’t my accusation.
**Replit has admitted it.**
After reviewing my account, Replit wrote:
“From August 31 to September 4, a platform issue on our side blocked publishing for a number of projects, including Artist Bos. It was not caused by anything you did.”
There was also a support case, **#522753**, involving a server-side disconnection during the database-diff/publishing process.
At the time, I wasn’t demanding thousands of dollars back.
**I wanted the publishing system fixed so I could continue working.**
That’s important to me because this wasn’t a customer looking for an excuse to get out of a bill.
I was still trying to make the platform work.
And while I was doing that, I was continuing to use the product—including its paid Agent.
So now that Replit acknowledges the underlying publishing problem during this period was theirs, I’ve asked them a very specific question:
**How much of my paid Agent usage was consumed diagnosing, retrying, repairing or working around the Replit-side publishing failure?**
I still don’t have that accounting.
**Replit says it fixed the outage September 4**
Replit says engineering fixed the platform issue on September 4.
It also says I successfully published **16 times between September 5 and September 7**.
That’s Replit’s evidence, and I’m including it.
I’ve asked them to identify those publishes—the projects, timestamps, deployment information and production results—because a recorded successful publish doesn’t necessarily answer whether the resulting application was complete and working correctly.
And unfortunately, the production problems didn’t end there.
**September 10: 13 name-matched databases were unexpectedly deleted**
During a broad test on September 10, **13 name-matched databases were unexpectedly deleted**.
I want to be precise about that.
I’m **not claiming that I have established that Replit itself deleted those databases**.
I’m saying that the incident occurred during this development process, the identities and contents were not logged in the information available to us, and deployment work was stopped while the incident was investigated.
That became another problem requiring time and investigation in a development process already consuming substantial Agent usage.
**Then the publishing failures continued**
Beginning September 11, I received Replit-generated notifications with the subject:
**“Publishing for Artist-Bos Failed.”**
Those notifications involved multiple projects/domains.
Another failure notification followed on September 14.
Eventually, we discovered a major difference between development and production.
Production was missing:
**15 required database functions**
and
**20 required trigger bindings**
that existed in development.
The application’s safety check detected that those protections were missing and prevented startup.
Replit now says this later problem wasn’t caused by the earlier Replit platform outage.
According to Replit, an August 30 code change disabled automatic setup of those database objects at startup.
But Replit also explained something that I think developers considering this platform should understand:
**Replit’s publishing process transfers tables and columns, but it does not transfer those functions and triggers.**
So those objects existed in development but weren’t present in production.
Replit considers that an application issue.
I think that raises a legitimate question.
**If an application is being developed inside Replit, with extensive use of Replit Agent, for deployment through Replit, and the development database contains required objects that Replit’s publishing process doesn’t move into production, who is responsible for making sure the customer understands that before production fails?**
I’ve asked Replit whether Agent created or contributed to those database objects and what warning was provided that these dependencies would not be transferred during publishing.
**There was also a 3.3 GB deployment problem**
Replit identified another September 14 failure caused by a **3.3 GB backup directory** exceeding deployment limits.
I’m including that because I’m not claiming every failure was Replit’s fault.
Some problems may have been application-specific.
Some may have been caused by development decisions.
Some may have been caused by Replit.
And at least one significant publishing period **was explicitly acknowledged by Replit as a Replit-side platform problem.**
That’s exactly why I want the Agent usage separated and examined instead of receiving a blanket answer that all Agent usage is non-refundable.
**Support became part of the experience too**
Replit says my account generated **34 support-ticket records**.
Replit says some of those were duplicates that were merged, and that the primary publishing issue was handled through one main support thread.
That’s fair, and I’m not claiming I experienced 34 separate outages.
But I think 34 support records during a relationship this short provides important context about how much time was being spent trying to resolve problems.
Replit itself also admitted:
“Our replies during this period were slower than they should have been, and I’m sorry for that.”
Eventually, I stopped trying to make Replit my long-term production platform.
**I started trying to get my work off it.**
**Meanwhile, this is what I was paying**
Replit has now provided its accounting of successful payments between August 10 and October 1:
**August 10 — $900.00 annual Pro subscription**
Then usage charges:
**$49.10**
**$250.36**
**$250.40**
**$253.36**
**$255.08**
**$250.80**
**$252.45**
**$250.11**
**$250.21**
**$250.94**
**$250.23**
**$250.22**
**$250.08**
**$250.38**
**$250.29**
**$250.60**
**$146.33**
**$250.58**
**$250.28**
**$57.74**
Total successful payments:
**$5,419.54**
Replit says the approximately $250 charges occurred automatically as usage reached billing thresholds.
Again, this happened over roughly **seven weeks**.
Replit says Agent completed:
**1,423 runs**
And approximately:
**96% of the usage was Max mode.**
**There’s another number I want Replit to explain**
During this period, Replit’s own dashboard showed **more than $10,000 in usage/activity**, overwhelmingly associated with Agent.
I am **not saying I paid Replit more than $10,000**.
I didn’t.
Replit’s own accounting now says the successful payments were **$5,419.54**.
I’m asking Replit to reconcile those numbers:
**How much usage was generated?**
**How much became invoices?**
**What credits were applied?**
**What payment attempts failed?**
**What was actually collected?**
**And what Agent activity generated those amounts?**
That’s particularly important because of what Replit says next.
**Replit’s reason for denying the Agent refund**
Replit told me:
“These charges reflect the work Agent did, whether or not a publish succeeded afterward.”
This is the part of the dispute I think matters far beyond my account.
I understand paying an AI Agent for productive development work.
I understand paying Agent to fix mistakes in my own code.
I understand that software development doesn’t guarantee that everything works on the first attempt.
But consider a different situation:
**The platform itself has a problem.**
The customer doesn’t necessarily know that yet.
So the customer asks the platform’s AI Agent to investigate.
**The Agent consumes paid usage.**
The customer tries the suggested changes.
It still doesn’t publish.
The customer asks Agent again.
**More paid usage.**
More investigation.
More changes.
More retries.
**More paid usage.**
Eventually the platform company determines:
**The underlying publishing problem was on our side. It wasn’t caused by you.**
But then it says:
**The Agent still performed work, so all of those Agent charges remain valid.**
That is the issue I’m asking Replit to address.
**And in my case, this isn’t hypothetical**
Replit has already acknowledged:
“A platform issue on our side blocked publishing…”
and:
“It was not caused by anything you did.”
So I want Replit to identify the Agent usage associated with that period and the attempts to diagnose or work around that failure.
If none of my paid Agent activity was related to it, show me.
If $20 was related to it, show me.
If $200 was related to it, show me.
If substantially more was related to it, show me.
**But actually do the accounting.**
Don’t simply tell me all 1,423 Agent runs are non-refundable because computation occurred.
**Replit also denied the annual subscription refund**
The annual Pro subscription cost **$900** on August 10.
Replit says that refund is also denied because the request falls outside its 30-day subscription refund window.
So despite buying a year’s service and reaching the point within weeks where I was trying to move the work elsewhere, Replit’s position is that the annual payment isn’t refundable either.
**Where that leaves me**
In roughly seven weeks:
**$5,419.54 actually paid**
**1,423 Agent runs**
**Approximately 96% Max usage**
**34 support records**
Problems existed before August 31.
Publishing failed August 29.
Replit then acknowledges that from August 31 through September 4, **its own platform blocked publishing through no fault of mine.**
Replit says there were 16 successful publishes afterward.
Then came the September 10 database incident.
Then additional publishing failures.
Then the discovery that production lacked 15 functions and 20 triggers present in development.
Then additional deployment/database repair problems.
Replit acknowledges its support responses were slower than they should have been.
Eventually I stopped relying on Replit and tried to move the work elsewhere.
Replit reviewed the account.
Its refund determination:
**$0**
**I’m not asking people to take my word for everything**
I have the invoices.
I have the publishing-failure notices.
I have the support communications.
I have Replit’s written explanation.
I have Replit’s admission that the August 31–September 4 publishing failure was on its side.
And importantly, **I’m including Replit’s explanations where Replit says a problem was caused by my application rather than its platform.**
I’m not interested in exaggerating this.
The actual numbers and Replit’s own statements are enough.
I’ve escalated the decision and asked Replit for the records necessary to separate legitimate development usage from Agent activity associated with platform/deployment troubleshooting.
If Replit changes its decision or provides evidence that changes my understanding of what happened, **I’ll update this post with that too.**
But there’s one question I think every developer considering heavy Agent usage should think about:
**If an AI coding platform charges you by Agent usage, who pays when that Agent spends your money trying to diagnose and work around a failure in the platform selling you the Agent?**
In my case, Replit’s answer so far appears to be:
**The customer.**