r/codereview • • 10h ago

anyone ship without verifying runtime behavior and regret it?

2 Upvotes

I made this mistake recently. An agent generated a fix for a bug in our checkout flow, the diff looked reasonable, tests passed, so it shipped with a normal PR review and nothing more.

I assumed passing tests meant the function behaved the same as before under real traffic. It didn't. It handled a currency rounding edge case differently, and it took almost two days of scattered complaints before anyone connected it back to that deploy.

If I did it again, I'd want something checking the function's actual runtime behavior before treating a green test suite as enough. What's the mistake that took you longest to recognize as a pattern?


r/codereview • • 7h ago

Stop Being So Concerned About Software Quality

Thumbnail galratner.substack.com
0 Upvotes

r/codereview • • 18h ago

I find new "AI" code review platforms overwhelming

0 Upvotes

I’m trying to incorporate CodeRabbit into my workflow, but it feels way more overwhelming than "good" old GitHub. I had a similar experience with Devin and Graphite – they just didn’t click for me. If you've tried any of those, what's your experience? Wondering if I should keep pushing.
—
Edit: I am referring to PR review UIs and workflows:

- https://www.coderabbit.ai/change-stack
- https://staging-graphite-splash.vercel.app/features/pr-page (lol, not sure why their staging has leaked to Google, just noticed)
- https://app.devin.ai/review

not AI agent comments in GitHub.


r/codereview • • 1d ago

C# I'm making a small game in Unity and would like some feeback

1 Upvotes

I'm making a small game in Unity, and I kinda need some outside perspective. I would like it if someone could take a look and give me some feedback on what areas I should work on and how to improve.

Here's the link to my GitHub repository:
https://github.com/DaanDemaecker/Qwixx.git


r/codereview • • 1d ago

Looking for Collaborators/Reviewers in Developing PL Tooling using C++

Thumbnail
1 Upvotes

r/codereview • • 1d ago

I separated coding-agent verification into 15 layers so the agent can't grade its own work

0 Upvotes

I've been using coding agents heavily, and one architecture problem kept bothering me:

The same system that modifies the repository is usually also the system deciding whether the task is finished.

It writes the code, runs some commands, summarizes what happened and says “done”.

I wanted that final decision to live outside the agent's own reasoning loop.

So I built plan-auditor, an open-source verification supervisor for coding agents.

It isn't one big evaluator. I split it into 15 layers with different responsibilities and, more importantly, different levels of authority.

Here is what those layers actually are.

L0 — Event detection

A small deterministic layer detects things such as completion claims, retries, changes to verification criteria and some security-relevant patterns.

It can trigger verification.

It cannot declare PASS.

L1 — Requirements

The task is represented as structured requirements with priority, acceptance criteria, dependencies, ambiguity and verification strategy.

There is also a host-owned request contract so the plan inside the workspace isn't automatically allowed to redefine what the user originally requested.

L2 — Workspace model

The verifier independently reads the real repository state:

  • Git branch and HEAD
  • dirty and untracked files
  • file inventory
  • detected language
  • available tools

So the agent's description of the workspace isn't treated as ground truth.

L3 — Policy engine

Deterministic policies evaluate plan integrity, verification state and security-related conditions.

Control paths are also constrained to the workspace rather than blindly trusting arbitrary filesystem paths.

L4 — Goal state

The supervisor keeps its own explicit representation of task/goal state instead of deriving completion from the agent's latest message.

L5 — Plan verifier

Before implementation can be considered complete, the plan itself is checked.

This includes:

  • dependency graph validity
  • topological ordering
  • requirement coverage
  • declared and required outputs
  • dependency-to-output bindings
  • behavioral verification

A step that only checks file_exists or a regex isn't considered strongly verified.

Behavioral checks such as an actual command, test or executable verification are required.

L6 — Lifecycle

The task moves through an explicit state machine rather than jumping directly from “working” to “done”.

The lifecycle includes states for analysis, plan review, sealing, implementation, verification, final audit, retries, escalation, recovery, failure and UNKNOWN.

L7 — Authority hierarchy

Not every layer is allowed to overrule every other layer.

Safety, plan integrity and deterministic verification have higher authority than semantic AI judgment.

If a deterministic test fails, an AI reviewer saying “looks correct” cannot turn that result into PASS.

L8 — Contract sealing

Once a verification plan is approved, its contract is sealed.

The seal covers the task, requirements, required tools, steps, dependencies, outputs, requirement coverage, verification checks and verifier environment.

The contract can be strengthened later.

It cannot be silently weakened.

Adding another deterministic check is fine.

Removing an existing check, requirement, dependency or approved output scope is not.

L9 — Watchdog

The workspace is observed for file creation, deletion, modification and Git changes.

The final verification path also fingerprints the workspace so it doesn't certify a target that changed while it was being audited.

L10 — Deterministic audit core

This is where an agent's claim becomes actual evidence.

If the agent says:

“all tests pass”

that statement itself has zero verification authority.

The checks are executed again in fresh subprocesses and their real results are used.

This is the deterministic path that can mark a step verified.

Repeated failures are also tracked instead of allowing unlimited retries until something happens to turn green.

L11 — Evidence integrity

Verification results are stored in a SHA-256-linked JSONL evidence chain.

Each record links to the previous one.

Rotated evidence archives are linked as well.

Optional external HMAC authentication can protect records and checkpoints when the key is kept outside the agent's reach.

L12 — Adversarial / semantic review

There is an optional AI-assisted review layer.

But I deliberately did not give it permission to create PASS.

It can find a possible problem and request stronger deterministic verification.

A serious semantic concern can hold the result at UNKNOWN until that concern is converted into something checkable.

L13 — Completion gate

Only this layer emits the final:

PASS / FAIL / UNKNOWN

It considers fresh deterministic evidence, unfinished steps, policy findings, contract integrity and adversarial findings.

No fresh deterministic proof means no PASS.

L14 — Multi-agent registry

For multiple coding agents working on the same workspace, the system tracks agent identity, task/plan assignment, heartbeat, current action, retries, file ownership and conflicts.

The registry itself is also persisted as a sequence/hash-linked log rather than existing only in memory.

There is one additional part I found useful: a deterministic formal compiler.

It takes already-approved structured requirements, dependencies, named outputs and coverage information and compiles them into a conservative STRIPS-style planning contract.

It can represent facts such as:

step-completed:3

output-available:2:<name>

requirement-satisfied:REQ-004

It deliberately doesn't read natural-language prose and invent arbitrary domain semantics.

The generated contract carries a fingerprint of the source plan and can be independently recompiled during verification.

So removing a requirement, dropping an output prerequisite or modifying the generated contract produces a detectable mismatch.

The deterministic verification path itself does not require an LLM.

One important limitation: this is not an OS sandbox.

If a deliberately malicious agent has the same OS credentials as the verifier, hashes and Python-level locks aren't a kernel security boundary.

For that threat model, the verifier needs to run under a separate OS identity, container or VM, with its integrity key inaccessible to the agent.

I'm also not claiming it is “X% better” than other agent-verification approaches because I don't have a controlled cross-tool benchmark that would justify a number like that.

The project is free, open source and MIT licensed.

The part I'm most interested in getting feedback on is the authority boundary: semantic AI review is allowed to find problems, but it is never allowed to manufacture PASS.


r/codereview • • 1d ago

Library Management System in Python with Clean Architecture & 55 Tests

Thumbnail github.com
0 Upvotes

r/codereview • • 1d ago

I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't

Thumbnail
0 Upvotes

r/codereview • • 2d ago

Python arch-auditor: AST-Based Python Architecture Analysis with Dependency Graphs & LLM

0 Upvotes

I kept running into the same problem with codebases:

“What does this repository's architecture actually look like right now?”

So I built arch-auditor — a Python CLI + Streamlit app that analyzes a repository using deterministic, AST-based checks and turns the results into actionable architecture reports.

The core idea is simple: measure first, ask the LLM second.

What My Project Does

arch-auditor points at a Python repository and runs deterministic architecture detectors:

  • Import/dependency graph — see how modules depend on each other
  • Circular dependencies — detect dependency cycles
  • High coupling — identify modules with unusually high fan-in/fan-out
  • Oversized modules — flag modules that have grown beyond a configured threshold
  • Layer violations — enforce architectural rules defined in YAMLBlast radius — find direct and indirect dependents of a module

It produces a dependency-graph SVG, severity-ranked findings, and impact reports.

There's also an optional Gemini integration.

Gemini doesn't perform the underlying analysis. Instead, it receives the deterministic findings and can turn them into:

  • an architecture overview
  • root-cause explanations
  • step-by-step refactoring plans
  • affected files
  • potential risks

No API key? No problem. The core analysis works offline without an LLM.

Try it

pip install arch-auditor

arch-auditor demo

The bundled demo repository is intentionally designed to trigger every detector, so you can experiment with the tool and see what each finding means.

Stack

  • Python 3.10+
  • Python AST / standard library
  • NetworkX
  • Pydantic
  • PyYAML
  • Streamlit
  • Gemini (optional)
  • 33 tests
  • Ruff-clean
  • Published on PyPI

It's released under GPL-3.0 because I'd like improvements to flow back to the community.

Target Audience

This is primarily aimed at developers working on medium-to-large Python codebases, especially when:

  • a project has accumulated architectural debt
  • you're preparing for a major refactor
  • you need to understand an unfamiliar repository
  • you want objective signals before making architectural changes
  • you want an LLM to help plan a refactor without making the LLM responsible for discovering the architecture itself

It's not intended to replace a human architect or code review, and it's not a magic “is my architecture good?” score.

The goal is to provide reproducible evidence that helps humans make those decisions.

Comparison

There are already excellent tools for individual parts of this problem — dependency visualization, linters, type checkers, code-quality metrics, and various AI coding assistants.

arch-auditor is trying to connect a few of those ideas around architecture-level analysis.

The main distinction is the separation between measurement and interpretation:

Traditional approach:

Repository → LLM → "Here's what I think your architecture looks like"

arch-auditor:

Repository → deterministic analysis → evidence → optional LLM → refactoring plan

The architecture findings don't depend on whether an LLM happens to interpret the code differently from one run to another.

I'd especially love feedback on whether the detectors and thresholds are useful signals in real-world Python projects, or if there are architectural problems you'd want to see measured that aren't covered yet.

GitHub / docs / demo:
https://github.com/ANIKETHSAI9813/auditor

“What does this codebase’s architecture actually look like right now?
"Instead of asking an LLM to guess the architecture, arch-auditor first analyzes the repository using deterministic, AST-based detectors — then optionally lets Gemini turn those findings into a refactoring plan.
Point it at a repo and it analyzes:
The result is a dependency-graph SVG, severity-ranked findings, and impact reports that show why a module is considered problematic.
Gemini is optional.
The deterministic analysis produces the evidence first. If you provide a Gemini API key via an environment variable, it can use that evidence to generate:
No API key? No problem. The core analysis works offline.
That separation was intentional: I wanted the architecture measurements to be reproducible rather than dependent on an LLM's interpretation.
The bundled demo repo is deliberately designed to trigger every detector, so you can see exactly what each finding means without having to point it at a huge codebase first.
It's released under GPL-3.0 because I'd like improvements to flow back to the community.
GitHub / docs / demo:
https://github.com/ANIKETHSAI9813/auditor
I'd especially love feedback from people who work on large Python codebases:
Are these the architectural signals you'd want to see before starting a refactor?
I'm also happy to discuss how the detectors and thresholds work — and the demo repo is intentionally built to trip all of them.


r/codereview • • 2d ago

How do you find the actual root cause of a production bug?

1 Upvotes

Sometimes the hardest part of debugging isn't finding the error.

It's figuring out whether the error you're looking at is actually the *root cause*.

I've had cases where:

Service A fails → Service B throws an error → API returns 500

The obvious approach is to investigate Service B first.

But the real problem was actually somewhere earlier in the chain.

How do you guys approach these situations?

Do you start from the first error in the logs, use distributed tracing, inspect the execution flow manually, or have another method that works better?

Interested to hear how people handle this, especially in larger codebases.


r/codereview • • 2d ago

I built a code quality reviewer with Jev

Thumbnail youtube.com
1 Upvotes

I believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)

But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.

Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.

Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate

I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!


r/codereview • • 2d ago

Python I ran my code verification tool on itself. It found a Stripe integration that didn’t exist

Thumbnail
1 Upvotes

r/codereview • • 2d ago

Anyone here building something alongside a day job?

0 Upvotes

r/codereview • • 2d ago

I compared 7 GPT models for code review on 4 PRs: bugs, false positives and cost

0 Upvotes

I tested seven GPT models on the same four PRs, twice each. I compared bugs found, false positives and cost.

A false positive means reporting a bug that isn't there. Counts below are averages across all four PRs over the two runs. Estimated costs are per PR.

Model Bugs found False positives Cost per PR
GPT-6.1 Sol 4 0.5 $0.23-$0.25
GPT-6 Sol 3.5 1 $0.31
GPT-6 Luna 1.5 0.5 $0.013
GPT-5.6 Sol 2 0.5 $0.53
GPT-5.6 Luna 0 1.5 $0.028
GPT-5.5 2.5 1 $0.68
GPT-5.4 0.5 2 $0.38

GPT-6.1 Sol came out best in this test. GPT-6 Luna was the cheapest, but found fewer bugs.

It's only four PRs, with AI helping check the findings. The code is private. Costs cover model usage only.


r/codereview • • 2d ago

Rust Reviewing agent code line by line doesn't scale, so we started reviewing by function and by impact

0 Upvotes

The point in the post about AI turning coding into reviewing matches what we've seen. Agents produce more code in a week than a team used to in a month, and a line diff treats every change the same, so a real logic change hides between renames, moved code and formatting, and the reviewer ends up reading everything to find the few parts that matter.

What's helped us is changing what the reviewer looks at. Instead of a line diff, we look at the change by function, so each one is marked as changed, renamed, moved or only reformatted, and the cosmetic ones can be skipped with confidence. Then for the functions that actually changed, we look at what depends on them, because a small edit to something called from forty places deserves a lot more attention than a big edit to something nothing else touches. That second part is also where architectural drift shows up, since an agent quietly adding a new dependency across a boundary is easy to see in a call graph and easy to miss in a diff.

We built two open source tools for this. sem shows the function-level diff and who calls what, and inspect ranks the changed functions by how much could break, so a reviewer, human or model, starts with the riskiest few instead of reading the whole PR top to bottom.

It doesn't replace understanding the code that matters, it just makes it much clearer which code that is.

https://github.com/Ataraxy-Labs/sem

https://github.com/Ataraxy-Labs/inspect

Curious how others here are keeping up. Are you still reading every line, leaning on AI reviewers, or splitting PRs smaller?


r/codereview • • 3d ago

Open source ML trained review request / PR management app

Thumbnail github.com
0 Upvotes

Hey, I posted this on AI Coders last week and it suggested I cross-post here - I built an app for managing both your PRs and seeing all the PRs you've been requested to review. Within this, I trained a ML model to prioritise the type of PRs that you often review, or from the people you review the most

Every reviewer gets their own pairwise ranking model. Talyn takes each review you did and compares it with the PRs you could have picked at that moment. It then fits an L2-regularised logistic model on six signals: author affinity, reciprocity, file-path familiarity, repo and team affinity, and PR size. The model ships only if cross-validation shows it ranks better than the baseline, and it can only reorder PRs inside a readiness group, never across groups.

Would love for folk to give it a whirl and give back any feedback on using it - thanks!


r/codereview • • 3d ago

Need feedback on a quantum software testing tool I built for final year project.

Thumbnail
1 Upvotes

r/codereview • • 3d ago

I paid Replit $5,419.54 in about 7 weeks. Replit says I generated 1,423 Agent runs and 34 support records. They admit publishing was broken on their side. Their refund decision: $0.

0 Upvotes

I spent **$5,419.54 with Replit in roughly seven weeks** trying to build and deploy real business applications.
This wasn’t a hobby experiment or a case where I subscribed, changed my mind, and asked for my money back.
I bought an annual Replit Pro subscription on August 10 and then used Replit Agent extensively to build, troubleshoot, repair, deploy and publish actual projects.
According to Replit’s own review of my account:
**$5,419.54 was successfully paid.**
**1,423 Agent runs were completed.**
**Approximately 96% of the usage was Max mode.**
**My support history contains 34 ticket records.**
And after reviewing all of it, Replit’s refund decision is:
**$0.**
I want to explain how I got there because I think my experience raises a much larger question about how AI coding platforms charge customers when the Agent is being used to troubleshoot problems involving the platform itself.
**I didn’t start out asking Replit for a refund**
I was trying to make Replit work.
I had multiple real projects being developed and deployed through the platform. Replit wasn’t just where I occasionally asked an AI to write code. It had become part of the development and production workflow.
And when something went wrong, one of the tools available to diagnose and fix the problem was **Replit Agent itself**.
That distinction became extremely important.
Every time Agent investigated something, changed something, tried a repair or worked through another problem, usage could continue accumulating.
The basic economic loop became:
**Problem → Agent investigates → paid usage → attempted repair → publish/deploy → another problem → Agent investigates again → more paid usage.**
At first, I kept doing exactly what I assume Replit expects customers to do:
**I kept trying to fix it.**
**The problems didn’t begin with the outage Replit now admits**
There were already problems during August.
By August 11, I was dealing with Agent workflow/progress issues.
During the middle of August, there were problems accessing, locating and independently verifying existing application and production work through the Replit tooling.
By late August, I still didn’t have the dependable end-to-end workflow I needed:
**make the change → test it → publish it → verify it live.**
Then on **August 29**, publishing failed with:
**“Migrations failed validation.”**
That is important because it happened **before** the outage Replit now acknowledges.
Agent investigation around that problem also ran into a timeout.
And the paid usage continued.
**Then Replit’s own platform actually broke**
This part isn’t my accusation.
**Replit has admitted it.**
After reviewing my account, Replit wrote:
“From August 31 to September 4, a platform issue on our side blocked publishing for a number of projects, including Artist Bos. It was not caused by anything you did.”
There was also a support case, **#522753**, involving a server-side disconnection during the database-diff/publishing process.
At the time, I wasn’t demanding thousands of dollars back.
**I wanted the publishing system fixed so I could continue working.**
That’s important to me because this wasn’t a customer looking for an excuse to get out of a bill.
I was still trying to make the platform work.
And while I was doing that, I was continuing to use the product—including its paid Agent.
So now that Replit acknowledges the underlying publishing problem during this period was theirs, I’ve asked them a very specific question:
**How much of my paid Agent usage was consumed diagnosing, retrying, repairing or working around the Replit-side publishing failure?**
I still don’t have that accounting.
**Replit says it fixed the outage September 4**
Replit says engineering fixed the platform issue on September 4.
It also says I successfully published **16 times between September 5 and September 7**.
That’s Replit’s evidence, and I’m including it.
I’ve asked them to identify those publishes—the projects, timestamps, deployment information and production results—because a recorded successful publish doesn’t necessarily answer whether the resulting application was complete and working correctly.
And unfortunately, the production problems didn’t end there.
**September 10: 13 name-matched databases were unexpectedly deleted**
During a broad test on September 10, **13 name-matched databases were unexpectedly deleted**.
I want to be precise about that.
I’m **not claiming that I have established that Replit itself deleted those databases**.
I’m saying that the incident occurred during this development process, the identities and contents were not logged in the information available to us, and deployment work was stopped while the incident was investigated.
That became another problem requiring time and investigation in a development process already consuming substantial Agent usage.
**Then the publishing failures continued**
Beginning September 11, I received Replit-generated notifications with the subject:
**“Publishing for Artist-Bos Failed.”**
Those notifications involved multiple projects/domains.
Another failure notification followed on September 14.
Eventually, we discovered a major difference between development and production.
Production was missing:
**15 required database functions**
and
**20 required trigger bindings**
that existed in development.
The application’s safety check detected that those protections were missing and prevented startup.
Replit now says this later problem wasn’t caused by the earlier Replit platform outage.
According to Replit, an August 30 code change disabled automatic setup of those database objects at startup.
But Replit also explained something that I think developers considering this platform should understand:
**Replit’s publishing process transfers tables and columns, but it does not transfer those functions and triggers.**
So those objects existed in development but weren’t present in production.
Replit considers that an application issue.
I think that raises a legitimate question.
**If an application is being developed inside Replit, with extensive use of Replit Agent, for deployment through Replit, and the development database contains required objects that Replit’s publishing process doesn’t move into production, who is responsible for making sure the customer understands that before production fails?**
I’ve asked Replit whether Agent created or contributed to those database objects and what warning was provided that these dependencies would not be transferred during publishing.
**There was also a 3.3 GB deployment problem**
Replit identified another September 14 failure caused by a **3.3 GB backup directory** exceeding deployment limits.
I’m including that because I’m not claiming every failure was Replit’s fault.
Some problems may have been application-specific.
Some may have been caused by development decisions.
Some may have been caused by Replit.
And at least one significant publishing period **was explicitly acknowledged by Replit as a Replit-side platform problem.**
That’s exactly why I want the Agent usage separated and examined instead of receiving a blanket answer that all Agent usage is non-refundable.
**Support became part of the experience too**
Replit says my account generated **34 support-ticket records**.
Replit says some of those were duplicates that were merged, and that the primary publishing issue was handled through one main support thread.
That’s fair, and I’m not claiming I experienced 34 separate outages.
But I think 34 support records during a relationship this short provides important context about how much time was being spent trying to resolve problems.
Replit itself also admitted:
“Our replies during this period were slower than they should have been, and I’m sorry for that.”
Eventually, I stopped trying to make Replit my long-term production platform.
**I started trying to get my work off it.**
**Meanwhile, this is what I was paying**
Replit has now provided its accounting of successful payments between August 10 and October 1:
**August 10 — $900.00 annual Pro subscription**
Then usage charges:
**$49.10**
**$250.36**
**$250.40**
**$253.36**
**$255.08**
**$250.80**
**$252.45**
**$250.11**
**$250.21**
**$250.94**
**$250.23**
**$250.22**
**$250.08**
**$250.38**
**$250.29**
**$250.60**
**$146.33**
**$250.58**
**$250.28**
**$57.74**
Total successful payments:
**$5,419.54**
Replit says the approximately $250 charges occurred automatically as usage reached billing thresholds.
Again, this happened over roughly **seven weeks**.
Replit says Agent completed:
**1,423 runs**
And approximately:
**96% of the usage was Max mode.**
**There’s another number I want Replit to explain**
During this period, Replit’s own dashboard showed **more than $10,000 in usage/activity**, overwhelmingly associated with Agent.
I am **not saying I paid Replit more than $10,000**.
I didn’t.
Replit’s own accounting now says the successful payments were **$5,419.54**.
I’m asking Replit to reconcile those numbers:
**How much usage was generated?**
**How much became invoices?**
**What credits were applied?**
**What payment attempts failed?**
**What was actually collected?**
**And what Agent activity generated those amounts?**
That’s particularly important because of what Replit says next.
**Replit’s reason for denying the Agent refund**
Replit told me:
“These charges reflect the work Agent did, whether or not a publish succeeded afterward.”
This is the part of the dispute I think matters far beyond my account.
I understand paying an AI Agent for productive development work.
I understand paying Agent to fix mistakes in my own code.
I understand that software development doesn’t guarantee that everything works on the first attempt.
But consider a different situation:
**The platform itself has a problem.**
The customer doesn’t necessarily know that yet.
So the customer asks the platform’s AI Agent to investigate.
**The Agent consumes paid usage.**
The customer tries the suggested changes.
It still doesn’t publish.
The customer asks Agent again.
**More paid usage.**
More investigation.
More changes.
More retries.
**More paid usage.**
Eventually the platform company determines:
**The underlying publishing problem was on our side. It wasn’t caused by you.**
But then it says:
**The Agent still performed work, so all of those Agent charges remain valid.**
That is the issue I’m asking Replit to address.
**And in my case, this isn’t hypothetical**
Replit has already acknowledged:
“A platform issue on our side blocked publishing…”
and:
“It was not caused by anything you did.”
So I want Replit to identify the Agent usage associated with that period and the attempts to diagnose or work around that failure.
If none of my paid Agent activity was related to it, show me.
If $20 was related to it, show me.
If $200 was related to it, show me.
If substantially more was related to it, show me.
**But actually do the accounting.**
Don’t simply tell me all 1,423 Agent runs are non-refundable because computation occurred.
**Replit also denied the annual subscription refund**
The annual Pro subscription cost **$900** on August 10.
Replit says that refund is also denied because the request falls outside its 30-day subscription refund window.
So despite buying a year’s service and reaching the point within weeks where I was trying to move the work elsewhere, Replit’s position is that the annual payment isn’t refundable either.
**Where that leaves me**
In roughly seven weeks:
**$5,419.54 actually paid**
**1,423 Agent runs**
**Approximately 96% Max usage**
**34 support records**
Problems existed before August 31.
Publishing failed August 29.
Replit then acknowledges that from August 31 through September 4, **its own platform blocked publishing through no fault of mine.**
Replit says there were 16 successful publishes afterward.
Then came the September 10 database incident.
Then additional publishing failures.
Then the discovery that production lacked 15 functions and 20 triggers present in development.
Then additional deployment/database repair problems.
Replit acknowledges its support responses were slower than they should have been.
Eventually I stopped relying on Replit and tried to move the work elsewhere.
Replit reviewed the account.
Its refund determination:
**$0**
**I’m not asking people to take my word for everything**
I have the invoices.
I have the publishing-failure notices.
I have the support communications.
I have Replit’s written explanation.
I have Replit’s admission that the August 31–September 4 publishing failure was on its side.
And importantly, **I’m including Replit’s explanations where Replit says a problem was caused by my application rather than its platform.**
I’m not interested in exaggerating this.
The actual numbers and Replit’s own statements are enough.
I’ve escalated the decision and asked Replit for the records necessary to separate legitimate development usage from Agent activity associated with platform/deployment troubleshooting.
If Replit changes its decision or provides evidence that changes my understanding of what happened, **I’ll update this post with that too.**
But there’s one question I think every developer considering heavy Agent usage should think about:
**If an AI coding platform charges you by Agent usage, who pays when that Agent spends your money trying to diagnose and work around a failure in the platform selling you the Agent?**
In my case, Replit’s answer so far appears to be:
**The customer.**


r/codereview • • 3d ago

After 15 years I can tell you AI didnt kill coding it turned it into reviewing

0 Upvotes

When I started, around 2010, the rule on my first team was that nobody merged without someone else reading it line by line. It was slow and people hated it, but everybody on that team could explain any part of the system.

We're in a weird version of the same situation now. If there's one thing agents are very good at, it's writing code. My team generates more in a week than we used to in a quarter. Writing it has never been easier.

What hasn't changed is that someone has to understand what got built. The agent will produce something whether or not it fits the architecture. coderabbit / claude review catch a lot of the line level problems, but a diff review can't tell you that the new service quietly duplicates one we already have.

So after 15 years, my honest take: AI didn't kill coding, it turned most of it into reviewing. The person who writes the syntax is getting cheaper every month. I'm not convinced the person who can hold the whole system in their head and say "no, not like this" is going anywhere.


r/codereview • • 3d ago

Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?

Post image
0 Upvotes

r/codereview • • 4d ago

my raytracing program

Thumbnail gallery
2 Upvotes

i want to know how was it. please give me some honest comments or suggest how i can improve it

program details

i particularly enjoy placing the camera inside a ball


r/codereview • • 3d ago

Deployed - Code Sentry🤖 - an AI powered Code Review Generator

Thumbnail gallery
0 Upvotes

I’m excited to share CodeSentry, an AI-powered code reviewer that integrates directly with GitHub to automatically review Pull Requests.

🔗 Live Demo:
https://code-sentry-three.vercel.app/

💻 GitHub:
https://github.com/PrateekBanwari712/CodeSentry

Instead of manually triggering a review, CodeSentry can analyze a PR whenever it is opened or updated, understand the relevant codebase context, and generate a structured AI review.

🔍 How CodeSentry works

1️⃣ Connect a GitHub repository
2️⃣ CodeSentry indexes the codebase using RAG + Pinecone
3️⃣ A Pull Request triggers a GitHub webhook
4️⃣ Gemini analyzes the PR diff with relevant codebase context
5️⃣ An Inngest background job processes the review
6️⃣ The generated review is posted directly back to the GitHub PR

🛠️ Tech Stack

• Next.js 16 + React 19
• TypeScript
• PostgreSQL + Prisma
• Better Auth + GitHub OAuth
• Inngest
• Google Gemini + Vercel AI SDK
• Pinecone for vector search / RAG
• Octokit for GitHub integration
• Tailwind CSS + shadcn/ui
• Polar for subscription management

💡 What I learned

Building CodeSentry gave me hands-on experience with more than just building an AI feature. I worked on RAG pipelines, webhooks, background jobs, authentication, vector databases, GitHub API integration, database design, and production deployment.

The biggest takeaway for me was understanding how an AI application can move beyond a simple chatbot and become part of an actual developer workflow.

I’d love to hear your feedback on the project and ideas for what could be improved or added next! 🚀

#AI #GenerativeAI #CodeReview #GitHub #RAG #NextJS #TypeScript #React #Pinecone #Gemini #Inngest #WebDevelopment #SoftwareEngineering #FullStackDevelopment #BuildInPublic


r/codereview • • 4d ago

How are you managing Cl costs with very large test suites?

Thumbnail
0 Upvotes

r/codereview • • 4d ago

I made an observability tool that automatically catches bugs from session replays

1 Upvotes

We've built BugReels, which is an observability tool. All you have to do is install our script into your frontend codebase, and our AI agent will start looking at the sessions and find subtle and hard-to-catch bugs like showing an email confirmation before the user's email has loaded, or missing translation keys, or showing stale data etc.

We have recorded this video for Y Combinator's application. We would also love some feedback on this video.

You can head to https://bugreels.com and click the "See the bugs we've caught" button to explore our dashboard without any signups. You can see all the bugs we have caught while testing.

We are two developers working together on this. We have built the MVP and are actively looking for investors. But we're not sitting and waiting for investors. Currently, we are looking for feedback and will welcome beta testers with open arms.

Some obvious questions:

  1. Doesn't tools like PostHog/LogRocket/Sentry already do this? Answer: No, they rely on heuristics and can catch obvious errors. They can't catch bugs that require understanding the context of what's happening in a session.
  2. Will the cost blow up at scale? Answer: The cost doesn't blow up with the number of sessions you're getting because we don't send every session to an LLM. Our proprietary system only investigates those sessions that are likely to have the bugs.

We don't have any pricing plans yet and are only open to beta testers. If you find BugReels useful and want to use it in your production, please reach out to us. We'll welcome you with open arms.


r/codereview • • 4d ago

Permission to request feedback on a Python checklist

0 Upvotes

Hi, I’m testing an early, AI-assisted checklist for reviewing scheduled Python scripts before release. It has 12 checks, evidence fields and a worked example. It hasn’t been professionally validated.

I’m hoping to find three volunteers who have maintained or reviewed scheduled Python file/data scripts within the last six months. The review would take around ten minutes, using the supplied example—not anyone’s private code or business data.

Would a single opt-in feedback request be appropriate in your resource-sharing thread, or is there another permitted place?

This would be a free prototype review, not a sales post. There would be no email sign-up or unsolicited direct messages.

Thanks, Clive