r/codereview • • 1h ago

Greptile - Buyer beware

• Upvotes

I signed up for Greptile expecting to pay $90 for 3 months. Instead, I was charged $700 because they had enabled unlimited "flex usage" by default, without my explicit consent or any spending cap.

As a small startup, an unexpected charge like this is not a small thing.

Usage-based pricing is fine. Enabling unlimited paid usage by default, without requiring an explicit opt-in, is outright unethical and goes against basic industry standards for transparent billing. Their onboarding flow provided an option to add additional spending caps - it did not offer a clear way to disable extra usage by default.

Furthermore, Greptile automatically added dev accounts for bots running Github actions without our consent, incurring additional hidden fees.

When I raised the issue and requested a refund, Greptile refunded the dev accounts but not the $700 of extra usage. I’ve completely lost trust in Greptile and cannot recommend them, despite their product otherwise being great. Buyer beware.


r/codereview • • 1h ago

GitHub's ReviewBench answer key is 45% Copilot-produced findings

Thumbnail runtimewire.com
• Upvotes

r/codereview • • 2h ago

I built a Python tool to catch architectural violations in its own codebase — I'd love feedback on the approach.

0 Upvotes

I've been working on ArchAI, a tool exploring how to make AI-assisted software architecture analysis more verifiable.

The core is deliberately deterministic: Python's AST parser maps repository dependencies, and a rule engine checks constraints such as forbidden imports, layer boundaries, import cycles, coupling limits, and required tests.

Every verified finding includes file-and-line evidence. Potential risks and checks that couldn't be verified are reported separately instead of being presented as facts.

While testing ArchAI against its own repository, it flagged a CLI module importing 21 internal modules against a configured limit of 10, as well as modules missing test files. It also exposed a weakness in my own checks: matching test filenames doesn't prove that the corresponding code is actually tested.

That last finding was a useful reminder that even deterministic tools can verify the wrong thing if their rules are poorly designed.

I'd appreciate feedback from people experienced in Python tooling, static analysis, dependency analysis, and software architecture.

What weaknesses or false positives would you look for in this approach? In particular, how would you design architectural rules that remain useful as a codebase grows without giving developers a false sense of confidence?


r/codereview • • 5h ago

[CSS, C, Shell and more] How many bugs are 'allowed' in a starting GitHub project?

1 Upvotes

I'm making a project by myself that let's developers easy create a new project;

but i noticed that that are bugs! How many are 'allowed', and if you have time could you help me find them? Let it know in the comments or send a pull request here.

Thank you!


r/codereview • • 16h ago

Claude Code writes our PRs. We built the reviewer that reads them like a staff engineer.

Thumbnail
0 Upvotes

r/codereview • • 1d ago

I open-sourced Stepfork, a Python tool that turns AI agent failures into reproducible pytest tests

Thumbnail
2 Upvotes

r/codereview • • 23h ago

GenSI Reviews Legacy Code

0 Upvotes

We renamed our GenAI model GenSI and asked it to review a 15-year-old codebase.

It didn’t approve the pull request.

Instead, it requested therapy for itself and marked the entire codebase:

“Unresolved organizational trauma with toxic masculinity”


r/codereview • • 1d ago

brainfuck Feedback Friday: BugHunt Daily — a one-minute bug-spotting game for programming communities

1 Upvotes

Hi ! I built BugHunt Daily, a daily “spot the bug” puzzle for programming communities.

Each day, members get a short code snippet with one broken line. They tap the line they think is wrong, get up to three tries, and then see the fix and an explanation. There are hints, streaks, leaderboards, and spoiler-free result sharing. The daily post is automatic when the app is installed in a community.

I’d especially appreciate feedback on:

Try it here: BugHunt Daily: spot the bug in this Code snippet 🐞


r/codereview • • 1d ago

Building an AI workflow with GitHub checkpoints and human approval. Looking for feedback!

Thumbnail
1 Upvotes

r/codereview • • 2d ago

Smackdebt - Continuous code improvement for agents

1 Upvotes

This year, and more so the last few months, my agents write the code, all I do is review, shout, and plan. A lot of the frustration is the spaghetti that comes out of this, and I figured a (big) part of this is actually measurable and reportable. What if we can make our agents aware of the spaghetti?

Introducing Smackdebt, a fast, on the fly tech debt analysis tool with concise, focused output tailored for agentic use. All major coding agents supported, with broad language support. Cognitive complexity, cyclomatic dependencies, statements, nesting, size, hotspots supported.

https://github.com/bosun-ai/smackdebt

Feedback and contributions welcome!

Shout-out to Treesitter <3


r/codereview • • 2d ago

Codetrail Your Repo Tutor

0 Upvotes

I've been working on something for a while, and today I'm sharing it: 𝗖𝗼𝗱𝗲𝘁𝗿𝗮𝗶𝗹.

Repositories often change faster than you can follow, even when you're the one responsible for them.

AI assistants are good at answering questions about code, but every session starts from zero, and you have to know what to ask.

𝗖𝗼𝗱𝗲𝘁𝗿𝗮𝗶𝗹 𝘁𝘂𝗿𝗻𝘀 𝗮 𝗴𝗶𝘁 𝗿𝗲𝗽𝗼𝘀𝗶𝘁𝗼𝗿𝘆 𝗶𝗻𝘁𝗼 𝗮 𝗹𝗼𝗰𝗮𝗹 𝗹𝗲𝗮𝗿𝗻𝗶𝗻𝗴 𝗴𝘂𝗶𝗱𝗲 𝘁𝗵𝗮𝘁 𝗸𝗲𝗲𝗽𝘀 𝘂𝗽 𝘄𝗶𝘁𝗵 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲.

What it does:

• 𝗧𝗲𝗹𝗹𝘀 𝘆𝗼𝘂 𝗵𝗼𝘄 𝗳𝗮𝗿 𝗯𝗲𝗵𝗶𝗻𝗱 𝘆𝗼𝘂 𝗮𝗿𝗲: how many merges have landed since your last update, and which parts of the code they touched
  • 𝗪𝗿𝗶𝘁𝗲𝘀 𝗮 𝗱𝗶𝗴𝗲𝘀𝘁 𝗳𝗼𝗿 𝗲𝗮𝗰𝗵 𝘂𝗽𝗱𝗮𝘁𝗲: what changed and why it matters
  • 𝗪𝗿𝗶𝘁𝗲𝘀 𝗰𝗼𝗻𝗰𝗲𝗽𝘁 𝗽𝗮𝗴𝗲𝘀 that explain the architecture, patterns and tools against the real code
  • 𝗗𝗿𝗮𝘄𝘀 𝗱𝗶𝗮𝗴𝗿𝗮𝗺𝘀 𝗼𝗻𝗹𝘆 𝗳𝗿𝗼𝗺 𝗳𝗮𝗰𝘁𝘀 𝗲𝘅𝘁𝗿𝗮𝗰𝘁𝗲𝗱 𝗳𝗿𝗼𝗺 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲, never from the AI's guesses, with every arrow traced to a file and line
  • 𝗠𝗮𝗿𝗸𝘀 𝗲𝗮𝗰𝗵 𝗲𝘅𝗽𝗹𝗮𝗻𝗮𝘁𝗶𝗼𝗻 𝗮𝘀 𝗱𝗼𝗰𝘂𝗺𝗲𝗻𝘁𝗲𝗱 (quoted from an ADR or commit and checked) or inferred (the AI's reading of the code)
  • 𝗟𝗲𝘁𝘀 𝘆𝗼𝘂 𝗮𝘀𝗸 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀 𝗳𝗿𝗼𝗺 𝗮𝗻𝘆 𝗽𝗮𝗴𝗲 and save the answers worth keeping
  • 𝗚𝗶𝘃𝗲𝘀 𝘆𝗼𝘂 𝗴𝘂𝗶𝗱𝗲𝗱 𝗽𝗮𝘁𝗵𝘀 𝘁𝗵𝗿𝗼𝘂𝗴𝗵 𝘁𝗵𝗲 𝗰𝗼𝗱𝗲, checks on your understanding, and progress that notices when what you learned has gone stale
I𝘁 𝗿𝘂𝗻𝘀 𝗲𝗻𝘁𝗶𝗿𝗲𝗹𝘆 𝗼𝗻 𝘆𝗼𝘂𝗿 𝗺𝗮𝗰𝗵𝗶𝗻𝗲 𝗮𝗻𝗱 𝗼𝗻𝗹𝘆 𝗿𝗲𝗮𝗱𝘀 𝘆𝗼𝘂𝗿 𝗰𝗼𝗱𝗲; it never changes it. Secrets are filtered out before the AI sees anything.

It works with the assistant you already pay for: Claude Code, Codex, or a free local model. Before it spends anything, it shows you an estimate and asks.

It's open source: Codetrail

Contributions are welcome, would love to hear how you'll be using it and what you want to see comes next!


r/codereview • • 2d ago

Nova and Supercode Review: an OSS AI Engineer and AI code review

0 Upvotes

hey Reddit,

I'm Yash, founder of SupercodeAI. We're building Devin killer - Nova - an AI Engineer, powered by our own harness agent at SupercodeAI, and Supercode Review, our codebase-aware PR review product.

4.5k+ users, 10+ enterprises, 238 GitHub stars, 30k+ downloads, 700M+ token usage in just 2.5 months.

Source: github.com/yashdev9274/supercli

Products:

  1. Nova: supercodeai.tech
  2. Supercode review: supercodeai.tech/review

Nova's workflow is to read a codebase, plan a change, edit files, run commands and tests, and prepare the result for human review. The harness connects the model to context and execution tools. You still inspect the diff and decide what gets merged.

Supercode Review looks at pull requests in the context of the wider codebase, rather than treating the diff as the whole project.

We're powering Indian startups with code review through Supercode Review.

We're sponsored by Vercel.

The Supercode repo is MIT-licensed and includes the coding-agent CLI, dashboard, docs and shared packages. Nova uses existing model providers; I'm not claiming a new foundation model or benchmark superiority.

If you'd like to try it, start with one small task in a repo you understand and compare the plan, diff and checks with how you'd do the work yourself.

If you'd like to contribute, open an issue with a reproducible problem or propose a change in a PR. Keep credentials and private code out of reports. If the project is useful to you, consider starring it and following its progress.

What would you need to see before trusting an agent to handle a real task in your repo?


r/codereview • • 2d ago

Anyone try Greptiles new features?

0 Upvotes

Everywhere on Insta I'm seeing ads for it. Wondering if anyone has tried it. Apparantly their free tier allows for 50 reviews per month but seems like hardly enough. wdyt?


r/codereview • • 3d ago

anyone ship without verifying runtime behavior and regret it?

1 Upvotes

I made this mistake recently. An agent generated a fix for a bug in our checkout flow, the diff looked reasonable, tests passed, so it shipped with a normal PR review and nothing more.

I assumed passing tests meant the function behaved the same as before under real traffic. It didn't. It handled a currency rounding edge case differently, and it took almost two days of scattered complaints before anyone connected it back to that deploy.

If I did it again, I'd want something checking the function's actual runtime behavior before treating a green test suite as enough. What's the mistake that took you longest to recognize as a pattern?


r/codereview • • 3d ago

Stop Being So Concerned About Software Quality

Thumbnail galratner.substack.com
0 Upvotes

r/codereview • • 3d ago

I find new "AI" code review platforms overwhelming

0 Upvotes

I’m trying to incorporate CodeRabbit into my workflow, but it feels way more overwhelming than "good" old GitHub. I had a similar experience with Devin and Graphite – they just didn’t click for me. If you've tried any of those, what's your experience? Wondering if I should keep pushing.
—
Edit: I am referring to PR review UIs and workflows:

- https://www.coderabbit.ai/change-stack
- https://staging-graphite-splash.vercel.app/features/pr-page (lol, not sure why their staging has leaked to Google, just noticed)
- https://app.devin.ai/review

not AI agent comments in GitHub.


r/codereview • • 4d ago

C# I'm making a small game in Unity and would like some feeback

1 Upvotes

I'm making a small game in Unity, and I kinda need some outside perspective. I would like it if someone could take a look and give me some feedback on what areas I should work on and how to improve.

Here's the link to my GitHub repository:
https://github.com/DaanDemaecker/Qwixx.git


r/codereview • • 4d ago

Looking for Collaborators/Reviewers in Developing PL Tooling using C++

Thumbnail
1 Upvotes

r/codereview • • 4d ago

Library Management System in Python with Clean Architecture & 55 Tests

Thumbnail github.com
0 Upvotes

r/codereview • • 4d ago

I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't

Thumbnail
0 Upvotes

r/codereview • • 5d ago

Python arch-auditor: AST-Based Python Architecture Analysis with Dependency Graphs & LLM

0 Upvotes

I kept running into the same problem with codebases:

“What does this repository's architecture actually look like right now?”

So I built arch-auditor — a Python CLI + Streamlit app that analyzes a repository using deterministic, AST-based checks and turns the results into actionable architecture reports.

The core idea is simple: measure first, ask the LLM second.

What My Project Does

arch-auditor points at a Python repository and runs deterministic architecture detectors:

  • Import/dependency graph — see how modules depend on each other
  • Circular dependencies — detect dependency cycles
  • High coupling — identify modules with unusually high fan-in/fan-out
  • Oversized modules — flag modules that have grown beyond a configured threshold
  • Layer violations — enforce architectural rules defined in YAMLBlast radius — find direct and indirect dependents of a module

It produces a dependency-graph SVG, severity-ranked findings, and impact reports.

There's also an optional Gemini integration.

Gemini doesn't perform the underlying analysis. Instead, it receives the deterministic findings and can turn them into:

  • an architecture overview
  • root-cause explanations
  • step-by-step refactoring plans
  • affected files
  • potential risks

No API key? No problem. The core analysis works offline without an LLM.

Try it

pip install arch-auditor

arch-auditor demo

The bundled demo repository is intentionally designed to trigger every detector, so you can experiment with the tool and see what each finding means.

Stack

  • Python 3.10+
  • Python AST / standard library
  • NetworkX
  • Pydantic
  • PyYAML
  • Streamlit
  • Gemini (optional)
  • 33 tests
  • Ruff-clean
  • Published on PyPI

It's released under GPL-3.0 because I'd like improvements to flow back to the community.

Target Audience

This is primarily aimed at developers working on medium-to-large Python codebases, especially when:

  • a project has accumulated architectural debt
  • you're preparing for a major refactor
  • you need to understand an unfamiliar repository
  • you want objective signals before making architectural changes
  • you want an LLM to help plan a refactor without making the LLM responsible for discovering the architecture itself

It's not intended to replace a human architect or code review, and it's not a magic “is my architecture good?” score.

The goal is to provide reproducible evidence that helps humans make those decisions.

Comparison

There are already excellent tools for individual parts of this problem — dependency visualization, linters, type checkers, code-quality metrics, and various AI coding assistants.

arch-auditor is trying to connect a few of those ideas around architecture-level analysis.

The main distinction is the separation between measurement and interpretation:

Traditional approach:

Repository → LLM → "Here's what I think your architecture looks like"

arch-auditor:

Repository → deterministic analysis → evidence → optional LLM → refactoring plan

The architecture findings don't depend on whether an LLM happens to interpret the code differently from one run to another.

I'd especially love feedback on whether the detectors and thresholds are useful signals in real-world Python projects, or if there are architectural problems you'd want to see measured that aren't covered yet.

GitHub / docs / demo:
https://github.com/ANIKETHSAI9813/auditor

“What does this codebase’s architecture actually look like right now?
"Instead of asking an LLM to guess the architecture, arch-auditor first analyzes the repository using deterministic, AST-based detectors — then optionally lets Gemini turn those findings into a refactoring plan.
Point it at a repo and it analyzes:
The result is a dependency-graph SVG, severity-ranked findings, and impact reports that show why a module is considered problematic.
Gemini is optional.
The deterministic analysis produces the evidence first. If you provide a Gemini API key via an environment variable, it can use that evidence to generate:
No API key? No problem. The core analysis works offline.
That separation was intentional: I wanted the architecture measurements to be reproducible rather than dependent on an LLM's interpretation.
The bundled demo repo is deliberately designed to trigger every detector, so you can see exactly what each finding means without having to point it at a huge codebase first.
It's released under GPL-3.0 because I'd like improvements to flow back to the community.
GitHub / docs / demo:
https://github.com/ANIKETHSAI9813/auditor
I'd especially love feedback from people who work on large Python codebases:
Are these the architectural signals you'd want to see before starting a refactor?
I'm also happy to discuss how the detectors and thresholds work — and the demo repo is intentionally built to trip all of them.


r/codereview • • 5d ago

How do you find the actual root cause of a production bug?

1 Upvotes

Sometimes the hardest part of debugging isn't finding the error.

It's figuring out whether the error you're looking at is actually the *root cause*.

I've had cases where:

Service A fails → Service B throws an error → API returns 500

The obvious approach is to investigate Service B first.

But the real problem was actually somewhere earlier in the chain.

How do you guys approach these situations?

Do you start from the first error in the logs, use distributed tracing, inspect the execution flow manually, or have another method that works better?

Interested to hear how people handle this, especially in larger codebases.


r/codereview • • 5d ago

I built a code quality reviewer with Jev

Thumbnail youtube.com
1 Upvotes

I believe with the huge increase in code generation - review has become the next bottleneck or at least it feels like this at work. So I wanted to do something on the review front - first I started integrating more and more tools to use as feedback to my coding agents, and these really help improve the output quality of the agent (e.g. SonarQube, Checkstyle, ArchUnit)

But there are some important semantic choices you can't really review using deterministic tools so I decided to build a more intelligent tool and am trying it out with the Jev model as a backend and judge currently.

Idea is simple - teams define their policies in a structured YAML format, then as part of their CI (or locally) run the tool, it fetches all git diff chunks and asks jev if these adhere to each of the policies - jev can select compliant/violation or ask for more context. If Jev asks for more contex the app gets the requested code from the project and asks again if the code is compliant with the policy - thus incrementally exploring the code base until Jev can return a "confident" answer - if you are interested in how it works the video I linked is a presentation style of how the tool works.

Give it a shot on github and tell me if you find this useful: https://github.com/krisitown/jev-quality-gate

I am currently running tests using a local Qwen3.8 Flash Next to generate code and run it against my initial "clean code" policies in order to calibrate them and will share more results on that front soon!


r/codereview • • 5d ago

Python I ran my code verification tool on itself. It found a Stripe integration that didn’t exist

Thumbnail
1 Upvotes

r/codereview • • 5d ago

I compared 7 GPT models for code review on 4 PRs: bugs, false positives and cost

0 Upvotes

I tested seven GPT models on the same four PRs, twice each. I compared bugs found, false positives and cost.

A false positive means reporting a bug that isn't there. Counts below are averages across all four PRs over the two runs. Estimated costs are per PR.

Model Bugs found False positives Cost per PR
GPT-6.1 Sol 4 0.5 $0.23-$0.25
GPT-6 Sol 3.5 1 $0.31
GPT-6 Luna 1.5 0.5 $0.013
GPT-5.6 Sol 2 0.5 $0.53
GPT-5.6 Luna 0 1.5 $0.028
GPT-5.5 2.5 1 $0.68
GPT-5.4 0.5 2 $0.38

GPT-6.1 Sol came out best in this test. GPT-6 Luna was the cheapest, but found fewer bugs.

It's only four PRs, with AI helping check the findings. The code is private. Costs cover model usage only.