Every AI tool can comment on a pull request now. That is precisely the problem.
A plausible remark is not free. Someone who knows the code has to read it, weigh it and decide whether it matters. When the remark turns out to be noise, the tool did not save review time. It spent it. In most software that is an annoyance. In systems code, where the person fielding the comment maintains a driver or a kernel tree and a bad merge costs weeks, it is a tax nobody agreed to pay.
So we built Agentic SQA around a different unit of output. Not the comment. The claim with its evidence.
What a finding actually is
When the agent reviews a change, it is not looking for things to say. It is building a case, and the case has a shape.
A finding opens with a verdict and its confidence. "This change likely introduces a bug" is a higher-confidence finding, worth a close look before merge. "This change might introduce a bug" is a lower-confidence finding where a second look is a good idea. The verdict also notes whether the issue would be reachable now or only if other conditions change, because a real defect behind an impossible precondition is a different conversation than one on the hot path.
Under the verdict sits the evidence. The suspected failure mode names how the code would break: memory errors, bounds violations, concurrency issues and the other defect classes that are easy to miss in review. Each claim is laid out next to what in the code prompted it, with links to the exact lines involved. What to check states the specific conditions that would need to be true for the bug to manifest. And the next steps split into what to do first and what deserves deeper inspection.
The point of the structure is that a maintainer can disagree with it. A vibe cannot be rejected on the merits. A claim with evidence can. That is what makes it worth reading, and it is also what makes it worth writing: a system forced to show its work cannot hide behind plausibility.
Silence is a feature
The second design decision was harder to hold: when nothing clears the bar, the product says nothing.
That runs against every instinct a tool has. Output feels like value. A review with no comments looks, on a dashboard somewhere, like a product that did not do anything. But the maintainer reading the pull request knows better. A quiet review on a clean change is the product working. Every finding we do not post is a promise that when we do post one, it is worth your attention.
Noise is a tax. We decided early that we would not charge it, even when saying something would have made us look busier.
From finding to fact
Investigation is probabilistic. The agent reads a change in the context of the codebase it lives in, forms a hypothesis about what could fail and backs it with evidence. Done honestly, that is as far as judgment goes, which is why every finding carries its reasoning instead of just its conclusion.
The layer we are building underneath it is not probabilistic. The verifier takes eligible findings and confirms them by reproducing the defect. Not every finding can be mechanically reproduced, and the ones that cannot still carry their evidence. But the ones that can stop being opinions entirely. A finding becomes a fact you can check.
That is the direction of the whole product. It keeps getting cheaper to generate code and cheaper to generate opinions about code. The scarce thing is knowing what is true before it merges. We think the tools that matter in systems code will be the ones that show their work.
Agentic SQA reviews pull requests in C codebases today. Connect a repository and reviews start on your next pull request.
