AI code review has found its real job

The most interesting AI story this week is not a new chatbot, a shinier benchmark or another demo in which an agent books a restaurant table without setting fire to the calendar.

It is Linux having too many bug fixes.

Linux 7.2 reached its seventh release candidate with an unusually large flow of late fixes. Linus Torvalds described this as the possible “new normal”, with many fixes resulting from reviews by AI tools. That sounds like a straightforward success story. More bugs found, more bugs fixed, safer software.

It is not quite that tidy.

The Linux kernel is one of the largest, most scrutinised and most consequential open-source projects in the world. Its development process is deliberately demanding. Patches travel through public mailing lists, specialist reviewers and subsystem maintainers before they reach the mainline kernel. If AI-assisted review is altering that process, we are looking at something far more useful than a model producing a plausible block of code on command.

We are seeing AI become part of quality control.

The less glamorous job may be the better one

Most discussion about AI and software focuses on generation. Can the model build the feature? Can it write the application? Can somebody with no coding experience make a subscription business before lunch?

Code review is a different proposition. Instead of asking a model to invent a solution from a blank page, you give it a proposed change and ask it to look for omissions, unsafe assumptions, resource leaks, race conditions and interactions the author may not have considered.

This plays to a genuine strength of current models: examining a large body of material from several angles without becoming bored on the fourth pass.

Sashiko, an agentic review system operated under the Linux Foundation and funded by Google, monitors kernel mailing lists and runs proposed changes through specialised review stages. Its published evaluation says it detected 53 per cent of bugs in an unfiltered set of 1,000 recent upstream issues identified using Fixes: tags. Those were bugs that had already passed through human review before later being corrected.

That figure is promising, but it needs careful handling. It is a project-reported evaluation, not proof that Sashiko catches 53 per cent of every real-world kernel bug. False positives are also harder to measure. The useful fact is simpler: the system is finding real problems in code that experienced people have already examined.

That makes AI review a second pair of eyes, except this pair can look repeatedly, use different specialist prompts and inspect a volume of patches no individual could reasonably cover.

Finding a problem is not the same as helping

Here is the awkward part. Every warning creates work for somebody.

A maintainer must decide whether the report is valid, understand its implications, check for duplicates, reproduce the issue where necessary, judge the proposed fix and ensure the correction does not break something else. A tool can generate twenty suspicious findings in minutes. The human cost of dismissing nineteen weak ones may erase the value of the useful twentieth.

Linux developers have already experienced that pressure. AI-assisted security reports have increased the volume of duplicate, trivial and poorly timed submissions. In community discussion, practitioners repeatedly draw a useful distinction between AI-generated patches and AI-assisted review. The latter is generally seen as more promising, but only when a responsible human owns the output.

This is the hidden constraint in agentic work. The bottleneck moves.

If an agent makes research ten times faster, verification becomes the bottleneck. If it produces code ten times faster, review and testing become the bottleneck. If it finds possible vulnerabilities ten times faster, triage becomes the bottleneck.

The agent has not removed the workflow. It has changed where the expensive thinking happens.

What businesses should learn from Linux

You do not need to maintain an operating system to use this lesson.

Many businesses are introducing AI at the point of production: write the email, build the page, generate the campaign, create the report. That can save time, but production is also where confident mistakes escape into customer-facing work.

Review is often a safer and more valuable place to begin.

An AI reviewer can check a website brief against the delivered build, compare a proposal with the client’s requirements, inspect a contract for missing fields, test whether a report’s claims match its sources, look for accessibility problems or flag inconsistencies across a set of product pages.

The important design choice is that it should not silently approve its own work. Generation and review need different passes, different instructions and, for meaningful decisions, human judgement.

A practical review workflow looks like this:

  1. Give the reviewer a bounded object, such as a document, patch, page or spreadsheet.

  2. Define what counts as a material issue, not merely anything that could be changed.

  3. Require evidence for every finding, including the exact location and reason.

  4. Separate high-confidence defects from suggestions and stylistic preferences.

  5. Deduplicate findings before they reach a person.

  6. Keep a human owner accountable for acceptance, rejection and action.

  7. Measure useful findings against the time spent checking bad ones.

The final point matters. Counting everything the agent flags rewards noise. The better metric is useful defects found per hour of human review.

The real product is the filter

AI companies naturally sell capability: more reasoning, more autonomy, longer tasks and broader tool access. In practice, the difference between a clever demo and a dependable system is often the filter around it.

Who is allowed to submit work? What evidence must be attached? Which findings reach a human? What happens to duplicates? Who can make a change? What gets logged? How do you learn from false alarms?

Linux is confronting these questions in public because its process is public. Most businesses will encounter the same questions quietly, inside inboxes, project boards and increasingly crowded approval queues.

The lesson is not that AI review is unreliable. Nor is it that human review is obsolete. It is that useful automation can create more work before it creates less, particularly when the system detects possibilities faster than people can resolve them.

AI code review appears to have found a genuinely valuable job. Now the harder work begins: designing a human system capable of using what it finds.

Key takeaway

AI creates the most value when it catches material problems before they escape, but every automated finding consumes human attention. Build the triage, evidence and ownership system at the same time as the agent.

Sources

Linux 7.2-rc7 coverage and Torvalds’ release comments (https://www.techradar.com/pro/linus-torvalds-says-huge-linux-kernel-updates-are-now-the-status-quo-and-its-all-thanks-to-ai)

Sashiko official project (https://sashiko.dev/)

LWN: The Sashiko patch-review system (https://lwn.net/Articles/1063292/)

LWN: Reviewing kernel patches with LLMs (https://lwn.net/Articles/1073583/)

Linux kernel mailing-list patch documenting AI-assisted review tooling (https://www.mail-archive.com/linux-kernel%40vger.kernel.org/msg2622390.html)

Discuss your website or AI workflow with LIFT Brandworks.

Lift BrandWorks

Hi, I’m Rob, a Squarespace Gold Circle Partner and designer with a focus on fast, high-converting websites for growing businesses. I specialise in clean design, persuasive copy, and smart user flows that turn traffic into trade.

I’ve worked across hospitality, wellness, coaching, and creative sectors, delivering custom builds without templates, every site is tailored, mobile-optimised, SEO-ready, and built for growth. I’m also a trained coach and great communicator, making the whole process easy and collaborative from start to finish.

Need help launching or refreshing your site? I offer full design, development, and optional chatbot integration to help your site work harder for you.

https://www.liftbrandworks.com
Previous
Previous

AI is becoming a utilities business