Defence Before Fix: Preventing Bug Classes with Static Analysis
Defence Before Fix is a phase that runs before a defect is fixed. Rather than dropping straight into remediating the specific instance in front of you, you first treat that instance as evidence of a class, and you build the automated defence that detects every occurrence of that class across the whole codebase. The defence is only trusted once it has been seen to fire.
That is the definition from the specification, which is the canonical statement of the method and the place to go if you want the precise version. This article is where I first published the term, on 22 February 2026, and I have since rewritten it so that it agrees with the specification rather than with my earlier, looser description of the idea. Where the two still differ, the specification is correct, and I would suggest reading it as the source of truth and this as the long-form worked example that it deliberately leaves out.
The name is meant literally. The defence comes before the fix in time, because the moment you fix the bug the evidence you would have built the defence from is gone, and I have found that this is the part people most often skip whilst believing they have done it.
I didn't arrive at this from nowhere, and it is probably worth saying where I did arrive from. I have spent a long time, going back well before I had a name for any of it, believing that tooling rather than discipline is what actually makes quality stick in a codebase: php-qa-ci and the precursors that came before it are the practical result of that belief, a project's own checks wired into one entry point that nobody can forget to run because the pipeline runs them regardless of who is under deadline pressure that week. What has changed, and changed quite recently, is not the belief but the cost of acting on it. Writing a bespoke rule used to take long enough that only the most obviously recurring problems ever earned one, and everything smaller went into a code review comment and quietly reappeared a few months later under a different ticket number. An agent can now draft, prove and wire in a rule from a single reported bug in roughly the time it used to take to write that comment, so the calculation that used to favour fixing it and moving on has more or less flipped. Defence Before Fix is the name I have given to actually acting on that, every time, rather than only on the bugs that were annoying enough to justify the old cost.
The method, in six clauses
The specification states the method as six clauses, in order, and they are worth having in front of you before the worked example, because the example is really just these six applied to one bug.
- Attribute the defect to a class. The question is not "what went wrong here" but "what kind of thing is this an instance of", drawn within a lower bound (a rule that only catches the original instance is probably too narrow) and an upper bound (a rule that fires on code which does not carry the hazard is too broad).
- Build the net. Express the class as a rule in a detector, meaning a tool that reads code rather than executing it, in whatever your toolchain offers and bespoke where nothing off the shelf will take it.
- Prove the net by making the rule fire. Red before green, on the originating instance, and committed before the fix so that the red run survives in the history.
- Sweep the codebase, then fix every instance. Record the count before fixing anything, corroborate it by a search that does not depend on the rule, and then fix them all, with each fix addressing the hazard rather than merely satisfying the rule.
- Enforce permanently, and block. The rule becomes part of the project's quality checks, through the project's own entry point for accepting changes, and it fails rather than warns.
- Make the failure message terse, and point it at real documentation. The message carries a stable identifier that resolves to documentation shipped and versioned with the rule, stating what the rule is about, why it exists and how to fix a violation correctly.
Three of those clauses turn on judgements the specification deliberately does not close: whether code carries the hazard, whether a search was comprehensive, and how broadly to draw the class. Any threshold it gave would be calibrated to one codebase and one generation of tooling, so instead it asks the project to settle those judgements and record them somewhere the next person can find them. I think that is the right call, but it does mean the method asks more of you than a checklist would.
The clause I did not write down the first time
The original version of this article did not mention clause three at all, and it described the method as inverting the usual order, with the static analysis rule replacing the reproduction test. Both of those were wrong, or at least badly put, and the specification corrects them.
Defence Before Fix does not replace test-driven development and does not compete with it. You still reproduce the specific defect with a test, executed by a runner, and you still prove it fixed, exactly as normal. What the method adds is a phase before that work begins, operating a level above it: the test pins the instance, and the rule, evaluated by a detector, catches the class. A test proves that one input produces one wrong output. A rule finds the pattern wherever it occurs, including in code nobody thought to test, which is why a test must not serve as the detector.
The proof is the part that looks like a formality and is not. A new rule has to be seen to fire before it is trusted, and at minimum it has to fire on the defect that sent you looking. If your brand new rule comes back green, the honest reading is that the rule is broken rather than that the codebase is clean, and it is never the goal to write the rule and be instantly green. The specification goes further than I originally would have and requires that proof to survive as a commit of its own: the rule is committed with the originating instance still present, and the fix is committed afterwards, so that a reviewer can check out the first commit and watch it go red rather than take your word for it. Where the pattern is genuinely absent from the codebase, because it was already fixed or because you are defending pre-emptively, the rule is proven against fixture code instead, and that fixture is kept as the rule's own test.
What a bug hands you
A defect you have just found is a real, confirmed, impactful example of a harmful pattern. It is nothing like a pattern you read about in a style guide or one you suspect might cause trouble some day, because this one has already cost somebody something. That is precisely the raw material a good custom rule needs, and it is what speculative rules never have, because the hardest part of writing a rule is usually being sure that the thing it detects is genuinely worth detecting.
Every defect is therefore an opportunity to extend the codebase's permanent defensive coverage, and that opportunity exists only in the window before the fix. The specification is careful about what counts as a defect: a bug, a code review finding, a performance observation, an incident, an inconsistency. The method does not much care which, and it does not care whether the hazard is a failure either, since error hiding and code that nobody can safely change are hazards too.
The error-hiding patterns
The bugs that first made me want a name for this were all of one kind: the test suite is green, PHPStan reports nothing, CI passes, and somewhere in production a customer's payment has silently failed or their name has been replaced with a blank. It is not the sort of crash that sets off an alert, it is the failure that carries on running quite happily whilst producing wrong results. The cause was almost always code written to hide errors rather than handle them, and three patterns account for most of it in the PHP and TypeScript codebases I have worked in.
Pattern 1: the silent default
Null coalescing to a falsy value is the most pervasive form. It looks like defensive programming, and it is very nearly the opposite.
<?php
// Anti-pattern: converts a bug into empty data
$customerName = $order->getCustomer()->getName() ?? '';
$emailBody = "Dear {$customerName},\n\nYour order has shipped.";
// When getName() returns null because of a bug:
// "Dear ,\n\nYour order has shipped."
// The email sends, the test passes, and the customer is confused.
// Anti-pattern: the same problem in TypeScript
const customerName = order.customer?.name ?? '';
const emailBody = 'Dear ' + customerName + ', your order has shipped.';
// When name is undefined because of a data mapping bug:
// "Dear , your order has shipped."
// TypeScript is satisfied, the test passes, and the customer gets a broken email.
The distinction between "this value is legitimately empty" and "this value is missing because of a bug" has been erased. A renamed API field, a failed database lookup and a wrong property path all produce the same result, an empty string, and an empty string looks valid enough to pass any test that only checks "does this return a string".
Pattern 2: the empty catch
Exception handling exists so that errors propagate up the call stack until something can meaningfully deal with them. An empty catch block intercepts the error and discards it.
<?php
// Anti-pattern: the payment disappears silently
try {
$this->paymentGateway->charge($order->getAmount(), $card);
$order->markAsPaid();
} catch (\Exception $e) {
// TODO: handle this properly
}
// markAsPaid() never runs and the exception is gone.
// The user sees nothing unusual: no error, no retry, and no payment.
These usually start life as temporary scaffolding during rapid development, with every intention of adding proper handling later. Later rarely comes, because the code appears to work: no uncaught exceptions, no test failures, and nothing to draw anyone's attention back to it.
Pattern 3: implicit type coercion
Languages that perform implicit coercion absorb type mismatches instead of raising errors. PHP without strict_types will happily convert the integer 42 to the string "42" rather than flagging a type error at the function boundary where the two collide.
<?php
// Without strict_types, PHP silently coerces
function processOrderId(string $id): void
{
// $id becomes "42" even when called with the integer 42,
// so the type bug at the call site is invisible
}
processOrderId(42); // No error, no warning, silently wrong.
<?php
declare(strict_types=1);
// With strict_types, the bug surfaces immediately
processOrderId(42);
// Fatal error: Argument 1 must be of type string, int given
Strict typing turns every function signature into a validation checkpoint, so that a mismatch is caught where it occurs rather than three layers downstream when the wrong shape of data finally produces some unexpected behaviour.
Why green tests lie
The three compound. Silent defaults hide missing data at one layer, loose types let the wrong shape through the next, and an empty catch swallows the exception that would have revealed the problem at the third. The result is a system in which every test passes because every error has been converted into a valid-looking result, and a test that checks "the API returns a string" passes whether that string is the customer's name or the empty string a renamed field produced.
I would go as far as saying this is worse than having no tests, because untested code is at least obviously unverified, whereas code covered by tests that cannot see the error produces active false confidence. Diagnosing these bugs is also slow out of all proportion to their size, since the error and the symptom are separated by several silent conversions and tracing that chain backwards is closer to archaeology than to debugging.
The worked example
The specification says how a rule must behave and is deliberately silent on what one looks like, so this is the part this article exists to do. The example is a real shape of incident, tidied up, and it is the one the specification refers to when it mentions a reported defect that turned out to be one of twenty-three.
The incident
A support ticket arrives: customers are receiving emails that begin "Dear ," because the name is blank. You trace it to this code:
<?php
$name = $this->customerRepository->find($id)?->getFullName() ?? '';
$email->setBody("Dear {$name},\n\n{$body}");
The customer had been soft-deleted, so find() returned null, getFullName() was never reached, and ?? '' converted the null into an empty string. The email sent successfully as far as the application was concerned, no exception was thrown, and no test caught it. The fix is a one-liner and everything about the situation is telling you to make it and move on. This is the point where I would suggest resisting the urge, hopefully for reasons that will become clear.
Clause 1: attribute the defect to a class
This is not a missing null check, or not only that. It is an instance of coalescing an absent value into a falsy one, which is a pattern, and patterns can be detected mechanically. The class could be drawn wider still, as "any coalesce to a falsy default", and in the toolchain I maintain that wider class is in fact two rules, one for ?? '' and one for ?? false, because the fixes differ. For this example the class is the empty string. It is worth noticing that the class sits one level below the thing that was reported: the ticket was about blank names in emails, which no detector can read, and what the rule catches is the mechanism that produced them. The specification asks you to say, alongside the rule, whether the reported behaviour itself is also pinned by a check that actually runs the code, and here it is, by the test we come back to at the end.
Before writing the rule, search for other instances by means that do not depend on it. A text search for the token is one technique; reading the code paths that build customer-facing text is a second, independent one, because it would catch a spelling the search missed. The specification's stopping criterion is saturation rather than effort: use at least two independent techniques, and stop when the last one added found nothing the earlier ones had missed. Here the reading turned up nothing the text search had not, which is the signal that the search is done, and it also gave me a number to check the rule against later.
Clause 2: build the net
PHPStan custom rules implement its Rule interface and operate on AST nodes. Here is one that bans $value ?? '', with a stable identifier and a tip that points at the documentation:
<?php
declare(strict_types=1);
namespace App\QA\PHPStan;
use PhpParser\Node;
use PhpParser\Node\Expr\BinaryOp\Coalesce;
use PHPStan\Analyser\Scope;
use PHPStan\Rules\Rule;
use PHPStan\Rules\RuleErrorBuilder;
/**
* Bans null coalescing to the empty string: $value ?? ''
*
* The pattern hides bugs by converting missing data into empty data.
* Handle null explicitly so that the bug surfaces at its source.
*
* @implements Rule<Coalesce>
*/
final class NoNullCoalesceToEmptyStringRule implements Rule
{
// The identifier is stable, so the docs can be found from it
// long after the class has been renamed.
public const IDENTIFIER = 'app.nullCoalescingEmptyString';
public function getNodeType(): string
{
return Coalesce::class;
}
public function processNode(Node $node, Scope $scope): array
{
assert($node instanceof Coalesce);
if (
$node->right instanceof Node\Scalar\String_
&& $node->right->value === ''
) {
return [
RuleErrorBuilder::message("Null coalescing to empty string (?? '') hides missing data.")
->identifier(self::IDENTIFIER)
->tip('See docs/rules/null-coalescing-empty-string.md for the correct construction.')
->build(),
];
}
return [];
}
}
Register it in the PHPStan configuration:
# phpstan.neon
services:
-
class: App\QA\PHPStan\NoNullCoalesceToEmptyStringRule
tags:
- phpstan.rules.rule
Note the rule is drawn to the empty string specifically. A non-empty default such as ?? 'unknown' is a real decision that states what the absent case means, and a rule that flagged it would be firing on code that does not carry the hazard, which the upper bound forbids. In my experience a single report on innocent code is enough to make people stop trusting a rule, so I wouldn't tolerate any rate of false positives at all.
Clause 3: prove the net by making it fire
Run it. It must go red, and at minimum it must catch the line in the incident. Run through the project's own entry point rather than by invoking PHPStan directly, because the point is to prove that the project's checks will run the rule and not merely that the rule can fire:
bin/qa phpstan
[ERROR] Found 23 errors
src/Email/CustomerNotification.php:31
Null coalescing to empty string (?? '') hides missing data.
🪪 app.nullCoalescingEmptyString
src/Email/OrderNotification.php:47
Null coalescing to empty string (?? '') hides missing data.
🪪 app.nullCoalescingEmptyString
src/Report/CustomerSummary.php:83
Null coalescing to empty string (?? '') hides missing data.
🪪 app.nullCoalescingEmptyString
... 20 more instances
Twenty-three. The one you were sent to look at is a symptom, and the other twenty-two are the same bug sitting in twenty-two other places, waiting to surface in different contexts, reported by different customers, at different times, each arriving as its own support ticket months apart with nothing to connect them.
Now commit, and commit the rule on its own, with all twenty-three instances still present. That commit is your red run, and it is the only one anybody can go back and check. It is what a reviewer checks out to reproduce the proof, and it is the only record that the rule fired on real code rather than on a fixture. If the rule and the fixes share a commit, the red run can only be reconstructed by guessing at which lines to revert, and the proof rests on the guess.
Clause 4: sweep, then fix every instance
Proving and sweeping are two questions and not necessarily two runs. The same execution has already answered both: the rule works, and there are twenty-three of them. What matters is not to confuse the answers, since a large count does not make the rule more proven, and a rule that fired does not tell you the sweep is complete. The count is corroborated by the independent search from clause one, and if that search had found instances the rule missed, the rule was too narrow and would have to be widened until it caught them. The search is what you trust, and the rule has to earn its way up to it rather than the other way round.
Then fix them all, and this is where the method delivers most of its value. Each fix needs to deal with the actual hazard, and not simply do whatever makes the rule go quiet. For the incident line, the absent customer is an error and should be treated as one:
<?php
$customer = $this->customerRepository->find($id);
if (null === $customer) {
throw new CustomerNotFoundException($id);
}
$email->setBody("Dear {$customer->getFullName()},\n\n{$body}");
Not every instance wants the same answer. Some propagate the null explicitly so the caller decides; a few turn out to be places where absence is a legitimate state whose meaning the code can define, and there a non-empty default that states that meaning is the right fix. What is forbidden is a change that turns the rule green whilst leaving the failure mode intact, and suppressing the rule at the call site, which is not a fix at all. Supplying an empty default is usually the hazard in another form, because it makes the absent case look present and moves the failure downstream to where nobody expects it.
Twenty-three is a lot of fixing, of course, but it is a job to get through rather than a number to negotiate with. Baselining the existing instances so the rule only blocks new ones is, under the specification, a decision for whoever owns the codebase and not for the person doing the work, and a project that knows about twenty-three instances and fixes one has really just written down a list of defects it has chosen to keep.
Clause 5: enforce permanently, and block
The rule goes into the gate that blocks the build, for every contributor, permanently, and it fails rather than warns. A warning is a suggestion, and suggestions decay under deadline pressure, which is the condition under which the original defect was written. Where the checks run is the project's business; on my projects that gate is a single bin/qa that CI calls as a thin shim, so a CI failure reproduces locally without any ceremony.
Clause 6: the message, the identifier, and the documentation
The message is read under interruption by somebody trying to get on with something else, so it stays short: what was detected, where, and the identifier. The documentation carries the reasoning and the remedy, and it is read once, by somebody who has decided to understand the rule. The split is by job rather than by length, and anything that will grow over time belongs in the documentation. A message that names a pattern and leads nowhere is not much use to anyone, because "pattern X detected" on its own teaches you precisely nothing.
The identifier is the only string that reaches the reader, so it has to be stable across renames and it has to resolve on its own, from a log or a ticket, without the message around it. The documentation lives in the same repository as the rule and is committed with it, so that the two cannot drift apart. The best custom rules I have written are opinionated documentation encoded as automation, and this clause is what makes that literally true.
Only then, the fix you came for
Now write the failing test for the original bug: a soft-deleted customer's order should throw when the email is prepared, not send a blank-named message. Watch it fail, make it pass and commit it, which is exactly the ordinary TDD loop you would have done anyway, only now it comes after the defence rather than instead of it. After this process you have one rule that prevents the class permanently, one test that documents the correct behaviour at the instance, and twenty-three latent bugs fixed rather than one.
The same rules in other tools
The equivalent ESLint rule for a TypeScript codebase operates on the same idea and a different AST:
/** @type {import('eslint').Rule.RuleModule} */
module.exports = {
meta: {
type: 'problem',
docs: {
description: 'Disallow null coalescing to empty string (?? "")',
url: 'https://example.com/docs/rules/no-empty-string-fallback',
},
messages: {
noEmptyStringFallback:
'Null coalescing to empty string hides missing data. '
+ 'See no-empty-string-fallback for the correct construction.',
},
schema: [],
},
create(context) {
return {
LogicalExpression(node) {
if (
node.operator === '??' &&
node.right.type === 'Literal' &&
node.right.value === ''
) {
context.report({ node, messageId: 'noEmptyStringFallback' });
}
},
};
},
};
And a PHPStan rule targeting Catch_ nodes catches empty exception handlers before they ship:
<?php
declare(strict_types=1);
namespace App\QA\PHPStan;
use PhpParser\Node;
use PhpParser\Node\Stmt\Catch_;
use PHPStan\Analyser\Scope;
use PHPStan\Rules\Rule;
use PHPStan\Rules\RuleErrorBuilder;
/** @implements Rule<Catch_> */
final class NoEmptyCatchRule implements Rule
{
public const IDENTIFIER = 'app.emptyCatchBlock';
public function getNodeType(): string
{
return Catch_::class;
}
public function processNode(Node $node, Scope $scope): array
{
assert($node instanceof Catch_);
if (count($node->stmts) === 0) {
return [
RuleErrorBuilder::message('Empty catch block silently swallows exceptions.')
->identifier(self::IDENTIFIER)
->build(),
];
}
return [];
}
}
For TypeScript, ESLint's built-in no-empty rule covers that one at error level with allowEmptyCatch off. Where an off-the-shelf rule or a tightening of existing configuration genuinely detects the class, using it conforms, and I would always reach for that first. The rules that matter most, though, are the ones tightly coupled to your project and carrying knowledge specific to it, which is why the specification requires a conforming toolchain to allow bespoke custom rules at all.
It is also worth saying that the defaults of the excellent tooling that exists for PHP and TypeScript are too permissive to do much of this on their own. declare(strict_types=1) in every file and PHPStan at level max with the strict rules extension is the highest-value change I know of in a PHP codebase; in TypeScript, strict is table stakes and noUncheckedIndexedAccess and exactOptionalPropertyTypes catch classes of runtime error that base strict mode misses. Those steps alone will usually surface a backlog of latent bugs in tested, passing code, and they are a good way to find out whether the method is for you before you write a rule.
Rules that reason across the whole codebase
Most rules examine a single file in isolation, but some of the most valuable custom rules cross file boundaries, checking whether code is properly connected to the rest of the system rather than only whether it is internally correct. The sharpest illustration I have of why that matters is a service that was entirely correct, thoroughly tested, and never called.
On a production Symfony project processing supplier product data, a preprocessing service existed with working SQL logic and a green test suite, but it had never been injected as a constructor dependency into the pipeline meant to call it. Autowiring does not wire in a service nobody asks for, so it lived in isolation. During a scheduled Christmas shutdown in which stock quantities were zeroed, the prices that service should have cleared stayed set. The code was correct, the tests verified the code, and the pipeline never ran it.
A custom PHPStan rule now catches that entire class. At analysis time it reads the production source tree and checks whether a class that the tests exercise is referenced as a dependency anywhere in production code:
<?php
declare(strict_types=1);
namespace App\QA\PHPStan;
use PhpParser\Node;
use PHPStan\Analyser\Scope;
use PHPStan\Rules\RuleErrorBuilder;
// Simplified from a production PHPStan rule.
// Detects service classes used in tests but never in production code.
final class ServiceOnlyUsedInTestsRule extends AbstractClassRule
{
public function processNode(Node $node, Scope $scope): array
{
$className = $node->getClassReflection()->getName();
$shortName = $this->getShortClassName($className);
// Read every production source file and look for the class
// being referenced as a dependency anywhere outside itself.
$usedInProduction = false;
foreach ($this->productionSourceFiles() as $file) {
if ($file->getPathname() === $scope->getFile()) {
continue;
}
if (str_contains($file->getContents(), $shortName)) {
$usedInProduction = true;
break;
}
}
if (!$usedInProduction && $this->isUsedInTests($className)) {
return [
RuleErrorBuilder::message(sprintf(
'Service %s is only used in tests, never in production code.',
$shortName
))
->identifier('app.serviceOnlyUsedInTests')
->build(),
];
}
return [];
}
}
No test can reach this bug, because the tests exercise the service directly and correctly. Only something that reasons about the full production dependency graph can notice that the service is never invoked when the application actually runs. Codebases that accumulate rules of this kind tend to develop clusters of them around a single pattern: for a domain-specific database access layer I have ended up with a rule that prevents query objects being created inside loops, a companion that catches prepared statements created inside loops, a third that detects a prepared statement used only once in a method (it should be a simpler query class), and a fourth requiring a performance-monitoring dependency in every prepared statement class. Each catches a different failure mode of the same pattern, and together they make misuse structurally difficult.
The same approach works in TypeScript. An ESLint rule can read route definitions from a separate file at lint time and validate every internal link against them, so that renaming a route without updating its references fails the build without any test covering that navigation:
const fs = require('fs');
const path = require('path');
// Load the valid routes from the route registry at lint time
const routeSource = fs.readFileSync(path.resolve('src/routes.ts'), 'utf8');
const validRoutes = parseRoutesFromSource(routeSource);
module.exports = {
create(context) {
return {
Property(node) {
if (isLinkProp(node) && node.value.type === 'Literal') {
const href = node.value.value;
if (typeof href === 'string' && href.startsWith('/') && !validRoutes.has(href)) {
context.report({
node,
message: 'Link "' + href + '" points to a route that does not exist.',
});
}
}
},
};
},
};
Rules that read the wider codebase are more expensive to write and slower to run than single-file rules, so I would save them for failure modes that are severe and that tests genuinely cannot reach: services disconnected from pipelines, broken internal navigation, documentation that has drifted from the pages it describes. Those are the bugs that slip through green suites because tests model code in isolation rather than how the system is assembled.
What exists now
When I first published this article the method was a description and a name. It is now two specifications, both at defence-before-fix.github.io under Creative Commons Attribution 4.0: the method specification, version 1.0.0, which states what a practitioner does when a defect is found, and the toolchain specification, version 0.1.0, which is addressed to anyone who maintains a linter or a QA pipeline that other people install and says what a tool has to offer so that the projects using it can follow the method at all. The two quality toolchains I maintain, php-qa-ci and ts-qa-ci, each record in their own manifest the specification versions they conform to, can list their active defences and resolve a printed identifier to its documentation, and name the method with a link to the specification whenever the pipeline fails.
#!/usr/bin/env bash
# php-qa-ci: resolve a printed identifier to its documentation
bin/rule-doc app.nullCoalescingEmptyString
# ts-qa-ci: list every active defence, derived from the resolved ESLint config
npx ts-qa rules
I want to be careful about what I am and am not claiming, because the territory next to this is well populated. Defensive programming is decades old, preventing classes of bug rather than instances predates this by a long way, and static analysis, custom lint rules and blocking quality gates are all long-established practice, so I claim none of them. What I am claiming is the term, which I couldn't find in use as a named practice when I published it, the placement of the work before the fix, and the requirement that a rule be proven by firing before it is trusted. It is a method I named and published, and that is the whole of the claim; I am not asserting that anyone else has adopted it.
Why this matters more under AI-assisted development
For a human developer a failure message that teaches is good practice. They may read it, may internalise it, may ignore it, and which of those happens depends on their seniority, their workload and how many times they have seen the message before.
For a coding agent, as far as I can tell from watching a lot of them work, the failure message is more or less the whole of the correction loop. It gets consumed as instruction, in the same turn, every time, without fatigue and without any seniority gradient, and a message that resolves to documentation explaining the correct approach does not simply block the agent but tends to turn it around and point it the right way, and in my experience it does that about as well on the hundredth occurrence as on the first. That reframes the usual complaint about AI-written code somewhat. The difficulty was never really that agents make mistakes, people do too, but rather that nobody had built a channel for correcting them at the level of the class instead of one instance at a time.
Most of my working time for the last couple of years has been spent directing agents rather than hand-writing code, and the specification's section on operating a defence under AI-assisted development is the part I would most stand behind, because every clause in it was learned the hard way. The practitioner has to be able to run the defence themselves, or the loop never closes in the turn where the mistake was cheap to fix. The result has to arrive in the output they are already reading. The identifier has to resolve without a human, which for an agent means mechanically. The documentation has to state the correct construction and not only the prohibition, because a prohibition alone leaves the agent to guess at the replacement, and it will guess. And the project's defences have to be enumerable, because an agent arriving at a codebase has no colleague to ask and no memory of last time, so without a listing its standards can only be learned by violating them one at a time.
The cost of fixing has changed as well. The historical case for baselining a large sweep was the cost of human hours, and that cost has largely collapsed, since an agent can work through hundreds of instances at a price that would have made a baseline unavoidable a few years ago, so reaching for one now is usually a habit rather than a judgement, and the specification stops short of forbidding baselines outright, but only just. The other half of that coin, and the half that is easier to forget, is who actually gets to decide. An agent executing this method decides how the defence is built, and it does not decide what the codebase is permitted to keep; suppressing an instance, baselining, or leaving a known instance unfixed belong to whoever owns the codebase, whatever the count turns out to be, and an agent that reaches for an exception has almost always found a shortcut rather than a genuine obstacle. I saw both failure modes, the agent that stalls at the first judgement call and the agent that quietly takes the decision itself, when I had early drafts of the specification read cold, and the second is much harder to notice than the first.
Hopefully that makes the case on its own. If you intend to have agents writing code in your codebase, then the rules in your quality gate, and the messages attached to them, are the main channel you have for teaching them your project's standards at all.
The ratchet
The goal is not zero bugs, which I don't think is achievable for anyone. The goal is that every bug leaves the system better defended than it found it, so that each incident leaves behind a defence as well as a fix and a test, and the categories of bug that can survive in the codebase shrink over time. A codebase with a mature set of custom rules has a different character from one without: code review spends its attention on logic and architecture rather than on patterns the linter could find, new contributors are held to the established safe patterns from their first commit, and the mistakes of the past become structurally impossible to repeat rather than merely discouraged.
So the next time a bug reaches production, before you write the test, ask what pattern allowed it and whether a machine could be made to recognise that pattern everywhere. Do not decide in advance whether it can; attempt it, because failing to write a rule within the bounds is itself the evidence that the defect is out of scope, and it is cheaper and more reliable than a judgement made before trying. Where it can, write the rule first, prove it by making it fire, commit that proof, sweep, fix every instance, enforce it, and document it. Then, and only then, fix the bug in the ordinary way. The specification has the precise version of all of that, and if you maintain tooling that other people install, the toolchain specification alongside it is the one addressed to you.
This is what the QA-gates tooling automates.
See php-qa-ci and ts-qa-ci