We’ve Run This Experiment Before

Most evenings lately I write code with AI agents inside RISE, the spectral renderer I’ve written about here before. The agents write most of the lines now. What they also do, constantly, is propose workarounds: plausible, tidy, well-commented shortcuts that would each mortgage some future evening of mine. I almost never take them. The rule in my codebase is that an agent either gets it right or documents the limitation and turns it into an explicit cleanup step, and that rule is now codified into the skills and prompts I use, to the point where about four times out of five the agents enforce it without me. Last week, working through a robust spectral treatment of how light interacts with grand feu enamel, the discipline earned its keep several times over.

I can hold this line for two reasons. The first is that I wrote every line of the original RISE, so I can steer agents at the level of a principal engineer; I know where the bodies are buried because I buried them. The second is blunter: every bill in that codebase is addressed to me. There is no future teammate to inherit a shortcut, no reorg to discharge the debt. When the only person who can be mortgaged is you, you read the fine print.

Neither advantage survives contact with somebody else’s code. At work, if I drop into an unfamiliar corner of a large codebase (some part of the Android framework I’ve never touched, say), both evaporate at once. I can’t steer because I don’t know the terrain, and I can’t price a proposed shortcut because I don’t know what anything costs there. No prompt library fixes that; on a shared system the relevant context lives in hundreds of heads and a dependency graph that no one person holds. I am as much at the mercy of plausible-looking output as an engineer in their first week.

I’ve come to think my home setup is about as healthy as AI-assisted development gets, and that the reasons have almost nothing to do with the AI.

The amplifier is old news

The 2025 DORA report put its central finding right in the announcement: “AI doesn’t fix a team; it amplifies what’s already there.” Strong teams get stronger, struggling teams struggle faster. This is now close to conventional wisdom, and I’ve argued adjacent versions of it myself: that AI won’t fix your culture issues, and that when you accelerate the parallelizable part of engineering work, Amdahl’s law hands the whole game to the serial fraction that remains. Code generation is the parallel fraction. Review, integration, architectural judgment, maintenance: serial.

But “amplifier” explains less than it appears to. The amplifier is roughly the same everywhere; the outcomes are not. Some amplified systems correct themselves within quarters. Others spiral for a decade while everyone writes think pieces about them. If amplification is uniform and outcomes diverge, the interesting variable is somewhere else.

Conveniently, we have already run this experiment once, at civilizational scale.

What the last amplifier taught us

Social media collapsed the cost of producing content whose entire value depends on scarce downstream attention. Before the collapse, effort was doing quality-control work: publishing anything required enough investment that what got published mostly came from people with some stake in being right, or at least in being read twice. Remove that friction and you don’t amplify everything equally. Careful work is bottlenecked by judgment, which the tool doesn’t supply; careless work was bottlenecked only by effort, which the tool eliminates. The amplification skews, structurally, toward whatever was previously rate-limited by effort alone.

Code has exactly this shape. A post is worthless until it’s read; code is worthless until it’s reviewed, integrated, and maintained. And the numbers coming out of the current transition look less like a productivity story than like a feed filling up. Telemetry from Faros across 22,000 developers shows median PR review time up 441%, 31% more PRs merging with no review at all, and incidents per PR up 242%. GitClear‘s analysis of 623 million code changes finds duplication up 81% since 2023 while refactoring has collapsed to under 4% of changed lines, and the share of work that touches code more than a year old has fallen 74%. The codebase grows outward while the older strata calcify. These are vendor numbers with vendor-shaped selection effects; I lean on them anyway because independent datasets keep converging on the same shape. Production outran judgment, which is what posting did to moderation.

The analogy has real limits, and they’re worth naming before leaning on it. Social media was also algorithmic distribution, ad-market incentives, and identity games, none of which have clean counterparts in a pull request; code doesn’t go viral, and maintainers hold gates that moderators never did. What transfers is the load-bearing part: cheap production of artifacts whose value depends on scarce downstream judgment, and the question of who absorbs the difference.

None of this means the amplifier is bad, and these are shock-phase numbers besides, drawn from the first few years after the friction dropped; part of what follows is an argument that some of them will recover. What the numbers do establish is that the amplifier moved the bottleneck. The serial fraction didn’t shrink; it just acquired a much longer line outside its door. Which raises the only question that decides how this ends, and which of those curves bend back: whose budget does the serial fraction come out of?

Follow the invoice

Social media never answered that question, and that failure is most of its story. The gains (engagement, ad revenue) were booked by the platforms. The costs (attention, civility, adolescent wellbeing; pick your favorites from the literature, there are plenty) were diffused across society with no line item anywhere. Nothing on any platform’s income statement measured the damage, so no feedback loop existed to correct it. The only forces that ever changed platform behavior were advertiser boycotts and regulators: the rare occasions when a cost found its way onto the P&L. The system didn’t fail to self-correct for fifteen years out of negligence. Structurally, it couldn’t.

Enterprise software runs the same amplifier with one enormous difference: the bill usually arrives at the right address. Incidents page the organization that shipped them. Review load lands on the team that merged. Comprehension debt slows the same roadmap that booked the throughput win. The feedback arrives in weeks and months, not decades, addressed to the people holding the amplifier. This is why I expect enterprise engineering to climb out of its current instability dip in a way social media never climbed out of anything. You can already see the loop starting to close: the same reports documenting the instability are converging on the countermeasures (small batches, review-side agents, merge gates, provenance tracking), and organizations are adopting them for the oldest reason there is. It hurts.

The climb won’t be uniform, though. Amplifiers widen variance between organizations the same way they widen it between developers. Loosely coupled architectures with fast feedback will close the loop quickly and bank the gains; calcified monoliths will respond the way large bureaucracies usually respond to pain, with process scar tissue that trades the new throughput away for the old stability. Both are the loop working. One of them just works ugly, and if you’ve sat through the meetings you know which one.

But enterprises leak, and the leaks are where the trouble concentrates. The organization’s loop can close while the individual’s stays open, and in a large enough org the insulation can hold for years: the author’s team shielded from consequences by separate on-call rotations, SRE teams that absorb the pages, middle-management metrics that count velocity but not repair. The debt gets collectivized onto the engineering balance sheet while the individual invoices go unsent. Performance systems that reward authorship volume make it worse; the engineer collects on the throughput while the debt matures on somebody else’s watch, and median tenure is shorter than the half-life of the debt being created. A reorg is many things, and one of them is a bankruptcy proceeding for technical debt in which the creditors are not invited. Even the loop between a developer and their own experience turns out to be open: in METR’s 2025 randomized trial, experienced open-source developers using AI tools were 19% slower on their own repositories while believing they’d been 20% faster. Sixteen developers, early-2025 tools; the specific number will not survive, but the direction of the self-assessment error is the durable finding. And past the enterprise boundary sits contract and agency development, the ship-and-leave end of commercial software, where the loop is nearly fully open and, I’d predict, quality is currently decaying fastest with the fewest people measuring it.

So the honest version of the claim isn’t a binary; it’s a dial. The speed at which an amplified system corrects itself is proportional to how tightly its costs loop back onto whoever holds the amplifier. My home setup sits at the tight end: one developer, one payer of debts, one architect, the invoice and the decision belonging to the same person. The same loop does double duty. Twenty years of paying my own bills in that codebase is what lets me price the debt, and it’s also what built the judgment that lets me steer the agents at all. The loop that routes consequences is the loop that manufactures taste. Enterprises sit in the middle of the dial, leaks and all.

And at the far end, the loop is open all the way around.

The commons

In January 2026, inside three weeks: curl announced the end of its bug bounty program, which since 2019 had paid over $100,000 for 87 confirmed vulnerabilities. In earlier years better than 15% of submissions had turned out to be real; starting in 2025 the rate fell below 5%, with roughly a fifth of submissions being what the ecosystem now calls slop. Daniel Stenberg’s conclusion was pure incentive economics: the bounty itself had become too strong an inducement to fabricate problems. He also noted, tellingly, that pull requests had never been a problem for curl, because two hundred CI jobs filter those before a human ever looks. The slop pooled exactly where verification still requires human judgment. The same month, Ghostty’s Mitchell Hashimoto moved from AI-disclosure requirements to a zero-tolerance policy; his commit message said the quiet part, that agentic programming had eliminated the natural friction of effort that used to filter contributions, with bad PRs up roughly tenfold by his estimate. Weeks later Ghostty added a vouching system requiring first-time contributors to introduce themselves “in your own words, not written by AI.” tldraw began auto-closing all external pull requests. NetBSD now requires core-team approval for AI-generated code; QEMU declines it outright.

The immune response has a familiar shape. Disclosure requirements are content labels; vouching is identity verification. Auto-closed PRs are closed registration, and permanent bans are deplatforming. Killing the bounty demonetizes the behavior that attracted the spam, while the muttering about decamping to Codeberg is the Mastodon exodus. Nobody coordinated any of this. A few thousand maintainers independently reinvented the content-moderation toolkit in about eighteen months, which is what happens when the same disease meets the same kind of host.

The commons moved faster than most enterprises can change a perf rubric, which can look like the open loop closing quickest of all. It isn’t, because defense is not correction. Maintainers can’t send invoices; exclusion is the only tool available to someone who bears the cost but can’t reprice it. curl didn’t close its loop, it amputated the surface the slop arrived on, and paid for the speed with the thing that made the commons a commons. The velocity is itself the mechanism showing through: people who bear the costs act decisively when they also hold the authority to set the rules, which is Ostrom’s central observation. What they cannot do, from inside a commons, is make the polluters pay.

The epilogue drives it home. A few months after the shutdown, Stenberg reported that the slop reports had stopped entirely, replaced by a rising stream of genuinely good, AI-assisted security reports arriving at a frequency he’d never seen and putting the team under comparable load. The quality improved; the workload didn’t. Generation stays cheap relative to review no matter what is being generated, which is why better models don’t dissolve the problem.

That’s because the host is a commons. Baltes and colleagues make this argument carefully in a recent position paper, “AI Slop and the Software Commons”, built on a companion empirical study of developer discourse: generating slop is cheap, reviewing it is not, and the costs of AI-accelerated contribution are externalized onto reviewers, maintainers, shared knowledge resources, and the talent pipeline. Hardin’s pasture, with pull requests. Their Ostrom-inspired prescriptions (communities free to set their own rules, gating, norms of accountable use) are the right medicine and worth reading in full. The commons isn’t uniform, of course; Kubernetes, with foundation backing and salaried corporate maintainers, has very different loop geometry from a solo-maintained library, which is part of why the defenses range from gentle disclosure rules to full closure.

What I’d add to their diagnosis is the comparative point. The commons is the one region of software where the loop is open all the way around, and it is therefore the one region faithfully reproducing the social media trajectory, platform incentives included. GitHub doesn’t run ads, but the structure rhymes: it sells the amplifier by the seat, and its value inside Microsoft is narrated in growth (36 million new developers arriving in 2025) while the review costs land on maintainers it doesn’t pay. Its own 2026 report described the flood as “a denial-of-service attack on human attention,” in the same breath as noting that maintainer capacity was not keeping pace. The platform books the volume; the commons reviews it. Meanwhile the ecosystem is bolting on, after the fact, the funding layer it always free-rode on: an Open Source Endowment launched this year, gathering north of $700,000 from its founding donors, backed by the people behind curl, Vue, and HashiCorp. Society funding the trust-and-safety function a decade late is, at this point, a tradition.

What the analogy predicts

If the loop is the variable, the predictions come cheap, and they’re checkable. Enterprise instability metrics recover quickly, on the order of months once an organization feels the pain and has the plumbing to respond; the industry-wide averages will trail by a few years, but that’s diffusion being slow, not correction. The spread between the decoupled and the calcified widens either way, because the correction has an address to arrive at and some addresses are easier to deliver to than others. Authorship-volume metrics become embarrassing in performance systems, roughly the way engagement metrics became embarrassing after 2018. Open contribution stops being the default posture of open source; vouching, provenance, and gated contribution become table stakes, and we will be nostalgic about drive-by patches the way we’re nostalgic about the open early web. Agency-built software gets recognized as its own quality tier, and AI-era debt starts showing up in acquisition due diligence. And wherever the market can’t close a loop on its own, regulation will close it clumsily; the EU’s Cyber Resilience Act is one of the first broad attempts to force software’s externalities back onto their producers, and it will not be the last.

Underneath them all is a less glamorous one. The industry is going to spend the next few years rediscovering, at considerable expense, that our practices (code review, ownership, on-call, blameless postmortems) were never process for its own sake. They were plumbing for routing consequences back to decisions. AI didn’t break that plumbing. It turned up the water pressure and showed us where the joints were loose. And agents reviewing agents won’t retire it either; that just moves the human job up a level, from judging the code to auditing the judgment, and the serial fraction migrates there right behind it.

Closing the loop

Tonight an agent will almost certainly propose another persuasive workaround somewhere in my next project in RISE, and about four times out of five a skill I wrote will bounce it before I even see it: get it right, or document the limitation and put the cleanup on the books. The reflex took twenty years of tripping over my own shortcuts to build. Writing it down took an afternoon, and unlike the twenty years, the writing down scales. Not without limits: rules transfer in proportion to how coherent the system underneath is, and RISE accepts mine because a single mind shaped it. But making larger systems rule-shaped is not a mystery; it’s what style guides, static analysis, and good platform teams have been doing for decades, one hard-won judgment at a time. That’s the encouraging part. The work ahead, for platform builders and engineering leaders alike, isn’t to slow the amplifier down. It’s to route the invoices (every shortcut with a price tag, every price tag with a name on it) and to encode the judgment of the people who’ve actually paid into the system everyone else uses.

Leave a comment