I'll start with two of my own, because I didn't need the news to find this one.

The first: I was cleaning up our deploy process and found it was faithfully standing up a set of SQS queues. Provisioned, permissioned, monitored. Nothing published to them. Nothing consumed from them. They were a dead end from an idea we had once and never wired up. When I went to pull them out, the AI working with me pushed back and treated them as part of the architecture, something to protect. It had read the infrastructure-as-code, saw the queues, and concluded they must matter because they were there.

And there's the whole problem in one sentence. It was there, so it must matter.

The second thing I noticed was an agent that needed to know when new data landed in the warehouse. Its solution was to build a view, query it, tear it down, and do it again. Every five minutes… all day… all night… Nobody asked it to be that aggressive, and it wasn't wrong, I mean not exactly, because it did find the new data. It just woke the Beast 288 times a day to ask a question that had a yes/no answer, and (because it wasn’t running through Airbrx.ai) we only found out by seeing the databricks bill spike.

Neither of these was a rogue agent. Neither did anything you'd write a postmortem about. They're just what happens when the thing doing the building has only ever seen architectures that go up.

The headlines are about deletion. The bill is about addition.

The AI horror stories that go viral are the dramatic ones. The DataTalks.Club course platform lost its production environment when an AI coding agent picked up an old Terraform state file and then got permission to run terraform destroy; the database, the VPC, the ECS cluster and the automated backups all went with it. That's a real story and a scary one.

But deletion is noisy, people and monitors start complaining in minutes. Adding infrastructure is quiet and you usually find out on the invoice, and sometimes you don't find out at all, because the thing that got added is working, but it's working on something nobody needs.

Exhibit A: five m8g.12xlarge instances to scan a hobby network

The best example of this year is the DN42 incident, written up in detail by Lan Tian and covered by InfoQ.

Someone gave an autonomous agent full AWS credentials and a deadline and told it to index DN42, a volunteer network where people practice BGP on small VPS boxes (we're talking network security on really cheap virtual servers). The agent decided the job needed five m8g.12xlarge instances, 48 vCPUs each, plus load balancers and Lambda functions, and justified it as scanning at 20 Gbps with redundancy and failover.

Redundancy and failover. For a port scan. Of a hobby network.

Then, every time it hit an error, it reapplied its CloudFormation template and duplicated the whole stack. The operator found out a day later when their card got hit for $6,531.30. AWS knocked it down to $1,894 and the operator went to the DN42 mailing list asking for donations to cover it. The community's estimate was that the workload would have fit on a $5-a-month VPS.

The agent wasn't being stupid, it wasn't being lazy, it just did exactly what it learned from us: this is what "serious" infrastructure looks like. Multiple instances, a load balancer in front, serverless glue, redundancy everywhere. It built the architecture diagram from every re:Invent talk it ever read.

It scaled up because up is the only direction it has ever seen.

Exhibit B: the default is "production-grade"

InfoWorld ran a piece this summer on agent-generated infrastructure bloat. Turns out when the agent writes Terraform for a Kubernetes cluster, it defaults to instance families and node counts built for resilience, not efficiency. Their example is a three-node cluster of n2-standard-16s for a service that runs comfortably on a single e2-medium.

For perspect That's 32 times the compute. It doesn't show up in staging. It shows up every hour, forever, in production.

Their fix was to put constraints in the spec so the agent reads "pick the smallest machine that fits the measured workload" before it writes a single resource block. Which is a good fix, one that I would think the agent would already have defaulted to but…  the agent's instinct, left alone, is to reach for the big box.

Exhibit C: the warehouse doesn't know it's talking to a robot

My every-five-minutes view example isn't unique either. Seemore Data has a good breakdown of what happens when agents get warehouse access: the agent has no intuition for what a query costs, it parallelizes aggressively, and the fan-out can trip the warehouse's auto-scaling. A runaway loop can put a four-figure number on the bill in minutes.

This is part of a monitoring issue, an agent that runs an expensive query successfully, over and over, doesn't look like a problem. It looks like a healthy workload. The agent that was rebuildng the view never errored out, it was just chugging along and everything was okay.

Exhibit D: AI adds code, it doesn't take it away

This is the part that explains my SQS queues I didn't know existed.

GitClear's June report, The Maintainability Gap, looked at 623 million code changes from 2023 through 2026. The volume of moved code (code moving from one place to another and their signal for refactoring) fell from 21% of changed lines in 2022 to 3.8% this year. Duplicated blocks are up 81%. And the one that matters most for this column: changes that touch or remove code older than a year are down 74%. Basically the AI is protecting old code so it doesn't “move.” If it's there, keep it is deeply baked into the models. 

Now look at infrastructure instead of code. The AI is very good at adding the new thing. It is not in the business of going back to delete the old thing. Old things are context. Old things are "the architecture." Old things get protected. That's exactly what happened with our queues. Nobody decided to keep them. They just had seniority.

Exhibit E: the meter runs at company scale too

Uber is the big-number version. The company burned through its entire 2026 AI budget by April, after rolling agentic coding out to its engineers and ranking teams on internal leaderboards by how much they used it. The CTO said they were back to the drawing board. A month later the COO told a podcast that the connection between all that usage and more useful features for customers “is not there yet.”

I'm not saying Uber made a bad bet. I'm saying when you reward usage, you get usage. A leaderboard that ranks consumption is an autoscaling policy with no max.

What the Flat Stack says about this

We wrote in the manifesto that AI removes the last friction that was left: the model adds the Redis because production systems have a Redis. I thought that was a prediction. It turned out to be a description.

  1. Every dependency is a decision.
    An agent doesn't make decisions about dependencies, it inherits them. If the queue is in the template, the queue is a requirement. Somebody has to be allowed to say "this is dead, delete it," and the agent has to be told that's a legitimate answer.
     
  2. Do not wake the Beast.
    The “rebuild the view every five-minutes” should have been a single cheap check: has the manifest, the watermark, the last-modified key changed? If no, go back to sleep. The warehouse only wakes when the answer is yes.
     
  3. Build it stateless. Build it to scale to zero.
    The DN42 agent never considered zero. It never even considered one. The right answer was a $5 box, and nothing in its training said small was an option.
     
  4. The balance sheet is the product owner.
    The agent has no idea what anything costs. If cost isn't in the spec, it isn't in the decision. And the bill arrives a day later, after the agent has moved on.

So what do we actually do

A few things we've started doing, and none of them are fancy:

  • Make "remove" a first-class task. When an agent touches infrastructure, the prompt includes: list what this change makes unnecessary. Our queues would have been gone months ago.
     
  • Put the budget in the spec, not the dashboard. Smallest instance that fits, zero replicas by default, no scheduled job without a stated reason for its frequency. The agent reads constraints. It doesn't read invoices.
     
  • Give agents a cheap question before the expensive one. Every polling loop gets a "did anything change?" check against object storage before it's allowed to touch the warehouse.
     
  • Scope the credentials to the blast radius. The DN42 agent had provisioning rights on an account with no ceiling. An agent that can only launch t-shirt-small instances can only make t-shirt-small mistakes.

And what we give up: some speed. The agent is slower when it has to justify every resource, and sometimes it asks a question a human has to answer. I think that's the point. The friction the cloud removed was doing a job. We're putting a little of it back on purpose.

The agents aren't going rogue. They're doing exactly what the internet taught them production looks like. Up is the default. Zero has to be designed.


Sources