The Internet Archive is asking donors (that’s you and me) for a recurring donation of $25 or more in September. A 2:1 match turns it into $75 of support; $100 a month becomes $300. So before I get into the tech stuff, may I suggest this is a good time to donate and a good thing to do.
Their post is pretty dang candid about what the money is actually for… they aren’t adding flashy new features or doing marketing… Servers, storage, power, cooling, and the people who keep it all running is breaking the bank.
They don’t contract the core technology out to “hypersclaers” like AWS or GCP, they build and maintain their own systems which how they 210 petabytes of the open web, and it's also why they say in the post that “our needs are growing rapidly.”
This wasn’t a bad architecture decision or a vendor bill nobody noticed, it was just more usage and more infrastructure that keeps growing forever and it’s funded by people giving $25 a month one tiny chunk at a time.
Keep in mind, the Internet Archive is totally flat-stack philosophy… they are the pragmatic librarians with a mission and they aren’t splashing money around casually. Let’s run through a few of the Flat Stack Manifesto rules.
| Rule | Grade | Why |
| 1 — Do not wake the Beast | C | Hot-URL caching against a predictable key doesn't appear to be in place |
| 2 — Center of gravity low | A | Self-hosted, open, no vendor tower |
| 3 — Balance sheet is the product owner | B | Honest cost ask, but doesn't separate growth drivers |
| 4 — Security, availability, resilience | C | Single-datacenter power failure took the site down for hours |
| 5 — Every dependency is a decision | B+ | Mostly self-built; 2024 exposure ran through a 3rd-party helpdesk tool |
| 6 — Stateless / scale to zero | B | Long-tail reads fit well; the popular 1% still runs always-on |
| 7 — No shiny objects | A- | No sign of chasing new tools |
| 8 — Modernization baked in | B+ | Looks continuous, not a bolted-on overhaul |
| 9 — Document your schemas | A | CDX index format is public and has been stable for years |
| Overall | B- | Philosophy is fully aligned; execution on caching and resilience lags the mission |
Rule 2 — keep your center of gravity low. These guys are one of the few organizations we've evaluated that already embraces this rule better than most venture-backed companies do. There's no proprietary engine standing between the Wayback Machine and the what they’re serving. If every article here ends with "flatten your stack," the folks at the Archive already flattened the part that's usually hardest: it never built the tower in th because they never built the big stack in the first place.
Rule 1 — do not wake the Beast. They have petabytes of cold and warm storage that has to answer a lookup every time someone requests a URL and a timestamp. Some of those lookups are for a page nobody has asked for in a decade (and I’m constantly amazed how I can find remnants of work I did 25 years ago there). Sure, there are people looking at the same hundred thousand popular URLs, over and over, from every browser extension and research tool and citation checker on the internet, and that’s the classic cache thing we solve at airbtrx.ai with our deterministic keys.
They only wake the deep archive for requests that are actually new. I have no idea how much of the Archive's serving traffic already works this way. We'd guess it's less than it could be, because… rule 4…
Rule 4 — security, availability, and resilience are not features. — On August 16, 2026, a power feed failure at one of the Archive's own datacenters took the Wayback Machine, search, and uploads down for hours, with search and some archived pages still degraded days later.
When you own every layer yourself, the mandate that "no single vendor can take us down" is true, what we don’t say in the Flat Stack Manifest is that “no single utility feed should be able to take you down either.”
A thin, cached read layer in front of the hottest lookups like we do with the Airbrx gateway would have kept a lot of the Wayback Machine traffic alive even if the origin was dark. Resilience through fewer moving parts cuts both ways: fewer parts also means fewer places to absorb a hit when the one part that's up goes down.
Rule 3 — the balance sheet is the product owner. The blog post is a pretty straight forward balance sheet request… What it's missing is the question I’;m always asking… do we really need all that compute?
"Our needs are growing rapidly" didn't say whether that growing slope is be cause requests are growing, storage is growing or both. And I’m curious whether any of it could be decoupled from cost with a caching layer instead of just met with more hardware. A flat-stack read of their request needs both numbers: what a dollar buys today, and what a dollar buys after the hot path is flattened (if it can be flattened at all).
Rule 6 — build it stateless, build it to scale to zero. Crawling and ingestion are naturally chatty but there's no reason that work needs always-on capacity between crawls. And serving the long tail of archived pages (like that plumbing company site I built in 1998 that nobody has asked for since… well… probably 1998) is about as close to stateless request/response as infrastructure gets. The part of the Archive's workload that isn't a good fit for scale-to-zero is the popular 1%, and that's precisely the part Rule 1's caching layer is supposed to carry instead.
Rule 9 — document your schemas. Definitely points on this one… the CDX index format the Wayback Machine is built on is public, documented, and has been stable for years. That's the rule working exactly as intended — a documented schema that outside tools and researchers can build against instead of guessing at.
Rules 5, 7, and 8 — dependencies, shiny objects, modernization. Nothing in the public record suggests the Archive is chasing new tools for their own sake, and a mostly self-built stack is close to the minimum-dependency ideal by construction. The one flag worth naming: the 2024 credential exposure that reportedly ran partly through a third-party helpdesk tool is a reminder that even the leanest self-built stack has a handful of outside doors, and every one of them is a decision worth revisiting on a schedule, not just when it breaks.
Where we'd like to help
We build the gateway this argument keeps describing — a control plane that sits in front of an expensive origin, answers what it can from a predictable key, and only wakes the heavy system behind it when a request genuinely needs it. That's a different traffic shape than the SQL warehouses we usually put it in front of, but it's the same architectural problem: repeated reads against an origin that bills, or strains, every time it's touched.
The Internet Archive already believes in almost everything this column measures against. It publishes its schemas, owns its data, and turns every dollar it's given directly into infrastructure it controls. That's rare enough that we'd rather contribute engineering than just write about it. We'd like to offer Airbrx's gateway work and flat-stack consulting time to look at the read-heavy side of the Archive's traffic — as a public engineering project, scoped and published in the open, the same way the Archive publishes everything else it does. Not a pitch. A donation of a different kind, from one organization that tries to run flat to another that already does.