I launched this site yesterday…. Just a little idea I had in the morning and I had it up and running in an hour or so. Obviously no one, an nothing, knew it existed. Except within MNUTES of launching it we saw fifty bots looking for a way in.
First, I want to point out that this site has nothing to probe. This is an HTML page that sits on an S3 bucket with cloudfront. For the most part you don’t even get to the S3 bucket, it’s just a little html sitting on a CDN. I call this my “bullet proof web” architecture because there isn’t any compute to compromise, only think you’d hack would be your own browser.
So it’s actually kind of adorable watching them try to find a way in… They don’t know what’s here, just that SOMETHING just stood up, and often new stuff is the least secure, with little back doors open wide enough to drive multiple horse trailers through.
They’re basically looking for a couple things:
- Full stack admin tools: Mainly Wordpress but also Laravel and just a bunch of PHP admin tools, they assume we’re running a stack with PHP and a database and that we have some publicly accessible admin tools that might be easy to run without a password.
- Credentials: The number of variations on .env (.env.bak, .env.prod, /webapp/.env…) and on .git/config is an amazing brute force search for passwords or deploy tool tirggers, or other compute that might not be on this server but still get them access to your compute.
- CGI: I find this one amusing because it’s 2026 but literally things like mailer.cgi and other tools we should retired decades ago are still in the probes because PEOPLE STILL USE THEM.
The reason that this shiny brand-new website got scanned within minutes of being published is because I did the right thing (the kind of required thing in modern web apps) and added a TLS cert, that’s the bit that gives you the https:// in the domain and encrypts this article from the source (CloudFront in this case) to your browser. Overkill for a blog, but modern browsers really want secure connections.
Thing is, using a publicly listed cert service like AWS sends out an invitation to all the bots on the planet. You see, every publicly-trusted TLS certificate is written to Certificate Transparency logs when it gets created. CT is a good thing that gives us a way to catch certificates that should never have been issued.
Unfortunately, it also means that the moment you ask for a certificate, you've announced your domain to anyone subscribed to the firehose. There is no security by obscurity… You get https://yoursite and within minutes you’re going to get probed.
In Certifiably Vulnerable: Using Certificate Transparency Logs for Target Reconnaissance (EuroS&P 2023), researchers stood up honeypot web servers on IP addresses that had never been used for anything, and ran them for 200 days. They saw probes arriving "just seconds after publishing the entry."
But it’s not like a self-signed cert is going to stop the probes… the whole point of being online is to be connected, so eventually DNS, backlinks, SEO… even if it’s not a website (like an API or a gateway) your service will be discovered and probed.
So a big part of “scale down to zero” in my mind is “Zero Trust” in the way that I’m not giving you anything I don’ have to.
GreyNoise runs sensors in more than eighty countries and publishes what lands on them. Their 2026 State of the Edge report covers 162 days, July 23rd to December 31st 2025, and saw 2,969,010,478 malicious sessions from 3,804,232 unique source IPs. Yes, 2 BILLION sessions in 6 months from 3 MILLION systems. (Fun fact: Los Angeles population is a little over 3 million… so that’s like mobilizing ever man, woman and child in LA at your web app).
52% of remote code execution attempts came from IPs GreyNoise had never seen before. There's no stable cast of known-bad hosts to block here. The fleet recycles itself constantly. One credential-spraying botnet in the same report grew from 2,000 IPs to 300,000 in 72 days.
Secondly, and this goes back to the cgi scripts and all that PHP probing, pre-2015 CVEs generated four times more exploitation traffic than 2023-2024 CVEs. 7.3 million sessions against decade-old bugs, against 1.8 million for the recent ones.
Nobody really thinks a new blog on internet application architecture is going to be running Mailman. Nobody's thinking at all. Those paths stay on the list because keeping them costs nothing, and somewhere out there somebody still is still using it and then… we’re in.
That's what attack surface means in practice. It isn't a threat model. It's a dictionary, and you're spell-checked against it a few hundred times a day, forever.
Now don’t forget those environment files and .git stuff. Even scaling to zero needs SOME compute somewhere. In August 2024, Unit 42 published an analysis of an extortion operation where hackers had scanned more than 230 million unique targets and hit 110,000 domains with exposed .env files. The hackers were able to pull over 90,000 unique env variables, including 7,000 access keys for cloud services and 1,515 credentials belonging to social media accounts.
Then they were able to create Lambda functions with administrative IAM roles inside their victims' own AWS accounts, then they used those to run the next round of scanning. They also were able to start EC2 instances (full compute stacks) for cryptomining, get data out of S3 buckets, delete objects, and leave ransom notes where the data had been.
It isn't only credentials, either. GreyNoise logged 91,403 sessions against exposed Ollama servers between October 2025 and January 2026. Somebody is combing the internet for other people's GPUs to run inference on. The new attack surface is a machine you left switched on.
Thing is, you’re going to get scanned, it’s like weather. You can’t control if it’s going to rain, but you can make sure you don’t get wet. To me this is a big part of Flat-Stack… scaling down means you don’t have a bunch of stuff out in the yard when it starts raining.
This blog costs pretty much nothing to run, even with the scans. The content is cached in cloudfront for a very long time (if it changes, we invalidate, so we aren’t even going back to the source S3 storage very often). Those probe requests cost almost nothing to serve, and on a CDN they're a rounding error. On a $15-a-month always-on box they're free at the margin too, because the box costs $15 whether the internet scans it or not.
But even if they aren’t stealing your compute, probes can run up the bill if they’re hitting Lambda, Kubernetes, ECS, or some other autoscaling fleet – I literally was responsible for a WordPress stack that, don’t ask why, we had to scale up to 8 EC2 instances to deal with what was essentially a self inflicted denial of service attack… pretty much once a week we’d go from one tiny instance to 8 big ones… probes triggered the same thing.
You're paying to keep an interpreter, application runtime, database, containers, pods, or instances warm and reachable so that the scan has something to talk to when it arrives.
The hosted platforms look like the answer but not quite. Handing the stack to somebody else is genuinely valuable: you stop patching, and for a great many of the sites on the 41% of the web running WordPress today, having someone else patch is way better than you patching badly.
But nothing got subtracted. The stack is the same size or bigger. It moved somewhere you can't see it, can't audit it and can't make decisions about it, and your whole perimeter collapsed into one account credential. Fewer things you have to do isn't the same as fewer things that can go wrong.
All of the above is the argument for the architecture this site runs on, so here it is, in the order the probes would have hit it.
The site is files. Articles are JSON. A renderer turns them into HTML, one stylesheet, a feed and an index, and that output is the deployable artifact. There's no template engine at request time because the templates already ran, and you can't inject into a render that finished yesterday.
Two buckets, different jobs. Everything is written to staging first: drafts, uploaded media, the article store, the admin list. Publishing is a server-side copy from staging to live. The copy matters less than the split behind it: the bucket the internet reads is never the bucket anything writes to. There's no path from a GET to a write, and a path from a GET to a write is what every request in that log was ultimately shopping for.
Both buckets are private. Neither is an S3 website endpoint. CloudFront reaches the live one through an Origin Access Control and nothing else reaches it at all. Four requests that morning carried list-type and four carried max-keys, which aren't web parameters at all; they're the S3 ListObjectsV2 API, somebody checking whether the origin will enumerate itself. It won't. It doesn't take questions from the internet.
The deploy leaves no residue. There's no .env, because nothing runs that needs secrets at runtime. There's no .git in the document root, because the document root is a rendered directory rather than a checkout. There's no wp-config.php.swp, because there's no wp-config.php and nobody has ever had a reason to edit one there. Those files are missing as a consequence of the build's shape, not because we remembered to hide them, which is the kind of defense that works right up until the once it doesn't.
Now this isn’t to say that we don’t have compute – there’s a content management tool that let’s me edit this article, preview it, publish it… but that compute is a lambda hidden away in a very different spot again with a lot of layers of CDN and cloudfront rules keeping it asleep – the probes don’t reach it because it doesn’t need the public face that this blog needs.
So.. to sum up… It’s going to rain. The smaller your web app is, the easier it is to keep it safe from the rain. And the less you have to probe… the less you care about the weather…
Let it rain.