My Security Scanner Was Grading the Firewall, Not the Site
Sites behind Vercel's bot protection got failing grades for headers they had set correctly. The fix: detect when you reached a wall instead of the site, and grade only what you actually saw.
contents (6)
WebSentry scans a website and grades its security: HTTPS, security headers, content security policy, cookies and more. Most of its checks read the response the site sends back.
Every one of those checks made the same assumption: the response is the site. Usually it is. Then we scanned a site behind Vercel’s bot protection, and it wasn’t.
What went wrong
When a site is protected by Vercel’s Attack Mode or Cloudflare’s challenge, an automated request like a scanner’s does not reach the site. It reaches an interstitial page, the “checking your browser” screen, served by the firewall in front of it.
That page carries none of the site’s own headers. So WebSentry scored a real, well-configured site as having no content security policy, no HSTS and no X-Frame-Options. The report then told the owner to add X-Frame-Options: DENY, a header they had already set.
That is worse than an incomplete report. It is a confident, specific, checkable claim that is false, made about someone’s security. And it hit exactly the customers who had paid for extra protection.
Did we reach the site, or a wall?
The fix starts with one question, asked before any check scores anything: is this response the site, or something standing in front of it?
WebSentry now looks for evidence of a challenge before trusting a response. Firewall vendors tend to mark their challenge responses, challenge pages have recognisable content, and a challenge page has a different shape from a real page: short, built around a script to solve, and missing the structure a real site has. Any one strong signal is enough; weaker ones have to agree.
A status code alone is not enough
The tempting shortcut is to treat any 403 or 429 as a firewall. That would be wrong in the other direction.
A 403 can be an ordinary permission error, and a 429 an ordinary rate limit. Both are real findings about the site, not a wall in front of it. Treating them as “blocked” would hide genuine problems. So a status code only counts as a challenge when the response also looks like one.
The scanner’s mistake was a false finding. Hiding a real finding would just be the same mistake pointed the other way.
Grade only what you saw
Once a check knows it hit a wall, it does not score the wall. It reports itself as not scanned, with which vendor blocked it and why.
The scoring changed to match. The overall grade is the points earned divided by the points possible. A category that could not be read now contributes nothing to either side: its points leave the denominator. So the grade reflects only what the scanner actually saw, instead of being dragged down by headers it was never shown.
The report says so plainly: which categories were not scanned, and that a firewall stopped the scanner from reading them. A complete grade is given up on purpose. An honest partial grade is more useful than a confident wrong one.
The same wall, twice
This was the second time Vercel’s bot protection fooled one of my monitors. When I turned it on for Plus234Feed, my own health check started reporting the site as down while readers were being served normally. I wrote about that in A Green Checkmark Is Not Evidence.
Both bugs had the same root. A tool that measures a website from outside is also measuring everything in front of it: CDNs, firewalls and bot protection. As that layer becomes the default, any check that assumes it reached the real site will eventually report something false about it.
What I took from it
Check that you measured the thing you think you measured. Before scoring a response, confirm it came from the system you are grading.
“Not checked” is a valid result. A check that cannot run should say so, not fall back to a pass or a fail.
Watch errors in both directions. Calling a real problem “blocked” hides it. Calling a firewall “the site” invents one. The rules need to avoid both.
The people protecting themselves most are the ones a naive scanner hurts most. That alone made this worth fixing properly.