October 1, 2026
A WAF Pass Is Not an Exploit: Reading Cloudflare’s AI Test Carefully
1,107 adaptive requests produced 49 WAF-relevant findings. Those are different measurements from application compromise.
By Saroyan Zorayr
2 min read
1,107 adaptive requests produced 49 WAF-relevant findings. Those are different measurements from application compromise.
Cloudflare's September 29 write-up on adaptive WAF testing has an easily missed distinction: the model generated request variations, but a request that was not blocked was only a lead for human review. It was not a confirmed exploit.
That distinction is worth keeping, especially when "AI broke the WAF" would make a much louder headline.
What was actually tested
The team used an authorized customer staging environment with a specific Cloudflare configuration: Managed Rules enabled, OWASP Core Ruleset at paranoia level 3, and WAF Attack Score blocking at 30 or below. An allowlisted test User-Agent helped requests reach the WAF rather than being stopped earlier by automated-traffic controls. The model had no source-code view or access to WAF rules. Code controlled request replay, hostname allowlisting, redirects, and attempt limits.
Across 45 scenarios, it recorded 1,107 mutation attempts. Cloudflare says 558 requests were blocked and 49 survived human triage as WAF-relevant findings; 48 of those involved command injection or server-side request forgery (SSRF). Many other attempts were malformed, benign, duplicated, out of scope, or never reached the target. The 49 are not 49 successful application compromises, and this one staging configuration is not a universal WAF "bypass rate."
The useful SSRF example
In one recorded sequence, a changed representation of a cloud-metadata destination produced a redirect rather than the expected WAF block. That was a real edge-observation difference. Cloudflare also states there was no successful origin response, no response body, and no evidence that the application fetched metadata. The right next step was safe replay and origin-side validation, not a claim that credentials were stolen.
The measurement chain should be explicit: was a valid request sent; did it reach the WAF; was it blocked; did the application process it; and was any security boundary at the origin actually crossed? Collapsing those into one "bypass" count obscures the answer defenders need.
What I would change in a web-security review
Test staging with the same edge controls and application routes as production. Keep examples that can be reproduced, then verify the origin's behavior. Feed confirmed edge gaps into regression tests, while tracking application fixes separately from WAF rule changes. Cloudflare says this exercise contributed to new SSRF detections and an update to an existing rule; the findings still required validation and false-positive checks before deployment.
For SSRF, defense cannot stop at a WAF. OWASP's prevention guidance covers destination allowlisting where feasible, careful IP and domain validation, redirect handling, and network-layer controls. The application should enforce what destinations it is allowed to fetch even if a request passes the edge.
My takeaway: adaptive models can be useful test generators. The engineering value comes from the harness, stable evidence, human triage, and an honest separation between edge pass and origin exploit.
Which metric does your team report today: WAF blocks, validated edge gaps, or confirmed application impact?