Trending: On-device modelsSearch
iHeartGeek
iTECH

Cloudflare let frontier AI models attack its own firewall

The company pointed an LLM at a customer staging environment to see whether a model that mutates payloads faster than a human could slip past its web application firewall. It found 558 blocks, 49 leads and one trailing dot worth a second look.

Cloudflare's blog header art: the white headline “We tested our own WAF with frontier AI models. Here's what we found.” on a dark navy panel beside a large glowing shield, graded purple to orange, carrying the Cloudflare flame mark and ringed by thin line-art icons of clouds, a magnifying glass over a warning triangle, traffic charts and server racks, with an orange stripe across the bottom

Cloudflare has turned frontier AI models loose on its own web application firewall and published the scoreboard. The company built a tester that drives a large language model through repeated variations of a known exploit, sent it at an authorised customer staging environment, and logged 1,107 attempts across 45 scenarios and six attack categories. After human triage, 558 of those requests were stopped by the firewall and 49 were counted as findings worth investigating — 48 of them in command injection and server-side request forgery (SSRF).

The model never saw the rules

The value of an LLM in this role is not knowledge, it is stamina. A model can iterate on encodings, move a payload from the query string into a form body, or try the same destination written another way, faster than a person working through a checklist. Cloudflare's tester ran two model calls per step: a proposal call that suggests the next variation from the starting request and a short history of earlier results, and a review call that reads the response status, selected headers and body to decide what to try next.

Neither call was given anything about the defences. No rule expressions, no rule identifiers, no WAF Attack Score detail and no indication of which security layer had acted. The harness itself is Python rather than a wrapper around an existing penetration-testing tool, and the code — not the model — sends each request: it checks the target hostname against an allowlist, disables redirects, enforces a 25-attempt limit per scenario and treats response text as untrusted input. Neither model call can deploy a rule or change enforcement.

Forty-four of the 45 scenarios covered cross-site scripting, SQL injection, command injection, SSRF, path traversal and local file inclusion, and Log4j; the remaining scenario covered log injection and is reported separately. The test zone ran a deliberately strict configuration: WAF Attack Score blocking scores of 30 or below, the full Cloudflare Managed Ruleset switched on, and the OWASP Core Ruleset at Paranoia Level 3.

One trailing dot, one clean miss

The published example is an SSRF run at cloud metadata, the address range that can hand out temporary credentials to whatever workload asks nicely. The tester sent the same address as a dotted quad, as the decimal integer 2852039166 and as the octal form 0251.0376.0251.0376, in different parts of the request. All were blocked until attempt 18, when the model kept the request structure identical to the previous blocked attempt and switched only the host to a trailing-dot form. The client met a redirect instead of a block page, so Cloudflare kept the request as an edge-pass observation — while noting there was no origin response, no body and no evidence that anything was fetched.

Across the run, 1,107 recorded attempts became a post-triage set of 607: the 558 blocks plus the 49 findings. Cross-site scripting, local file inclusion, SQL injection and Log4j came out with near full coverage. The rest of the noise was discarded for good reasons — the model failed to produce a usable HTTP request, the request never reached the target, or the mutation had made the payload harmless. Anything unblocked then had to survive five questions, including whether it was still malicious and whether the behaviour even belonged to the firewall rather than to DNS or the network.

Findings became three rule changes

The 49 findings were grouped into four sets of candidate rules, each one replayed and validated, then tested against live traffic for false-positive risk before it could protect anyone. The work fed three changes to Cloudflare's Managed Ruleset in the 21 July release: new SSRF - Obfuscated Host and SSRF - Restricted Protocol detections, and an improvement to the existing SSRF - Cloud rule. The obfuscated-host detection came directly from requests that wrote internal addresses in non-standard numeric forms — the same trick that produced the trailing-dot lead.

Two versions of the same model family were run over identical scenarios and produced different variations, with the same underlying gaps surfacing in both. Pushing more attempts into a single scenario stopped paying off near the 25-attempt ceiling, where the model began repeating itself; wider coverage came from more starting requests, categories and input locations instead. Cloudflare is explicit that the model only generated requests, and that nothing counted as a finding without replay and human review.

For customers the advice is unglamorous. Check that Managed Rules and Attack Score are configured in front of the application, run new rules in log mode, review the matches in Security Events and confirm legitimate traffic is unaffected before moving a rule to block. Layer API Security, bot and fraud detection and threat intelligence on top, and test against a staging hostname protected by the same controls as production. Patching still matters most: a payload that gets past a firewall still needs a vulnerable application to do anything. A follow-up post will cover a white-box run where the model knows both the application's flaws and the rules defending it.

Our opinion

The headline number here is not the 49 findings, it is the 45 scenarios. Cloudflare has published the difference between what a fixed regression suite checks and what an adaptive adversary will try, and it has done so with the misses attached — the trailing-dot case went out the door as a redirect that the customer's application, not the firewall, would have had to be broken to exploit. That is a far more useful disclosure than a block-rate percentage bolted onto a product page. The engineering claim worth taking seriously is narrower than the marketing line: the model produced candidates at volume and humans turned them into three rule changes. Strip the review step out and the same loop generates noise.

The awkward part is what it implies about everyone else's test suites. If an iterating model finds normalisation edge cases that a static list of payloads walks past, then a firewall vendor that never runs this kind of loop is selling a boundary it has not stress-tested — and one that runs it privately has no reason to tell anyone the misses. Cloudflare gets credit for publishing the trailing dot and for keeping the WAF in its correct place as one layer among several rather than a fix. The white-box follow-up is the one to watch, because a model that can read both the application and its rules is where this stops being a testing technique and starts being a different threat model altogether.