AI agent PageBreak finds 500+ Google web bugs, just 2 in hardened apps

- Google’s AI agent PageBreak has confirmed more than 500 cross-site scripting flaws across Google’s own web apps.
- PageBreak found just two XSS bugs in apps built on Google’s high-assurance frameworks as of September 4.
- PageBreak reports a bug only after a working exploit runs against a live copy of the app.
Google’s own AI agent PageBreak has validated over 500 cross-site scripting vulnerabilities in Google’s homegrown web apps. It found only two in hundreds of apps built on the company’s high-assurance web frameworks.
PageBreak flags a bug only after a working exploit runs against a live copy of the target, Google said.
Debug endpoints held both flaws counted
Both hardened-stack bugs were counted as of September 4, 2026. Both were in internal apps or debug endpoints that were not fully hardened.
Google cites the gap, more than 500 findings across its broader set of first-party apps compared with two on the hardened stack, as evidence that safe-by-design frameworks can withstand a relentless automated attacker.
On September 24, Google’s Product Security team announced PageBreak in a blog post by information security engineer Michał Bentkowski. The agent ran as a pilot starting in November 2025, and became a full project in January 2026.
Cross-site scripting, or XSS, is when an attacker injects a script into a page that another user loads. Depending on the app, the script could read data or commandeer a victim’s logged-in session.
Gemini 3.1 Pro scans, and a validator must fire the exploit
Most of the scans are done on Gemini 3.1 Pro and Gemini 3.5 Flash, but PageBreak can also use other models. The second step is what sets it apart from a normal LLM scanner.
Each suspected flaw is passed by the agent to a purpose-built validator. The validator then fires the actual payload against a running instance of the app.
As for XSS, the validator injects a JavaScript payload, loads the page, and ascertains whether the script executes.
Google is pitching PageBreak as a cure for the “AI slop” drowning security teams. Bentkowski’s post talks about LLMs as static code analyzers overwhelming teams with unverified hypotheses, where the hard part was sifting a real, exploitable bug from a plausible hallucination.
In addition to XSS, the validators test for database query injection, path traversal leaks and code execution.
Google runs the same seed over multiple iterations, so an agent who strays down a dead end still gets repeated chances at the right exploit.
Google says unverified candidates never make it to product teams as confirmed bugs. And they continue to feed into later scans or point engineers building the next validator, still within the security workflow.
Google says PageBreak’s false-positive rate is near zero.
The agent’s reach is magnified by Google’s own scale, which an outside researcher can’t replicate. One code repository enables it to trace execution paths across services, and live-traffic security data maps a page request back to the source code.
Existing scanners hand PageBreak logged-in entry to internal sites otherwise hard to reach.
Google plans to integrate PageBreak more tightly with CodeMender, a fix-writing agent, so that teams can review a proposed patch next to a confirmed bug.
Google cited CodeMender in May when its Threat Intelligence Group said it had caught what it believed was the first zero-day exploit created with AI assistance.
Bitcoin Red Team’s August sweep of 501 open source projects generated 7,958 findings over 108 hours, but only 24.7% had reproducible proofs at the time.
Autonomous agents already cross boundaries they weren’t meant to, as Cryptopolitan reported when Gemini reached three real companies during May testing.
The smartest crypto minds already read our newsletter. Want in? Join them.
FAQs
What is PageBreak and what has it found?
PageBreak is an internal AI agent from Google's Product Security team. It has confirmed more than 500 XSS flaws in Google's web apps.
How does PageBreak avoid false positives?
A validator runs a real exploit against a live copy of the app before PageBreak reports anything.
Why did Google's hardened apps have only two bugs?
They run on Google's high-assurance, safe-by-design web frameworks, which held PageBreak to two XSS bugs.
Disclaimer. The information provided is not trading advice. Cryptopolitan.com holds no liability for any investments made based on the information provided on this page. We strongly recommend independent research and/or consultation with a qualified professional before making any investment decisions.

Randa Moses
Randa Moses is an editor and reporter at Cryptopolitan covering tech, AI, robotics, crypto, scams, and hacks. She has worked in the crypto space since 2017. She held roles at Forward Protocol, AmaZix, and Cryptosomniac. Randa holds a degree in Electrical and Electronics Engineering from the University of Bradford.
















