Back to blog

Blog Post

wp2shell: Discovering One of 2026’s Biggest Zero-Days, and the Future of Exposure Management

Share on social

Aug 19, 2026

Lorem ipsum

Table of contents

Share on social

Join the newsletter
wp2shell: Discovering One of 2026’s Biggest Zero-Days, and the Future of Exposure Management

In July 2026, our research team discovered an unauthenticated remote code execution vulnerability in WordPress Core, which we dubbed ‘wp2shell’. This was the first vulnerability of its kind in WordPress core for almost a decade, affecting hundreds of millions of websites. Using GPT-5.6 Sol Ultra, it cost about $25 in compute and took roughly ten hours to uncover, and is something exploit brokers are willing to pay orders-of-magnitude more for. Within hours of the patch publication, public PoCs were circulating, and weaponized exploitation was observed less than two days after the patch shipped.

The severity of the vulnerability, how it was discovered, and the speed at which the timeline accelerated, makes this one of the most impactful case studies for where vulnerability discovery and exposure management is heading.

What actually happened

I won't re-walk the technical chain here. Our researcher, Adam Kues, who found the vulnerability, has already done that far better than I ever could. If you want the full detail, you can read our initial advisory on the vulnerability, and Adam’s full write up of the technical detail and how he discovered it at these links:

On July 17, 2026, we published our advisory and a public exposure checking tool (www.wp2shell.com), as soon as WordPress shipped version 7.0.2 fixing the vulnerability. Searchlight customers had received notification and mitigations two days prior and auto-updates went out to a significant proportion of the customer-base upon the patch release, yet millions of sites remained vulnerable. In our advisory, we provided details of the patch, as well as temporary emergency mitigations for those unable to update immediately.

We deliberately held-back Adam’s full write-up of the exploit chain to give defenders time to patch over the weekend, only releasing it once we saw public PoC’s replicating our exploit.

As we’ve seen many times before with vulnerabilities of this magnitude, it was only a matter of time before the patch was reverse engineered, and in-the-wild exploitation had begun. 

The key factor behind everything that happened with wp2shell comes back to time. The time it took AI to discover the exploit chain, the time it took for exploitation to begin, and the time advantage defenders need to outpace attackers. 

Vulnerability discovery in the AI era

Wp2shell wasn’t discovered through months of manual code review. As you’ll see from the full write-up, Adam used GPT-5.6 Sol Ultra to discover the exploit chain, with a surprising and unprecedented level of creativity.

This isn’t just a one-off. AI and frontier models are fundamentally changing vulnerability discovery and research. A few months ago, our team’s use of frontier AI models helped uncover several critical vulnerabilities in cPanel, another platform sitting quietly behind millions of websites and mailboxes worldwide. AI helped us move faster through reconnaissance and code analysis, ultimately uncovering a string of vulnerabilities. But not before attackers, who, as it turned out, had been exploiting one of the vulnerabilities two months before.

If a $25 experiment can produce a working RCE chain in a piece of software running a third of the web, that same workflow is available to criminal groups and state-linked actors too. We know AI can find bugs like this. The question is who finds them first, and what they do next.

AI is changing how research is conducted, but the researcher is not disappearing from it. With wp2shell, they became the most important part of the process. Adam still had to know where to look, direct the model, interpret and validate what it produced, and then take it through a responsible disclosure process rather than a payday. As he said himself:

“As the models get stronger and stronger, it seems like security research will become a bit higher-level—deciding what products and surfaces to investigate and for how long, steering research direction with prompts, and nudging the LLM when it veers off track. These meta skills are currently still handled quite poorly by AI and will become more and more important as the LLM is able to handle the bulk of the technical exploit development work.”

AI didn't replace his judgment. It made the judgment more valuable, because the alternative, the same discovery happening with none of that oversight, is a materially worse outcome for everyone running WordPress. Offensive security research matters even more than it did a year ago, precisely because AI is going to keep expanding who can find bugs like this and how fast.

This is going to keep happening. We already saw Microsoft release its largest Patch Tuesday update of all time in July, in part due to its own efforts pointing AI models at its products. AI means exponentially more vulnerabilities are being added to the patching backlog; as it stands, the ability for many organizations to remediate has not kept up.

Why the timeline matters

If you want to understand why exposure management has to change, look at the clock.

  • July 17: WordPress ships the patch and we publish our advisory and public checker tool. Full technical detail is deliberately withheld at first, specifically to give defenders time to update before attackers could reverse-engineer the fix.
  • Within hours: Public proof-of-concept exploit code starts appearing. Researchers begin reverse-engineering the patch almost immediately, comparing the fixed code against the vulnerable version to reconstruct exactly what changed and why.
  • July 19th: Mass exploitation campaigns are reported (source)

The entire window from patch release to confirmed real-world exploitation is under two days. Auto-updates and WAF mitigations absolutely worked, for the sites that had them properly configured and switched on. But getting to it eventually was never a viable response to a clock that short.

What ‘beating the patching scramble’ actually looks like

Because our own research team found and reported wp2shell directly, our customers had visibility into it more than two days before the public advisory and patch even existed. That means they were assessing their exposure and beginning mitigation before the public clock had even started, let alone before exploitation was seen in the wild.

The rest of the industry responded fast once the advisory went public. Several respected security teams turned around detailed technical analyses, IOCs, and mitigation guidance within a day or two of disclosure. That's fast incident response, and it deserves credit.

But when it comes to response, everyone else is inherently running the same race as attackers. Our customers weren't in that race at all, they'd already started mitigating before the public timeline even began. A two-day head start against a vulnerability that went from patch to active exploitation in under two days makes a huge difference.

The Takeaway

As we wrote when we published our research into cPanel earlier in the summer, with AI, everyone is equally as elevated towards finding vulnerabilities. That doesn’t kill offensive security research, it makes it all the more important, particularly with the responsible disclosure discipline that turns a discovery into a patching window instead of a weapon. 

But discovery is only half the equation. The other half is what you do once something is found. How fast you can close the window of exposure. wp2shell showed both sides of that clearly: two days after disclosure were enough for global mass exploitation. But two days prior were enough for our customers to already be protected.

Read more of our latest research and how it feeds our Preemptive Threat Exposure Management (PTEM) platform, helping customers protect themselves against zero-days before public disclosure
Tom Duncan

Author

Tom Duncan

Head of Content and Communications in Marketing

Related Blog Posts

August 14, 2026

Beacon: OpenAI's Astra Paused Due to Hacking Use Concerns

August 13, 2026

Phishing and Takedown now managed entirely in Monitor

August 6, 2026

How to Measure Preemptive Threat Exposure Management (PTEM) Success

August 5, 2026

August 4th – This Week’s Top Cybersecurity and Dark Web Stories

July 31, 2026

How Does Preemptive Threat Exposure Management Improve Exposure Prioritization?

July 29, 2026

July 28th – This Week’s Top Cybersecurity and Dark Web Stories

Never miss a beat

Get all news and updates about Searchlight Cyber, directly in your inbox.

Subscribe
Please enter a valid email address.
Background Gradient