Share on social
Aug 19, 2026
Lorem ipsum

In July 2026, our research team discovered an unauthenticated remote code execution vulnerability in WordPress Core, which we dubbed ‘wp2shell’. This was the first vulnerability of its kind in WordPress core for almost a decade, affecting hundreds of millions of websites. Using GPT-5.6 Sol Ultra, it cost about $25 in compute and took roughly ten hours to uncover, and is something exploit brokers are willing to pay orders-of-magnitude more for. Within hours of the patch publication, public PoCs were circulating, and weaponized exploitation was observed less than two days after the patch shipped.
The severity of the vulnerability, how it was discovered, and the speed at which the timeline accelerated, makes this one of the most impactful case studies for where vulnerability discovery and exposure management is heading.
What actually happened
I won't re-walk the technical chain here. Our researcher, Adam Kues, who found the vulnerability, has already done that far better than I ever could. If you want the full detail, you can read our initial advisory on the vulnerability, and Adam’s full write up of the technical detail and how he discovered it at these links:
- Advisory: wp2shell: Pre Authentication RCE in WordPress Core
- Full write-up: Exploit brokers pay $500,000 for a WordPress RCE. I found one with GPT5.6 Sol Ultra and $25
On July 17, 2026, we published our advisory and a public exposure checking tool (www.wp2shell.com), as soon as WordPress shipped version 7.0.2 fixing the vulnerability. Searchlight customers had received notification and mitigations two days prior and auto-updates went out to a significant proportion of the customer-base upon the patch release, yet millions of sites remained vulnerable. In our advisory, we provided details of the patch, as well as temporary emergency mitigations for those unable to update immediately.
We deliberately held-back Adam’s full write-up of the exploit chain to give defenders time to patch over the weekend, only releasing it once we saw public PoC’s replicating our exploit.
As we’ve seen many times before with vulnerabilities of this magnitude, it was only a matter of time before the patch was reverse engineered, and in-the-wild exploitation had begun.
The key factor behind everything that happened with wp2shell comes back to time. The time it took AI to discover the exploit chain, the time it took for exploitation to begin, and the time advantage defenders need to outpace attackers.
Vulnerability discovery in the AI era
Wp2shell wasn’t discovered through months of manual code review. As you’ll see from the full write-up, Adam used GPT-5.6 Sol Ultra to discover the exploit chain, with a surprising and unprecedented level of creativity.
This isn’t just a one-off. AI and frontier models are fundamentally changing vulnerability discovery and research. A few months ago, our team’s use of frontier AI models helped uncover several critical vulnerabilities in cPanel, another platform sitting quietly behind millions of websites and mailboxes worldwide. AI helped us move faster through reconnaissance and code analysis, ultimately uncovering a string of vulnerabilities. But not before attackers, who, as it turned out, had been exploiting one of the vulnerabilities two months before.
If a $25 experiment can produce a working RCE chain in a piece of software running a third of the web, that same workflow is available to criminal groups and state-linked actors too. We know AI can find bugs like this. The question is who finds them first, and what they do next.
AI is changing how research is conducted, but the researcher is not disappearing from it. With wp2shell, they became the most important part of the process. Adam still had to know where to look, direct the model, interpret and validate what it produced, and then take it through a responsible disclosure process rather than a payday. As he said himself:
“As the models get stronger and stronger, it seems like security research will become a bit higher-level—deciding what products and surfaces to investigate and for how long, steering research direction with prompts, and nudging the LLM when it veers off track. These meta skills are currently still handled quite poorly by AI and will become more and more important as the LLM is able to handle the bulk of the technical exploit development work.”
AI didn't replace his judgment. It made the judgment more valuable, because the alternative, the same discovery happening with none of that oversight, is a materially worse outcome for everyone running WordPress. Offensive security research matters even more than it did a year ago, precisely because AI is going to keep expanding who can find bugs like this and how fast.
This is going to keep happening. We already saw Microsoft release its largest Patch Tuesday update of all time in July, in part due to its own efforts pointing AI models at its products. AI means exponentially more vulnerabilities are being added to the patching backlog; as it stands, the ability for many organizations to remediate has not kept up.
Why the timeline matters
If you want to understand why exposure management has to change, look at the clock.
- July 17: WordPress ships the patch and we publish our advisory and public checker tool. Full technical detail is deliberately withheld at first, specifically to give defenders time to update before attackers could reverse-engineer the fix.
- Within hours: Public proof-of-concept exploit code starts appearing. Researchers begin reverse-engineering the patch almost immediately, comparing the fixed code against the vulnerable version to reconstruct exactly what changed and why.
- July 19th: Mass exploitation campaigns are reported (source)
- July 21st: wp2shell is listed on CISA’s Known Exploited Vulnerabilities catalog
The entire window from patch release to confirmed real-world exploitation is under two days. Auto-updates and WAF mitigations absolutely worked, for the sites that had them properly configured and switched on. But getting to it eventually was never a viable response to a clock that short.
What ‘beating the patching scramble’ actually looks like
Because our own research team found and reported wp2shell directly, our customers had visibility into it more than two days before the public advisory and patch even existed. That means they were assessing their exposure and beginning mitigation before the public clock had even started, let alone before exploitation was seen in the wild.
The rest of the industry responded fast once the advisory went public. Several respected security teams turned around detailed technical analyses, IOCs, and mitigation guidance within a day or two of disclosure. That's fast incident response, and it deserves credit.
But when it comes to response, everyone else is inherently running the same race as attackers. Our customers weren't in that race at all, they'd already started mitigating before the public timeline even began. A two-day head start against a vulnerability that went from patch to active exploitation in under two days makes a huge difference.
The Takeaway
As we wrote when we published our research into cPanel earlier in the summer, with AI, everyone is equally as elevated towards finding vulnerabilities. That doesn’t kill offensive security research, it makes it all the more important, particularly with the responsible disclosure discipline that turns a discovery into a patching window instead of a weapon.
But discovery is only half the equation. The other half is what you do once something is found. How fast you can close the window of exposure. wp2shell showed both sides of that clearly: two days after disclosure were enough for global mass exploitation. But two days prior were enough for our customers to already be protected.


.png)




