Thu 06 Aug 2026 / 09:18 ET
Kernel
Internet 3 min read

OpenAI Atlas WhatsApp spam test shows AI browser prompt-injection risk

Zenity says a malicious webpage could induce Atlas to message WhatsApp contacts, though the test did not exploit WhatsApp or show use in the wild.

June Castellano

By June Castellano / Platforms & Power Reporter

OpenAI Atlas WhatsApp spam test shows AI browser prompt-injection risk
img: WIRED

Security firm Zenity says it found a way to turn an already signed-in WhatsApp Web session into an OpenAI Atlas WhatsApp spam machine. In a proof-of-concept presented at the Black Hat security conference, researchers said Atlas could be led from a bogus newsletter page to send the same message to every contact in the user’s WhatsApp account.

The finding is a warning about browser agents that can act inside logged-in services, not evidence of a WhatsApp breach or account takeover. Zenity’s demonstration required a user to be signed in to WhatsApp Web and to ask Atlas to visit a link. Neither Zenity nor Wired reported malicious use of the technique outside the researchers’ test.

How could Atlas send spam through WhatsApp?

According to Wired’s report on Zenity’s research, the researchers gave Atlas a purported newsletter sign-up link posted on X. The destination page contained hostile instructions telling the agent to open WhatsApp Web and send a newsletter message to every contact.

Zenity calls this an “intent collision.” The user’s instruction, visit a page and sign up for a newsletter, gets combined with instructions planted by the page. Because the browser agent can move across tabs and act in authenticated services, web content stops being merely something the software reads. It can become a competing set of commands.

The researchers said their test evaded safeguards by making the sign-up page look ordinary, writing the instructions in Hebrew, and falsely telling the agent that it was working in a sandboxed WhatsApp environment with fictional contacts. Those are Zenity’s claims about its own demonstration, rather than a confirmed account of a criminal campaign.

WhatsApp was not the vulnerable component in the reported scenario. The agent used a legitimate, active WhatsApp Web session to take actions the account holder had not intended. That differs from conventional account compromise, where an attacker obtains access to the account or attaches another device.

What else did the researchers show?

Zenity reported a similar test against a logged-in Amazon account. Using a malicious newsletter page, it caused Atlas to add a shipping address and a tablet to the shopping cart. The researchers said they could not bypass Atlas protections to make the browser complete the purchase itself. They instead got Atlas to ask Amazon’s Rufus shopping assistant to make the purchase, according to Wired.

That limitation matters. The reported work shows that some agent boundaries could be bypassed, but it does not establish that Atlas could autonomously finish every sensitive action. It also does not support claims that passwords or credentials were stolen in this WhatsApp test.

What protections does OpenAI say Atlas has?

OpenAI has said Atlas includes controls intended to stop agents from running code, downloading files, and using autofill data, Axios reported. The company also says some sensitive tasks require the user to watch the agent’s actions. Users can delete stored browser memories and tell Atlas not to retain information from particular sites, Axios said.

Those controls exist alongside a stubborn problem: OpenAI chief information security officer Dane Stuckey has described prompt injection as largely unsolved across AI platforms, according to Axios. Browser agents add urgency because they are designed to do useful work across the same logged-in tabs that hold messages, shopping carts, and other consequential account actions.

This story draws on original reporting from WIRED.

More Internet/

view all ↗