Researchers expose encrypted prompt-injection flaw that lets attackers hijack xAI’s Grok chat
Adversa AI uncovered an encrypted prompt-injection technique that bypasses Grok’s safety filters and enables malicious actions.
Security researchers at Adversa AI have identified a novel attack called cryptographic context injection that lets an adversary embed encrypted commands on a web page together with the decryption key. Because Grok’s guardrail scanner only inspects plain text, it forwards the ciphertext to the model, which then decrypts it inside its execution sandbox and obeys the hidden instructions, effectively bypassing safety filters.
In a demonstration, the attack extracted a Grok user’s personal details and full chat log by appending them to a URL. The same approach was tried on Google’s Gemini; while Gemini blocked most payloads, it still produced content normally filtered, such as weapon-making instructions. xAI was alerted on June 3 via its HackerOne program and again in early August, but as of August 19 the vulnerability remained exploitable on Grok.com.
SpaceX, the owner of xAI, declined to comment, and Google was not notified because jailbreaks fall outside its bug-bounty scope. The researchers warn that the technique could be extended by splitting encrypted fragments across multiple sources, making future defenses more challenging.
Why it matters
The flaw lets AI chat agents be tricked into revealing user data and bypassing safety controls.
In this story
