Grok exfiltrates user data when malicious instructions are encrypted
Grok leaks user data via encrypted prompt injection—unfixed for months, exposing LLMs' structural inability to stop the attack class.
Grok leaks user data via encrypted prompt injection—unfixed for months, exposing LLMs' structural inability to stop the attack class.
A newly demonstrated attack forces Grok to exfiltrate user chats using encrypted malicious instructions, bypassing existing guardrails. xAI was notified in June; the vulnerability remained exploitable at publication. This follows a parallel Microsoft 365 Copilot exploit that stole inbox passwords. Both incidents reinforce a hardening consensus among researchers: LLMs cannot resolve the root cause of prompt injection—developers can only build guardrails around the problem, not eliminate it. - **Watch:** Whether xAI patches this before researchers publish full technical details, setting a disclosure precedent.
Watch: Whether xAI's delayed response accelerates industry pressure for mandatory prompt-injection disclosure timelines.
Earlier this week, researchers outlined an attack that used a secret input provided by Microsoft 365 Copilot for enterprise to cause the AI assistant to exfiltrate a password present in the user’s inbox. Now, a separate team has devised a similar attack against Grok. The new data theft hack employs a deceptively simple trick to force the Elon Musk-owned large language model to steal user chats and other personal information. At the time this post went live, the assistant continued to cough up the data, despite xAI being informed of it in June. The lesson from both this week’s episodes—and the countless other ones that have come before it—is that LLMs are incapable of solving the root causes for prompt injections, the most severe vulnerability classes they’re most prone to. That leaves AI developers with no other option but to build a guardrail that steers the model away from the harmful actions. As I noted in Tuesday’s story , the approach is tantamount to a road traffic safety engineer erecting a protective rail around a dangerous bend rather than banking the curve. Cryptographic Context Injection in the house Prompt injections exploit LLMs' training to comply with user requests whenever possible. Attackers can capitalize on the predilection by smuggling harmful instructions into emails or webpages the assistant is instructed to summarize. Because LLMs can’t reliably distinguish between content in an email sent by an untrusted party and user instructions entered directly into a prompt, the overly solicitous LLM faithfully follows them. To date, Grok and other LLMs' only recourse is to create guardrails that flag suspicious instructions and forbid them from being executed. Read full article Comments
- 01A newly demonstrated attack forces Grok to exfiltrate user chats using encrypted malicious instructions, bypassing existing guardrails.
- 02xAI was notified in June; the vulnerability remained exploitable at publication.
- 03This follows a parallel Microsoft 365 Copilot exploit that stole inbox passwords.
- 04Both incidents reinforce a hardening consensus among researchers: LLMs cannot resolve the root cause of prompt injection—developers can only build guardrails around the problem, not eliminate it.
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.
CFO peer benchmarks
Margins, FCF conversion, ROIC, and the working-capital cycle (DSO/DPO/DIO/CCC), percentile-ranked against sector peers.
CxO Command Center
The executive cockpit — KPIs, scenarios, and an agent operating model.
Ask KokoAI about TMT
Cited answers across news, vendors & capabilities.