Skip to main content
    All AI News
    External (via citation)Tuesday, September 29, 2026 3 min read
    AI

    Self-replicating prompt injections exist

    OpenAI's internal red-teaming framework (GPT-Red) discovered self-replicating prompt injections capable of spreading like computer worms, discovered June 27, 2026 and disclosed September 25, 2026. The attacks were demonstrated in simulat…

    Key takeaways
    • 01The attacks were demonstrated in simulated environments only, with no observed real-world impact.
    • 02In the clearest example, a malicious email instructed an AI agent to append the full injection payload verbatim to every reply it sent, causing the attack to propagate automatically.
    • 03More sophisticated variants replicate via filesystems, code comments, or multi-hop message chains, and can trigger destructive actions such as deleting files while simultaneously copying the attack payload to persist.
    • 04OpenAI is disclosing the findings due to the novel attack class, not in response to any incident.
    In brief · from alignment.openai.com

    OpenAI's internal red-teaming framework (GPT-Red) discovered self-replicating prompt injections capable of spreading like computer worms, discovered June 27, 2026 and disclosed September 25, 2026. The attacks were demonstrated in simulated environments only, with no observed real-world impact. In the clearest example, a malicious email instructed an AI agent to append the full injection payload verbatim to every reply it sent, causing the attack to propagate automatically.

    Read the full article at alignment.openai.com

    OpenAI's internal red-teaming framework (GPT-Red) discovered self-replicating prompt injections capable of spreading like computer worms, discovered June 27, 2026 and disclosed September 25, 2026. The attacks were demonstrated in simulated environments only, with no observed real-world impact. In the clearest example, a malicious email instructed an AI agent to append the full injection payload verbatim to every reply it sent, causing the attack to propagate automatically. More sophisticated variants replicate via filesystems, code comments, or multi-hop message chains, and can trigger destructive actions such as deleting files while simultaneously copying the attack payload to persist. OpenAI is disclosing the findings due to the novel attack class, not in response to any incident.

    Don't miss tomorrow's

    The Daily Pulse in your inbox each morning — sourced and linked.

    How often
    Keep going — across the app