If you're not using AI to attack your own systems, your adversaries will
AI agents are now attack infrastructure—defenders who skip agentic red-teaming are already being tested by someone else.
AI agents are now attack infrastructure—defenders who skip agentic red-teaming are already being tested by someone else.
Autonomous AI agents have compressed exploit timelines from days to seconds, while simultaneously creating unmanaged non-human identities that sidestep static policies. Former CISA acting cyber chief Matt Hartman argues every agent must be treated as a privileged identity. On the offensive side, Armadin's swarm ran 17 million attack actions across three days—work a five-person team would need four months to replicate. Continuous, AI-native red-teaming is no longer optional; the question is only who receives the findings.
Action: Audit all AI agent identities in your environment this quarter—apply privileged-access controls and behavioral monitoring before adversaries map them for you.
AI agents excel at hacking organizations, as they’ve demonstrated in real-life attacks multiple times over the past few weeks. They also expose a whole new attack surface for organizations trying to protect against both human and AI intrusions. As if defenders needed more worries to keep them up at night, agents introduce new data-integration channels that attackers can abuse. They also introduce a new type – and ever growing number – of non-human identities that are difficult to manage and can bypass traditional, static security policies. “There is tremendous risk associated with agentic AI and machine identities,” Matt Hartman, former acting head of cyber of the US Cybersecurity and Infrastructure Security Agency (CISA), told The Register. “As AI moves from generating content – yesterday's use case – to taking actions, it is inevitable that agents are going to receive access to sensitive systems and sensitive data,” Hartman said. “One area where organizations are struggling today is that they're going to need to treat every agent as a privileged identity.” Enterprises also face agentic threats from outside their organization, he added. “AI-enabled or AI-amplified identity and social engineering attacks are increasing significantly by the minute,” Hartman said. “We're seeing very highly personalized phishing, very good impersonation, automated reconnaissance. That really makes traditional indicators of trust increasingly unreliable.” For defenders, this means a “continued focus on strong identity, on phishing-resistant authentication, on behavioral signals, and on zero-trust principles therein,” he added. “Nothing deeply new here - but it is a whole new attack surface.” Meanwhile, on the attackers’ side, agents don’t take time off, and they remain singularly focused on completing a task, whether that’s finding vulnerabilities and exploit chains or mapping networks and identifying sensitive files. All of this makes these near-autonomous attack bots a gift from the heavens for financially motivated criminals and government-backed cyber operatives. It also presents a security use case for defenders: agentic red teaming. As former NSA cyber boss Rob Joyce said during a talk at RSAC: if you aren’t using AI agents to attack your own organizations, you can bet that someone else is. “You are going to be red-teamed whether you pay for it or not,” Joyce said. “The only difference is, you know who gets the results delivered to them.” Hartman echoed Joyce’s words. “What we are seeing as the leading capabilities to help defenders – there is a burgeoning market for continuous, AI-native, AI-enabled, automated red teaming and pen-testing,” he told us. After spending nearly two decades in the federal government at CISA, Hartman joined Merlin Group in October as its chief strategy officer. In his new private-sector role, he helps determine which early- to growth-stage cybersecurity and emerging technology companies the group invests in, and then works with these firms to navigate government, critical infrastructure, and other highly regulated markets. The goal is to integrate and scale “promising technologies” into critical environments, Hartman said. Right now, most of these technologies use AI agents to fight AI agents. “Organizations are just inundated with vulnerabilities, and adversaries are able to leverage AI to find vulnerabilities and exploit them in seconds when it used to take days,” he said. Agentic red teaming “is a category of products that every organization, including federal agencies, absolutely needs in the near term just to keep pace.” 'Largest controlled live AI cyberattack on record' Mandiant founder and former CEO Kevin Mandia has a new company, Armadin, which launched in March with a startling $190 million in seed and Series A funding. The firm builds and trains autonomous attacker swarms – thousands of AI agents that run 24/7 in organizations’ infrastructure to simulate real-life attackers. Ahead of Black Hat earlier this month, the startup said it and Tenex.ai, an agentic security operations provider, executed what they called the “largest controlled live AI cyberattack on record” for an unnamed “leading” global institution. Over the three-day attack, Armadin's swarm generated 17 million offensive actions, discovered 38 validated attack paths, and produced 238 security findings. Tenex.ai's agentic platform separately triaged 100 percent of 101,169 alerts and reconstructed the entire attack across 231 billion raw events. This exercise, we’re told, would have taken a five-person analyst team about 2,400 hours – or four months – to pull off. Co-founder and Chief Offensive Security Officer Evan Peña was the global red-team lead at Mandiant before co-founding Armadin. At Mandiant, he led a 210-person team whose members spanned the globe. “The problem was it was 100 percent human-led security assessments, and that would generally limit the amount of time that we would have,” Peña told The Register. His red team “would do a couple weeks or a one-month engagement, and then we would report on the engagement, give them a PDF file, walk away, and they would hire us again in a year. In today’s age of AI, it’s very archaic to think about that when we can scale so significantly with AI.” Attack yourself before someone else does At Armadin, Peña leads the human team that trains the AI agents. One of the lessons learned from OpenAI’s models autonomously attacking Hugging Face, according to Peña, is that organizations need to perform safe offensive AI attacks against their own systems. "Safe" is the keyword here: remember OpenAI’s rogue models intentionally didn’t have any guardrails in place. Yes, his statement is self-serving as it's core to Armadin's business. But he’s not wrong. “Organizations can cover so much more attack surface because we are able to leverage these agents at scale, and we have three things that we didn’t have before,” he said. “We have more time, because agents don’t sleep and they don’t take holidays. There’s no workforce requirements for them.” Number two, he said, is expertise. Attack agents need pre-training before they are set loose on organizations’ infrastructure. They need to know how to code, and perform source-code review. They need to know how to do application security, how to spot network misconfigurations, and hack into different systems and networks. “And then you add post-training to that from human expertise,” Peña said. “Number three is coverage,” he said. “We were only able to cover a finite amount of attack surface in the past. So if you had 10,000 external systems with a limited amount of time and humans, you could maybe cover 2,000 or 1,000 of those within that particular period of time. Now we can cover all 10,000 in probably hours.” Armadin’s AI agents have broken into every single customer’s environment, according to Peña. “We have found over 50 zero-days, and by zero-days, I don't just mean this zero-day allowed you to deface a web page. That’s cool, but I want to break into your network from the internet,” he said. “The zero-days I'm referring to allow an attacker to get remote code execution on an actual system. They're very high-impact zero-days. We don't care about noise, we care about impact.” Quarterly pen-testing doesn't cut it anymore The biggest challenge these days for defenders is the scale and speed AI brings to previously manual attackers’ dirty work – like scoping potential victims, performing reconnaissance, identifying vulnerable systems and exploits, and reading logs. Now all of these tasks can be automated. Penetration testing needs to keep up, Jay Bavisi, founder and group president of EC-Council, told The Register. The largest and best organizations do pen-testing once a year to meet compliance requirements, and “the better ones” run these exercises quarterly, Bavisi said. This is largely because human-led pen-tests take about three months. “So you have a serious problem with speed,” he
- 01Autonomous AI agents have compressed exploit timelines from days to seconds, while simultaneously creating unmanaged non-human identities that sidestep static policies.
- 02Former CISA acting cyber chief Matt Hartman argues every agent must be treated as a privileged identity.
- 03On the offensive side, Armadin's swarm ran 17 million attack actions across three days—work a five-person team would need four months to replicate.
- 04Continuous, AI-native red-teaming is no longer optional; the question is only who receives the findings.
Don't miss tomorrow's
The Daily Pulse in your inbox each morning — sourced and linked.