In a recent discussion, Gray Swan co-founders Zico Kolter and Matt Fredrikson highlighted the pressing risks of AI security, especially in light of new US export controls on the Mythos model, which have intensified attention on indirect prompt injection vulnerabilities. They emphasized that AI security differs significantly from traditional cybersecurity, as AI systems present unique exploit classes, particularly with coding and computer-use agents. Their automated red-teaming tool, Shade, now surpasses human capabilities in exposing these flaws. Kolter and Fredrikson warned that the first major AI security breach could manifest as a “gray swan event,” meaning it is foreseeable yet still underappreciated, underscoring the critical need for heightened scrutiny and specialized defenses in AI deployments.
Shade: Shade is Gray Swan’s automated red-teaming system designed to evaluate AI model robustness against attacks such as prompt injection. It has been used by organizations like Anthropic to test coding agents and can outperform human red teamers in specific tasks. Shade is directly featured in the news as a key innovation for identifying vulnerabilities in frontier models.
Cygnal: Cygnal is Gray Swan’s guardrail model that enforces custom policies on AI agents by filtering inputs and tool calls. It sits between users, models, and actions to detect violations like data exfiltration risks. The product is discussed in the news as part of Gray Swan’s defense toolkit alongside red-teaming efforts.
Gray Swan: Gray Swan is an AI security company focused on red-teaming and safeguarding AI systems. Its cofounders, Zico Kolter and Matt Fredrikson, developed tools to evaluate and mitigate vulnerabilities in frontier models and agents. The company’s work is directly relevant to the news as it addresses the rising risks of prompt injection and automated adversarial testing for systems like those used in coding agents.
Zico Kolter: Zico Kolter is a Carnegie Mellon University professor and member of OpenAI’s board on the Safety & Security Committee. He co-founded Gray Swan and co-authored key research on indirect prompt injections. His expertise is central to the news discussion on why AI security requires a distinct approach from traditional cybersecurity and the development of tools like Shade.
Matt Fredrikson: Matt Fredrikson is a Carnegie Mellon University professor and CEO of Gray Swan. He co-founded the company with Zico Kolter after years of research on vulnerabilities in deep learning systems. His role is highlighted in the news through explanations of Gray Swan’s automated red-teaming capabilities and enterprise guardrails.
US Export Controls: The US Government recently issued an export control directive on the Mythos model, heightening industry focus on indirect prompt injection risks.
Agent Vulnerabilities: Coding and computer-use agents create new exploit classes through prompt injection that can lead to data leakage or unauthorized actions when processing untrusted content.
AI Security Distinction: AI systems introduce vulnerabilities that differ fundamentally from traditional software, requiring specialized red-teaming and guardrails rather than standard cybersecurity approaches.
