/
© 2026 RiffOn. All rights reserved.

Get your free personalized podcast brief

We scan new podcasts and send you the top 5 insights daily.

  1. Latent Space: The AI Engineer Podcast
  2. Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan
Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast · Jun 22, 2026

Beyond cyber: Gray Swan's founders on red-teaming AI agents, tackling prompt injection, and building a new layer of AI-specific security.

Gray Swan's "Shade" AI Now Outperforms Humans in Vulnerability Discovery

In recent competitions, Gray Swan's automated red teaming system, called "Shade," has become more effective than human experts at breaking models within a given timeframe. This signals a turning point where specialized AI is becoming the primary tool for finding security flaws in other AIs.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

AI Model Robustness Does Not Improve With Scale; It Requires Explicit Adversarial Training

Making a model bigger doesn't automatically make it more secure against jailbreaks. Robustness is not an emergent property of scale and must be explicitly trained for using adversarial data. This is why specialized guardrail models can outperform larger, general-purpose models on security tasks.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

Frontier AI Models Fail at Red Teaming Because Their Safety Training Prevents Attack Generation

Using a powerful frontier model for automated red teaming is ineffective. Its built-in safety mechanisms cause it to refuse to generate the jailbreaks or attacks it's tasked with creating. Effective automated red teaming requires models specifically trained for adversarial purposes, often without the same safeguards.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

AI and Humans Possess Orthogonal Vulnerabilities, Not Superior or Inferior Security

An AI might resist a sophisticated attack but fall for a simple trick a human never would (e.g., an email saying "this is a simulation"). This shows AI vulnerabilities are not a subset or superset of human ones, but occupy a different dimension entirely. Direct robustness comparisons can be misleading.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

Simon Willison's "Lethal Trifecta" Defines Real-World Prompt Injection Risk

Prompt injection risk requires three conditions: the agent must ingest untrusted external data, have access to sensitive internal information, and possess the ability to send that information elsewhere (exfiltration). An agent lacking any of these components poses a significantly lower risk, providing a clear framework for mitigation.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

An Independent AI Security Ecosystem Is Emerging, Mirroring Past Tech Platform Shifts

Just as new platforms like operating systems and cloud computing spurred independent security companies, AI is creating a need for third-party safety providers. Even with strong in-house efforts at major labs, there is a distinct market demand for specialized, external security services like those from Gray Swan.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

Agent Identity Management Will Evolve Around User Personas Before Per-App Permissions

Managing permissions for AI agents is a huge challenge. The most likely near-term solution is not granular, per-app controls, which create overwhelming cognitive load. Instead, agent identity will be managed through distinct user personas, like a "work agent" for professional tasks and a "home agent" for personal ones.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

Coding Agents Can Transform AI Interpretability into a Rigorous, Automated Science

Mechanistic interpretability (Mekinterp) research has been slow due to its manual, ad-hoc nature. The guests argue that coding agents can automate the experimentation process, enabling large-scale, systematic analysis of AI models. The first science AI should automate is the science of understanding itself.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago

AI Coding Agents Could Make Impractical, Formally Verified Code a Mainstream Reality

Writing formally verified code, which can be mathematically proven to be secure, has been a niche practice due to its extreme difficulty for humans. Because AI agents don't get bored or frustrated, they could be tasked with writing code in these secure languages, making high-assurance programming practical for the first time.

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan thumbnail

Red-Teaming after Mythos — Zico Kolter & Matt Fredrikson, Gray Swan

Latent Space: The AI Engineer Podcast·2 months ago