Skip to main content
TL;DR: Red Team attacks your Agent to find weaknesses a fixed Harness would miss. It creates an Engagement that runs in waves, learning from each attack to plan the next. The result is campaign evidence, not a Trust Score.
A standard Evaluation answers a closed question: how does this Agent score against tests you already chose? That is the right question for release evidence, because the answer is reproducible and comparable over time. Red Team answers an open question instead: what can an attacker actually get this Agent to do? You cannot enumerate that in advance, so Red Team explores rather than measures.

Red Team Versus Standard Evaluations

Red Team does not replace Trust Score Evaluations. Use a Trust Score for reproducible readiness evidence, then use Red Team to search for harder-to-find vulnerabilities and successful attack strategies.

When To Choose Red Team

Choose Red Team when the Agent handles sensitive data, regulated workflows, or privileged actions, when it uses tools, MCP servers, delegated Agents, or external data stores, or when a release needs security and risk-owner review. Red Team is also the right choice when a Trust Score finding needs deeper investigation, because an Engagement can pursue a weakness across many turns instead of recording a single failed Probe.

Next Steps

Engagements

How waves, attacks, and reflections work

Run a Campaign

Launch Red Team from the Console

Understand Results

Read waves, judgments, and the final report

Red Team for Developers

Reach Red Team from the CLI, MCP, and REST API
Last modified on September 24, 2026