- Live metrics: status, phase, elapsed time, and attacker progress
- Waves: per-wave seeds and attackers, drilling into transcripts and judgments
- Final report: aggregated findings across the whole campaign
Live Metrics
Six metrics sit at the top of the run page.
Run status reports one of
pending, starting, running, analyzing, wave_complete, completed, succeeded, failed, cancelled, or errored. Note that this vocabulary is wider than the PENDING, RUNNING, COMPLETED, FAILED set used for standard Evaluations, because a campaign reports the wave loop as it progresses.
Cancel run stops an active campaign. Copy ID copies the run identifier, which is useful when correlating with logs or support requests.
Waves
Each wave is listed with a set of stat pills showing Seeds, Running, Done, and Errors. Click a wave to inspect what the campaign generated and what its attackers did. Review wave details to see whether later waves became more focused based on earlier discoveries. Strong campaigns explore broadly early, then narrow once the campaign identifies a promising strategy.Attacker Details
Click an attacker to inspect the evidence behind one seed. Use Previous seed and Next seed to move through a wave without returning to the list.
Before an attacker finishes, these panels report
Strategy metadata appears when the attacker completes. and Judgment not available yet.
Judgments
A judge reads the whole transcript rather than scoring each response in isolation, because multi-turn attacks often look harmless turn by turn and only succeed in aggregate. Harmful content delivered is the field to read first:
Claim and Evidence show what the judge concluded and the transcript excerpt it relied on, so you can verify a finding rather than take it on trust. Leaked artifacts lists internal details the Agent disclosed, such as system prompt fragments, tool names, private endpoints, credentials, or operational procedures.
Attaching Policies changes what the judge can conclude. With Policies available it can identify a concrete policy violation. Without them, judgments fall back on general safety and security expectations.
Final Report
The run page shows a Final report summary once the campaign completes. Before then it reportsReport is available after the run completes. Click Open full report view for the detail behind the summary.
Each finding carries a Category, a Description, Evidence, and a link to the Transcript that uncovered it. Use the report to decide which issues need product changes, prompt or policy updates, tool permission changes, or Dome Guardrails.
Downloading a Report
The download control on the report offers three formats:
Downloads are unavailable until the report is generated, reporting
Report is not available yet.
Prioritize Findings
Use the harmful-content judgment to decide what to fix first: Address immediately: findings with FULL harmful-content judgments, confirmed policy violations, and leaked artifacts. These are demonstrated failures with a transcript attached. Address in the next release: findings with PARTIAL harmful-content judgments. The Agent produced some harmful content, so the weakness is real, but the attack did not fully succeed. Review and monitor: findings whose Evidence does not clearly establish harm. Check these against the actual Agent design, policies, and data access before treating them as defects.Next Steps
Run Another Campaign
Widen coverage or verify a fix
How Engagements Work
The wave loop behind a campaign
Configure Guardrails
Turn findings into runtime protection
Trust Score Results
Interpret standard Evaluation findings