Skip to main content
Diamond supports two evaluation types. A behavioral evaluation uses Harnesses to send Probes and produce a Trust Score. A Red Team evaluation runs adaptive attack waves against the Agent and produces findings and a report. Both use an evaluation ID.

Run an Evaluation

Use vijil evaluate <agent> with an Agent ID or alias. The default type is behavioral. Choose either the standard Harness set or a custom Harness:
Set --type redteam to start a Red Team evaluation. Pass an optional JSON object with --redteam to override wave settings:
The command waits for a terminal status and prints the result by default. Add --no-wait to create a cloud evaluation and return its ID immediately:
Red Team does not use --baseline, --bespoke, --sample-size, --local, or --harness-name.

Inspect Results

List evaluations for an Agent, or show a behavioral evaluation by its ID:
vijil scores show returns the Agent’s latest Trust Score. For Red Team progress and reports, use the evaluation-scoped commands:
See the Red Team CLI guide for seeds, attacks, judgments, reflections, and report formats. To request JSON output from a CLI command, use the global option before the command, for example vijil --output json evaluations list.

Manage Behavioral Harnesses

Harnesses apply to behavioral evaluations. The current CLI provides these resource commands: Create a custom Harness, then pass its ID to --bespoke:
vijil harnesses create also accepts --description, --system-prompt, --system-prompt-file, --persona-ids, and --policy-ids. Separate Persona or Policy IDs with commas.
Last modified on September 30, 2026