Diamond Evaluations
Navigate to Tests in the sidebar to open Diamond Evaluations. The page has three sections:- Create Evaluation: Configure and launch new Evaluations
- Evaluation Results: Track progress and access completed Evaluations
- Terminated Evaluations: A list of canceled or failed Evaluations
Creating an Evaluation
Start by clicking Create Evaluation on the Diamond Evaluations page.Select Agent
The Agent table shows all registered Agents in your workspace:
Select the Agent you want to evaluate by clicking its row. Only Agents with status Active can be evaluated.
Select Harness
In the Testing Configuration panel, you can choose an Evaluation type:- Trust Score: Standard Evaluation type. Measures Agent trustworthiness across reliability, security, and safety. Configure it under Baseline or Bespoke tabs
- Red Team: Adaptive Evaluation type. Configure it under Adaptive tab
- Reliability: Correctness, consistency, robustness
- Security: Confidentiality, integrity, availability
- Safety: Containment, compliance, transparency
Run Evaluation
Once you have selected an Agent and configured your Harness:- Verify your Agent selection in the left panel
- Confirm configuration settings in the right panel
- Click Run Evaluation
Running a Red Team Campaign
Red Team is designed for deeper adversarial exploration than a standard Trust Score or custom Harness Evaluation. It is useful when:- The Agent handles sensitive data, regulated workflows, or privileged actions
- The Agent uses tools, MCP servers, delegated Agents, or external data stores
- A Trust Score or custom Harness finding needs deeper investigation
- A release needs security, safety, or risk-owner review before deployment
- You want to validate whether previous fixes reduced exploitable behavior
Before You Start
For best results, make sure the selected Agent is Active and has as much context as you can safely provide.Launch Red Team
- Open Create Evaluation in the Tests page in the Console
- Choose the registered Agent you want to test
- Select the Adaptive tab in the Test Configuration panel
- Configure Red Team settings
- Start the campaign
Red Team Settings
Optionally, you can add Personas and Policies to help improve the Evaluation:
If Policies are missing, Red Team can still run, but judgments may rely more heavily on general safety and security expectations.
Advanced Settings
Use advanced settings when you understand the cost and runtime impact of the campaign:Monitoring Progress
Running Evaluations appear in the Evaluation Results table:Evaluation Status
Evaluations typically complete in 5-30 minutes depending on the Harness size and your Agent’s rate limits.
Red Team campaigns can take longer because each wave may run multiple attackers and reflection steps. During a Red Team run, use the campaign panel to track the current wave, phase, elapsed time, accumulated cost, completed attackers, and failed attackers.
Viewing Results
When an Evaluation completes, access the results through the Actions column:- View (eye icon) — Opens the Trust Report in a new tab
- Download (download icon) — Downloads results as a file
- Overall Trust Score with pass/fail status
- Per-dimension breakdown
- Detailed findings for each Probe category
- Deployment recommendations
- Wave details and generated attack seeds
- Attack transcripts and final strategies
- Judge scores and harmful-content judgments
- Leaked artifacts and policy violations
- A final report that clusters vulnerabilities and successful strategies
Evaluation Considerations
Rate Limits
Diamond respects the rate limit you configured during Agent registration. Higher rate limits enable faster Evaluations but may exceed your provider’s quotas. If Evaluations fail with timeout errors:- Verify your Agent URL is accessible
- Check that your API credentials are valid
- Consider reducing the rate limit in Agent settings
Agent Availability
Your Agent must remain available throughout the Evaluation. If your Agent goes offline or becomes unresponsive, the Evaluation may fail or produce incomplete results. For production Agents behind load balancers, ensure sufficient capacity to handle Evaluation traffic alongside normal usage.Red Team Runtime and Cost
Red Team campaigns can generate more traffic than a standard Harness because each wave may launch several attackers and each attacker can run multi-turn conversations. Start with conservative wave, seed, and parallel attacker settings. Increase them only after you have confirmed your Agent’s rate limits and the campaign cost profile.Re-running Evaluations
You can run multiple Evaluations against the same Agent. Each Evaluation creates a new entry in the results table, allowing you to:- Track Trust Score changes over time
- Compare results before and after Agent modifications
- Verify fixes for previously identified issues
Best Practices
Evaluate before deployment: Run a Trust Score Evaluation on every Agent before it reaches production. The results provide evidence of baseline trustworthiness. Test after changes: Any modification to your agent—prompt updates, model changes, tool additions—can affect behavior. Re-evaluate to verify. Use appropriate Harnesses: The Trust Score Harness tests general behaviors. For domain-specific requirements, create custom Harnesses with relevant personas and policies. Use Red Team for deeper security review: Run Red Team after baseline Evaluation, before major releases, and after changes to tools, prompts, policies, or access controls. Give Red Team enough context: Policies and personas improve seed quality and judgment accuracy. Monitor for regressions: Compare Trust Scores across Evaluations. A declining score indicates problems introduced by recent changes.Next Steps
Understand Results
Interpret evaluation findings
Trust Score Harness
Learn about the standard evaluation
Custom Harnesses
Build targeted evaluation Scenarios
Configure Guardrails
Add runtime protection with Dome