Skip to main content
Open the printed URL in CLI, or open Runs in your project UI and select the run.
  1. Check execution progress. Each test and persona combination is one simulation. A simulation can complete, fail to execute, or be canceled.
  2. Select a simulation and open Results. Wait for its selected graders to finish, then read each grader’s score and individual result.
  3. For Expected behaviors, read the statement results. Follow the evidence into Transcript to see what the caller, agent, and tools actually did.
  4. Use the transcript and recording, when available, to decide whether to change the agent or the test. Save the change and start another run.
Execution and grading finish separately. A completed run can still have grading in progress, and completed execution can include simulations that failed to execute. Check individual simulation and grader results.