Multi-Turn Conversation Evaluation
Review an assistant conversation one turn at a time, score each turn against a rubric, flag issues, highlight evidence spans, and give a conversation-level verdict.
- Conversation
- JSON
An agent run is not one answer to grade. It is reasoning, tool calls, subagents, artifacts, and errors, and the failure is usually somewhere in the middle. This interface lays the whole trace out so a reviewer can find it.
Steps render in order with their arguments, results, tokens, and latency, and a minimap sized by latency, tokens, or cost shows where the run spent its budget. Reviewers rate each step, set a verdict, score a configurable rubric (helpfulness, faithfulness, efficiency, tool use, and instruction following by default), tag failure modes like hallucination or wrong tool, and write a critique.
Each section saves independently, so partial reviews are fine. Verdicts, step ratings, failure modes, and rubric items are all parameters.