Agent Evaluation Lab
See how a measured agent run is evaluated on outcome and trajectory separately — task success is not the same as a good tool path.
Example evaluation
Select a measured Cookbook case, then Run to load its recorded trace and computed evaluation.
Task: Check the payments service. If it is degraded, inspect the relevant documentation and summarize what the user should know.
Related Labs
Tool calling and the agent loop produce traces; this Lab evaluates those traces. Planning is the next runtime lesson.
02 Agent Loop03 Agent Evaluation04 Planning
Continue Learning
Guides that explain outcome, trajectory, and evaluation.
Implementation
Run the measured Cookbook example locally.
Related Tools
Tools commonly used with agent evaluation stacks.
Related Rankings
See how this stack compares in DataAIHub rankings.
Lab ID agent-evaluation · guided / precomputed