Glossary

Trajectory evaluation

Glossary

What is trajectory evaluation?

Trajectory evaluation (or tool-call evaluation) scores the sequence of steps an agent took—tools chosen, arguments, ordering, and intermediate decisions—not only the final response.

Final-answer grading misses many harness bugs: wrong tool selection, unnecessary retries, stuck loops, or skipped retrieval. Trajectory evals inspect the path.

OpenLIT traces give you the raw trajectory (LLM and tool spans). Pair them with LLM-as-a-judge or programmatic checks to score paths and gate CI when trajectories regress.

Related terms

Get started

Ready to use the OpenLIT UI in production?

Self-host or connect your stack in minutes. Same harness UI for traces, dashboards, prompts, and evaluations.