Final-answer grading misses many harness bugs: wrong tool selection, unnecessary retries, stuck loops, or skipped retrieval. Trajectory evals inspect the path.
OpenLIT traces give you the raw trajectory (LLM and tool spans). Pair them with LLM-as-a-judge or programmatic checks to score paths and gate CI when trajectories regress.