# What is agent evals?

> Agent evals are systematic checks that score agent quality, safety, and cost on real traces or offline datasets—using LLM-as-a-judge, programmatic rules, or human review.

- HTML: https://openlit.io/glossary/agent-evals
- Markdown: https://openlit.io/glossary/agent-evals.md

Online evals run on production traffic. Offline evals and CI gates catch regressions before release. Both belong in a harness engineering loop: observe failures, score them, fix the harness, verify.

OpenLIT supports LLM-as-a-judge and programmatic evaluators on traces, plus human feedback. Use them to gate releases when accuracy, safety, or cost drifts.


## Related

- [Trajectory evaluation](https://openlit.io/glossary/trajectory-evaluation.md)
- [Agent observability](https://openlit.io/glossary/agent-observability.md)
- [Harness engineering](https://openlit.io/glossary/harness-engineering.md)
- Pillar: https://openlit.io/agent-harness-engineering.md
