Glossary

Agent evals

Glossary

What is agent evals?

Agent evals are systematic checks that score agent quality, safety, and cost on real traces or offline datasets—using LLM-as-a-judge, programmatic rules, or human review.

Online evals run on production traffic. Offline evals and CI gates catch regressions before release. Both belong in a harness engineering loop: observe failures, score them, fix the harness, verify.

OpenLIT supports LLM-as-a-judge and programmatic evaluators on traces, plus human feedback. Use them to gate releases when accuracy, safety, or cost drifts.

Related terms

Get started

Ready to use the OpenLIT UI in production?

Self-host or connect your stack in minutes. Same harness UI for traces, dashboards, prompts, and evaluations.