Online evals run on production traffic. Offline evals and CI gates catch regressions before release. Both belong in a harness engineering loop: observe failures, score them, fix the harness, verify.
OpenLIT supports LLM-as-a-judge and programmatic evaluators on traces, plus human feedback. Use them to gate releases when accuracy, safety, or cost drifts.