The industry consensus is that reliability lives in the harness, not only in the model. Teams run agents on real tasks, observe failures in traces, classify them, fix prompts, tools, rules, or guardrails, and verify the fix with regression evals.
The core loop is run → observe → evaluate → fix the harness → verify. Every repeated agent failure becomes a harness defect to ratchet closed.
OpenLIT is an open-source agent harness engineering platform for that loop: agent observability and OpenTelemetry tracing, agent evals, guardrails, prompt management, and cost and GPU monitoring.