Concepts
Glossary
Agent harness glossary
Short definitions for agent harness engineering terms: agent harness, harness engineering, agent observability, agent evals, guardrails, and trajectory evaluation.
Agent harness
An agent harness is everything in an AI agent except the model: the tools, context, prompts, memory, hooks, guardrails, and feedback loops that turn a model into a working agent.
Harness engineering
Harness engineering (or agent harness engineering) is the discipline of designing, measuring, and improving everything around the model in an AI agent so the agent is reliable in production.
Agent observability
Agent observability is the practice of tracing and monitoring the full agent harness: LLM calls, tool calls, MCP requests, retrieval, multi-step trajectories, cost, latency, and outcomes.
Agent evals
Agent evals are systematic checks that score agent quality, safety, and cost on real traces or offline datasets—using LLM-as-a-judge, programmatic rules, or human review.
Guardrails
LLM guardrails are runtime controls that detect or block unsafe prompts and responses—such as prompt injection, sensitive topics, and topic restriction—before they reach users or tools.
Trajectory evaluation
Trajectory evaluation (or tool-call evaluation) scores the sequence of steps an agent took—tools chosen, arguments, ordering, and intermediate decisions—not only the final response.