Observability & Monitoring
Tracing with OpenTelemetry, token-cost FinOps, production metrics, and the tooling that lets you see what your agents are actually doing.
Articles
Agent observability: the five signals and the alert rules that watch them
What to measure on a production AI agent — outcome rate, cost per run, loop depth, tool errors and liveness — with copyable alert rules, the OpenTelemetry GenAI conventions that standardise the data, and two failures from operating a real portfolio.
Read →Agent cost control: FinOps for autonomous AI
Why cost is a first-class signal for agents, what actually drives the bill, and how token tracking, loop detection and budgets keep an autonomous system affordable.
Read →The production metrics that tell you an agent is working
Six metrics — completion, accuracy, hallucination, latency, cost and satisfaction — that turn an agent from a black box into a system you can actually run.
Read →Tracing AI agents with OpenTelemetry: spans, attributes and working code
How to instrument an agent with the OpenTelemetry GenAI conventions — invoke_agent and execute_tool spans, gen_ai.* attributes, a Python example you can run — and why a trace with a correlation id once saved this portfolio from rotating a healthy credential.
Read →Ship agents you can trust
Get the AgentOps Production-Readiness Checklist — free, vendor-neutral, practical.
Get the checklist