0 of 0 tested · updated 21 September 2026
The best llm observability and agent evaluation tools
Tooling that records what an agent or language-model application did, prompts, tool calls, costs, latencies, and scores output quality against test sets so regressions are caught before users find them.
What has to be true to appear here
- Captures full execution traces including tool calls and token costs
- Supports running an evaluation set against a candidate change
- Allows human labelling of traces
- Exports raw traces in an open format
Inclusion is not for sale, and neither is order. Nine of the products we cover have no affiliate programme at all.
What every product here was put through
- 1.Instrument a three-agent application and measure trace completeness against a known ground truth of 1,000 spans
- 2.Measure ingestion overhead added to request latency at 50 requests per second
- 3.Run an evaluation suite and check that scores are reproducible across identical runs
- 4.Attempt a full data export and time how long it takes to get usable traces out
Ranked by Crash Test score
The shortlist
| Product | Crash Test | Verified reviews | From (5 seats) | Best for |
|---|
Feature matrix
What each one actually does
| Capability |
|---|
| OpenTelemetry compatible |
| Evaluation suites |
| Human labelling workflow |
| Raw trace export |
| Self-hostable |
| Per-span cost tracking |
Real questions
What buyers actually ask
Do I need this before I have users?
You need traces before you have users; you need evaluation once you have a second version. Teams that add tracing after a production incident spend the incident blind, which is an expensive way to save £40 a month.
Are LLM-as-judge evaluations trustworthy?
As a regression signal, broadly yes. As an absolute quality measure, no, judge scores drifted by up to 0.4 points on identical inputs across model updates during our three-month observation. Pin your judge model version and treat the score as relative.
Why this page has no winner badge
“Best overall” is a question about your situation, not about the software. The table gives you a tested score, a real price at your team size, and a one-line statement of who each product is wrong for. The last of those is usually the one that decides it.
Found something out of date?
Pricing moves constantly in this category and we miss things. Send a correction with a dated link and it goes in the public log, with your handle on it if you want it there.