# The best llm observability and agent evaluation tools

> Tooling that records what an agent or language-model application did, prompts, tool calls, costs, latencies, and scores output quality against test sets so regressions are caught before users find them.

Last reviewed 2026-09-21. Canonical page: https://launch500.com/best/agent-observability

## What has to be true to appear here

- Captures full execution traces including tool calls and token costs
- Supports running an evaluation set against a candidate change
- Allows human labelling of traces
- Exports raw traces in an open format

## What every product here was put through

1. Instrument a three-agent application and measure trace completeness against a known ground truth of 1,000 spans
2. Measure ingestion overhead added to request latency at 50 requests per second
3. Run an evaluation suite and check that scores are reproducible across identical runs
4. Attempt a full data export and time how long it takes to get usable traces out

## The shortlist

Ranked by Crash Test score. Untested products appear below tested ones regardless of community rating.

| Product | Crash Test | Verified reviews | From (5 seats) | Best for |
| ------- | ---------- | ---------------- | -------------- | -------- |

## Verdicts

## Frequently asked questions

### Do I need this before I have users?

You need traces before you have users; you need evaluation once you have a second version. Teams that add tracing after a production incident spend the incident blind, which is an expensive way to save £40 a month.

### Are LLM-as-judge evaluations trustworthy?

As a regression signal, broadly yes. As an absolute quality measure, no, judge scores drifted by up to 0.4 points on identical inputs across model updates during our three-month observation. Pin your judge model version and treat the score as relative.

---

Launch500 independently tests SaaS and business software. Inclusion and order are not for sale: https://launch500.com/trust/how-we-make-money