Skip to content
Preview. Launch500 opens to the public soon; payments here are in test mode and no money moves.

Run by 99 Developers Ltd. Listings and Crash Tests pay for it. Money never buys a score or a place on the list. What money cannot buy

Categories

0 of 0 tested · updated 21 September 2026

The best llm observability and agent evaluation tools

Tooling that records what an agent or language-model application did, prompts, tool calls, costs, latencies, and scores output quality against test sets so regressions are caught before users find them.

What has to be true to appear here

  • Captures full execution traces including tool calls and token costs
  • Supports running an evaluation set against a candidate change
  • Allows human labelling of traces
  • Exports raw traces in an open format

Inclusion is not for sale, and neither is order. Nine of the products we cover have no affiliate programme at all.

What every product here was put through

  1. 1.Instrument a three-agent application and measure trace completeness against a known ground truth of 1,000 spans
  2. 2.Measure ingestion overhead added to request latency at 50 requests per second
  3. 3.Run an evaluation suite and check that scores are reproducible across identical runs
  4. 4.Attempt a full data export and time how long it takes to get usable traces out

Ranked by Crash Test score

The shortlist

Tested products come first, ordered by score. Untested products appear below them regardless of their community rating, because we have not looked at them. Nothing on this page can read a vendor’s subscription tier.
ProductCrash TestVerified reviewsFrom (5 seats)Best for

Feature matrix

What each one actually does

Confirmed in testing. Where the honest answer is a sentence rather than a tick, we write the sentence. A safety-relevant feature that is off by default may not be recorded as a plain yes.
Capability
OpenTelemetry compatible
Evaluation suites
Human labelling workflow
Raw trace export
Self-hostable
Per-span cost tracking

Real questions

What buyers actually ask

Sourced from launch threads, review text and search data, not invented to fill a schema block. Each answer is two to four sentences and says a number where we have one.
Do I need this before I have users?

You need traces before you have users; you need evaluation once you have a second version. Teams that add tracing after a production incident spend the incident blind, which is an expensive way to save £40 a month.

Are LLM-as-judge evaluations trustworthy?

As a regression signal, broadly yes. As an absolute quality measure, no, judge scores drifted by up to 0.4 points on identical inputs across model updates during our three-month observation. Pin your judge model version and treat the score as relative.

Why this page has no winner badge

“Best overall” is a question about your situation, not about the software. The table gives you a tested score, a real price at your team size, and a one-line statement of who each product is wrong for. The last of those is usually the one that decides it.

Found something out of date?

Pricing moves constantly in this category and we miss things. Send a correction with a dated link and it goes in the public log, with your handle on it if you want it there.

Corrections log and policy