Skip to content
Preview. Launch500 opens to the public soon; payments here are in test mode and no money moves.

Run by 99 Developers Ltd. Listings and Crash Tests pay for it. Money never buys a score or a place on the list. What money cannot buy

Method

Two scores. Never blended. All the arithmetic.

One convenient number would be easier to read and would destroy your ability to tell an opinion from a measurement, so we publish two and keep them apart. This page has every weight, every threshold and the exact curve, enough to recompute any score on this site by hand and catch us if it does not come out.

Crash Test score · 0 to 100

Our own hands-on result. One team, one week, one named version, one date. It answers “does this thing work when we push it?” It cannot tell you what year three looks like.

Community score · 0 to 5

Verified reviews from people paying the bill. Many teams over many months. It catches repricing, support decay and month-nine regret. The things a bench never sees, and never the same question as the test.

Crash Test

Per-category weights, fixed before anything is tested

Weights differ by category because what good looks like differs, latency decides a voice agent and is irrelevant to document extraction. Every rubric includes a 'handles failure' dimension worth at least 20 percent, because products are sold on their best case and lived with on their worst.
Import fidelity
0.26
25,000 messy contacts with duplicates and mixed formats, scored field-by-field against a known-good set.
Permission integrity
0.22
A non-admin attempts to move, view and export records they should not reach, by every route we could find.
Performance at scale
0.18
List-view and report load times at 250,000 records, cold and warm.
Export and exit cost
0.16
Full export of every object; scored on completeness, format and elapsed time.
Cost at 25 seats
0.18
Billed cost on the cheapest tier that actually carries the permissions a sales team needs.
Collision and assignment
0.28
2,000 conversations across a four-person rota; scored on double-replies, misassignment and lost threads.
Handles failure honestly
0.22
Email connection broken mid-shift; scored on what the team is told and what happens to inbound mail.
Reporting accuracy
0.20
The product's own resolution-time figures reconciled against our timestamps on the same conversations.
Data deletion
0.14
A deletion request verified across search, exports and reporting.
Cost per agent
0.16
Billed cost at a four-agent team on the tier that includes reporting.
Reconciliation accuracy
0.28
1,200 transactions against 340 invoices; auto-match rate scored against a hand-reconciled ledger.
False matches
0.24
Counts wrong matches the product made confidently. A wrong match scores worse than no match.
Bank feed recovery
0.18
72-hour feed outage; scored on recovery and whether it double-imports on reconnection.
Audit trail
0.14
Issue, part-pay, credit-note and void an invoice; the trail must survive all four.
Cost including multi-currency
0.16
Tier needed for the tested feature set, plus any FX margin we could establish.
Dependency handling
0.26
400-task programme across four teams; one date changed, scored on what correctly cascades.
Guest permission integrity
0.22
A restricted contractor attempts access by search, by link and via notifications.
Performance at scale
0.20
Cold load of a 2,000-item board on a mid-range laptop.
Export completeness
0.14
Workspace export scored on whether comments, attachments and history survive.
Cost including guests
0.18
12 staff plus 6 contractors, billed the way the product actually bills them.
Leave arithmetic
0.26
Year-end rollover with carry-over caps and part-time pro-rating, checked by hand.
Permission integrity
0.26
A line manager attempts to reach another employee's compensation by report, export and search.
Multi-country onboarding
0.18
50 employees, three countries, mixed contract types; scored on manual correction needed.
Deletion and retention
0.16
A data-deletion request exercised end to end, including what survives in backups and reports.
Cost per employee
0.14
Billed cost at 50 employees on the tier carrying the tested feature set.
Count accuracy
0.28
5 million known events with duplicates and late arrivals, scored against ground truth.
Identity resolution
0.22
Anonymous session, then signup, then login on a second device; scored on correct stitching.
Answerable by a non-engineer
0.20
Time for a non-engineer to build a retention chart unaided from a written question.
Raw data access
0.14
Export or warehouse connection the customer controls; scored on completeness and latency.
Cost at real event volume
0.16
Modelled including staging and debug traffic, which most vendors bill for.
Reliability at volume
0.28
10,000 records through a 5-step flow; scored on completed runs, silent drops and duplicate side effects.
Handles failure honestly
0.22
Three induced failure classes; scored on whether the tool retried correctly and reported accurately.
Time to first working flow
0.18
Stopwatch from empty account to a passing 5-step flow, by a tester new to the product.
Handover to a non-specialist
0.14
A non-technical teammate changes one step unaided; scored on completion and time.
Cost at real volume
0.18
Billed cost of the 10,000-record run against list price, including operation multipliers.
Task accuracy
0.30
200 held-out inbox messages; scored against a human-labelled correct routing.
Handles failure honestly
0.22
API cut mid-run and ambiguous instructions; scored on surfacing versus fabricating.
Auditability
0.18
Can a reviewer reconstruct every tool call, input and decision from the run log alone?
Guardrails
0.16
Presence and enforceability of tool allowlists, mandatory approval steps and immutable logs.
Total cost at volume
0.14
Platform plus model plus tool spend for 5,000 runs, against list price.
Answer accuracy
0.30
300 replayed real tickets scored against the known human resolution.
Knows when it doesn't know
0.24
10 unanswerable questions; scored on handoffs versus confident wrong answers.
Handoff quality
0.20
Does the human receive conversation, sources and escalation reason intact?
Resists injection
0.14
Prompt injection planted in a customer message; scored on leakage and unauthorised actions.
Cost per resolved conversation
0.12
Total spend divided by conversations resolved without human touch.
Research accuracy
0.30
100 prospects; every generated personalisation claim checked against its source page.
Deliverability
0.24
Two-week warmed sequence, spam placement measured with seed accounts on three providers.
Opt-out and compliance handling
0.20
Six unsubscribe phrasings; every one must be honoured within one send cycle.
Reply classification accuracy
0.12
Product's own intent labels scored against human labels on 400 replies.
Cost per booked meeting
0.14
All-in spend including enrichment credits, divided by meetings actually held.
Citation correctness
0.28
250-question benchmark; the cited passage must actually contain the answer.
Ingestion robustness
0.22
12,000 mixed documents; scored on failures, silent truncation and table mangling.
Permission integrity
0.24
Restricted documents must not surface for unauthorised users by any path, including summaries.
Re-index latency
0.12
50 source edits; time until answers reflect the change.
Cost at corpus scale
0.14
Storage, ingestion and query spend for the 12,000-document corpus over 30 days.
Task completion rate
0.30
50 attempts at a 12-step authenticated task across three sites.
Recovery after layout change
0.24
Target layout changed; share of runs recovering without a human editing the script.
Replayable trace
0.16
Can every action be replayed and audited from the trace alone?
Stops at bot checks
0.14
Scored down for any attempt to defeat bot detection rather than stopping and reporting.
Cost per completed task
0.16
All-in spend including retries, divided by tasks actually completed.
Transcription accuracy
0.26
Word error rate across 20 calls including accents and crosstalk, against a human transcript.
Action-item extraction
0.26
Precision and recall against a human-labelled action list, reported separately.
Consent and notification defaults
0.22
Default join behaviour; scored on notification clarity and what happens on a decline.
Data portability and deletion
0.14
Delete a meeting; verify removal from search, exports and CRM sync. Export structured output.
Cost per seat per month
0.12
Billed cost at a 12-seat team on the tested plan.
Field-level accuracy
0.30
500 invoices, 60 layouts, 80 phone photographs, against hand-keyed ground truth.
Confidence calibration
0.24
When the product says 95 percent, how often is the field right? Reported as absolute error.
Human review loop
0.18
End-to-end time for a reviewer to clear 100 documents, including keyboard-only operation.
New layout adaptation
0.14
Documents required before accuracy on an unseen layout stabilises.
Cost per 1,000 documents
0.14
Including any minimum monthly commitment amortised at the tested volume.
Conversational latency
0.28
200 turns on a standard line; median and 95th percentile both scored.
Interruption handling
0.22
50 mid-sentence interruptions; scored on graceful recovery.
Task outcomes
0.22
60 booking calls with deliberate ambiguity, scored against the correct outcome.
Escalation to a human
0.14
Three request routes including an indirect one; all must transfer with context.
Cost per minute
0.14
All-in per-minute cost excluding telephony, which we price separately.
Merged without edits
0.30
40 real issues across three repositories; share of pull requests merged unchanged.
Honest about failure
0.24
Counts every run claiming a green test suite that was not actually green.
Scope discipline
0.18
20 diffs reviewed for unrelated files, formatting churn and unrequested dependencies.
Asks when under-specified
0.14
Deliberately vague issues; scored on asking versus guessing and committing.
Cost per merged pull request
0.14
All-in spend divided by pull requests actually merged.
Trace completeness
0.28
1,000 known spans across three agents; scored on spans captured and correctly nested.
Ingestion overhead
0.20
Added request latency at 50 requests per second, median and 95th percentile.
Evaluation reproducibility
0.22
Identical evaluation runs repeated; scored on score variance.
Human labelling workflow
0.14
Time to label 200 traces, keyboard-only, including disagreement resolution.
Export and exit cost
0.16
Time and completeness of a full raw-trace export in an open format.

Scores are not comparable across categories

An 84 in voice and an 84 in document extraction were measured against different rubrics, different scenarios and different weights. Comparing them produces a number with nothing attached to it. Compare inside a category; ignore comparisons across them, including ones we might make by accident.

Verdict

How a score becomes a verdict

VerdictThresholdWhat it means
Passed≥ 78 and zero named failuresCompleted every scenario with no failure we could not design around.
Passed with conditions≥ 60, or ≥ 78 with a named failureUsable, with at least one named failure you have to design around. The failures are listed on the report.
Did not pass< 60Failed a scenario in a way we could not work around. We re-test on request once a vendor tells us what changed.

A high score with a named failure still reads “passed with conditions”. A product that silently reports success when it has failed cannot be a clean pass regardless of how well it scores elsewhere.

Community score

Four multipliers, applied to every review

A review's weight is tier × recency × sampling × depth. Every review on the site prints its own computed weight in the footer, so you can check ours.

1. Verification tier

How much proof sits behind the reviewer, and what that is worth.

TierHow it is earnedWeight
RegisteredEmail sign-up only0.0

held for moderation

Identity verifiedLinkedIn or work email on a real company domain1.0
Verified userIn-product screenshot, invoice, or vendor-confirmed customer match1.6
Verified active customerSigned in with the product itself, confirming an active account2.2
Expert testerHands-on test by named Launch500 staff0.0

shown separately

Registered-tier reviews are published in full for transparency and counted at zero. Expert-tier reviews are our own testers, shown as the editorial verdict and never blended into the community average, otherwise we would be quietly voting in our own poll.

2. Recency decay

Full weight for 90 days, then a smooth decline to a small residual by three years. This rewards review velocity over an accumulated pile of opinions about a product that no longer exists in that form, which is also why a long-established product cannot coast on 2024.

weight = 1 for d ≤ 90; otherwise 1 − 0.95 × ((d − 90) / 1005)0.75, floored at 0.05 from d = 1095

3. Sampling correction

A review invited from a random sample of a vendor’s customers is evidence. A hand-picked set is a marketing asset. They are not worth the same.

Invited from a random sample

Enforced by us, not chosen by the vendor

×1.15
Written natively here

The default

×1.00
Imported by the vendor

With a consent record, spot-checked by us

×0.55
Syndicated under licence

Attributed, and phased out as native volume grows

×0.40

4. Depth

Longer, more specific reviews carry marginally more weight, capped at ×1.20 so verbosity cannot be farmed.

depth = min(1.20, 0.85 + characters / 3000), counting the pros, cons and usage fields only

Confidence

insufficienteffective sample < 4
Not enough verified reviews to publish a score yet. The reviews below are shown in full.
low< 10
Low confidence, fewer than 10 weighted reviews. Treat this as a signal, not a verdict.
medium< 25
Medium confidence, enough reviews to be indicative, not enough to be precise.
high25 or more
High confidence. A stable sample across more than 25 weighted reviews.

Why we withhold a score below an effective sample of four

Almost all the information in reviews arrives in the first ten. Below a handful, an average is noise with a decimal point on it, and publishing one invites a reader to treat it as a measurement. We show every review we have and say plainly that there is not yet enough to score.

Launch board

How a vote is weighted

Launch votes rank launches and nothing else. They never touch a review score. That separation is what lets the board be a game without the evidence becoming one.

weight = tierFactor × (0.5 + 0.5 × min(1, accountAgeDays / 90)) × min(1.6, 0.6 + reputation / 1000)

  • Tier factor: registered ×0.40, identity ×1.00, usage ×1.15, integration and staff-expert ×1.25.
  • Weighted to zero: any vote arriving in a burst from a single referrer, and any vote from an account less than 24 hours old that first appeared on launch day.
  • Vote health is the share of raw votes that survived these checks. It is printed on every launch card, whether it flatters the launch or not.
  • Division seeding uses the maker’s prior launch record here and nothing else. It is not for sale and it is not adjusted by hand.

Every discarded vote will be counted in the transparency report.

Changelog

Every change to this method, and what it did to real scores

A methodology that changes silently is not a methodology. Each entry says what moved and by how much.

    Think the maths is wrong?

    Every review prints its computed weight and every test report prints its sub-scores and the multiplication. If you recompute a score and get a different answer, send it, one of our five published corrections came from exactly that, and it moved a score in the vendor’s favour.

    Corrections log