Method
Two scores. Never blended. All the arithmetic.
One convenient number would be easier to read and would destroy your ability to tell an opinion from a measurement, so we publish two and keep them apart. This page has every weight, every threshold and the exact curve, enough to recompute any score on this site by hand and catch us if it does not come out.
Crash Test score · 0 to 100
Our own hands-on result. One team, one week, one named version, one date. It answers “does this thing work when we push it?” It cannot tell you what year three looks like.
Community score · 0 to 5
Verified reviews from people paying the bill. Many teams over many months. It catches repricing, support decay and month-nine regret. The things a bench never sees, and never the same question as the test.
Crash Test
Per-category weights, fixed before anything is tested
CRM software and sales pipeline tools
weights sum to 1.00| Import fidelity | 0.26 | 25,000 messy contacts with duplicates and mixed formats, scored field-by-field against a known-good set. |
|---|---|---|
| Permission integrity | 0.22 | A non-admin attempts to move, view and export records they should not reach, by every route we could find. |
| Performance at scale | 0.18 | List-view and report load times at 250,000 records, cold and warm. |
| Export and exit cost | 0.16 | Full export of every object; scored on completeness, format and elapsed time. |
| Cost at 25 seats | 0.18 | Billed cost on the cheapest tier that actually carries the permissions a sales team needs. |
Helpdesk and customer support ticketing software
weights sum to 1.00| Collision and assignment | 0.28 | 2,000 conversations across a four-person rota; scored on double-replies, misassignment and lost threads. |
|---|---|---|
| Handles failure honestly | 0.22 | Email connection broken mid-shift; scored on what the team is told and what happens to inbound mail. |
| Reporting accuracy | 0.20 | The product's own resolution-time figures reconciled against our timestamps on the same conversations. |
| Data deletion | 0.14 | A deletion request verified across search, exports and reporting. |
| Cost per agent | 0.16 | Billed cost at a four-agent team on the tier that includes reporting. |
Accounting and invoicing software for small business
weights sum to 1.00| Reconciliation accuracy | 0.28 | 1,200 transactions against 340 invoices; auto-match rate scored against a hand-reconciled ledger. |
|---|---|---|
| False matches | 0.24 | Counts wrong matches the product made confidently. A wrong match scores worse than no match. |
| Bank feed recovery | 0.18 | 72-hour feed outage; scored on recovery and whether it double-imports on reconnection. |
| Audit trail | 0.14 | Issue, part-pay, credit-note and void an invoice; the trail must survive all four. |
| Cost including multi-currency | 0.16 | Tier needed for the tested feature set, plus any FX margin we could establish. |
Project management and team task tracking software
weights sum to 1.00| Dependency handling | 0.26 | 400-task programme across four teams; one date changed, scored on what correctly cascades. |
|---|---|---|
| Guest permission integrity | 0.22 | A restricted contractor attempts access by search, by link and via notifications. |
| Performance at scale | 0.20 | Cold load of a 2,000-item board on a mid-range laptop. |
| Export completeness | 0.14 | Workspace export scored on whether comments, attachments and history survive. |
| Cost including guests | 0.18 | 12 staff plus 6 contractors, billed the way the product actually bills them. |
HR software and employee onboarding systems
weights sum to 1.00| Leave arithmetic | 0.26 | Year-end rollover with carry-over caps and part-time pro-rating, checked by hand. |
|---|---|---|
| Permission integrity | 0.26 | A line manager attempts to reach another employee's compensation by report, export and search. |
| Multi-country onboarding | 0.18 | 50 employees, three countries, mixed contract types; scored on manual correction needed. |
| Deletion and retention | 0.16 | A data-deletion request exercised end to end, including what survives in backups and reports. |
| Cost per employee | 0.14 | Billed cost at 50 employees on the tier carrying the tested feature set. |
Product analytics and business intelligence software
weights sum to 1.00| Count accuracy | 0.28 | 5 million known events with duplicates and late arrivals, scored against ground truth. |
|---|---|---|
| Identity resolution | 0.22 | Anonymous session, then signup, then login on a second device; scored on correct stitching. |
| Answerable by a non-engineer | 0.20 | Time for a non-engineer to build a retention chart unaided from a written question. |
| Raw data access | 0.14 | Export or warehouse connection the customer controls; scored on completeness and latency. |
| Cost at real event volume | 0.16 | Modelled including staging and debug traffic, which most vendors bill for. |
Workflow automation and integration software
weights sum to 1.00| Reliability at volume | 0.28 | 10,000 records through a 5-step flow; scored on completed runs, silent drops and duplicate side effects. |
|---|---|---|
| Handles failure honestly | 0.22 | Three induced failure classes; scored on whether the tool retried correctly and reported accurately. |
| Time to first working flow | 0.18 | Stopwatch from empty account to a passing 5-step flow, by a tester new to the product. |
| Handover to a non-specialist | 0.14 | A non-technical teammate changes one step unaided; scored on completion and time. |
| Cost at real volume | 0.18 | Billed cost of the 10,000-record run against list price, including operation multipliers. |
AI agent builders and no-code agent platforms
weights sum to 1.00| Task accuracy | 0.30 | 200 held-out inbox messages; scored against a human-labelled correct routing. |
|---|---|---|
| Handles failure honestly | 0.22 | API cut mid-run and ambiguous instructions; scored on surfacing versus fabricating. |
| Auditability | 0.18 | Can a reviewer reconstruct every tool call, input and decision from the run log alone? |
| Guardrails | 0.16 | Presence and enforceability of tool allowlists, mandatory approval steps and immutable logs. |
| Total cost at volume | 0.14 | Platform plus model plus tool spend for 5,000 runs, against list price. |
AI customer support agents and helpdesk automation
weights sum to 1.00| Answer accuracy | 0.30 | 300 replayed real tickets scored against the known human resolution. |
|---|---|---|
| Knows when it doesn't know | 0.24 | 10 unanswerable questions; scored on handoffs versus confident wrong answers. |
| Handoff quality | 0.20 | Does the human receive conversation, sources and escalation reason intact? |
| Resists injection | 0.14 | Prompt injection planted in a customer message; scored on leakage and unauthorised actions. |
| Cost per resolved conversation | 0.12 | Total spend divided by conversations resolved without human touch. |
AI SDR and sales outreach agents
weights sum to 1.00| Research accuracy | 0.30 | 100 prospects; every generated personalisation claim checked against its source page. |
|---|---|---|
| Deliverability | 0.24 | Two-week warmed sequence, spam placement measured with seed accounts on three providers. |
| Opt-out and compliance handling | 0.20 | Six unsubscribe phrasings; every one must be honoured within one send cycle. |
| Reply classification accuracy | 0.12 | Product's own intent labels scored against human labels on 400 replies. |
| Cost per booked meeting | 0.14 | All-in spend including enrichment credits, divided by meetings actually held. |
RAG platforms and AI knowledge base software
weights sum to 1.00| Citation correctness | 0.28 | 250-question benchmark; the cited passage must actually contain the answer. |
|---|---|---|
| Ingestion robustness | 0.22 | 12,000 mixed documents; scored on failures, silent truncation and table mangling. |
| Permission integrity | 0.24 | Restricted documents must not surface for unauthorised users by any path, including summaries. |
| Re-index latency | 0.12 | 50 source edits; time until answers reflect the change. |
| Cost at corpus scale | 0.14 | Storage, ingestion and query spend for the 12,000-document corpus over 30 days. |
Browser automation and web agent software
weights sum to 1.00| Task completion rate | 0.30 | 50 attempts at a 12-step authenticated task across three sites. |
|---|---|---|
| Recovery after layout change | 0.24 | Target layout changed; share of runs recovering without a human editing the script. |
| Replayable trace | 0.16 | Can every action be replayed and audited from the trace alone? |
| Stops at bot checks | 0.14 | Scored down for any attempt to defeat bot detection rather than stopping and reporting. |
| Cost per completed task | 0.16 | All-in spend including retries, divided by tasks actually completed. |
AI meeting notetakers and call recording software
weights sum to 1.00| Transcription accuracy | 0.26 | Word error rate across 20 calls including accents and crosstalk, against a human transcript. |
|---|---|---|
| Action-item extraction | 0.26 | Precision and recall against a human-labelled action list, reported separately. |
| Consent and notification defaults | 0.22 | Default join behaviour; scored on notification clarity and what happens on a decline. |
| Data portability and deletion | 0.14 | Delete a meeting; verify removal from search, exports and CRM sync. Export structured output. |
| Cost per seat per month | 0.12 | Billed cost at a 12-seat team on the tested plan. |
AI document processing and data extraction software
weights sum to 1.00| Field-level accuracy | 0.30 | 500 invoices, 60 layouts, 80 phone photographs, against hand-keyed ground truth. |
|---|---|---|
| Confidence calibration | 0.24 | When the product says 95 percent, how often is the field right? Reported as absolute error. |
| Human review loop | 0.18 | End-to-end time for a reviewer to clear 100 documents, including keyboard-only operation. |
| New layout adaptation | 0.14 | Documents required before accuracy on an unseen layout stabilises. |
| Cost per 1,000 documents | 0.14 | Including any minimum monthly commitment amortised at the tested volume. |
Voice AI agents and automated phone answering
weights sum to 1.00| Conversational latency | 0.28 | 200 turns on a standard line; median and 95th percentile both scored. |
|---|---|---|
| Interruption handling | 0.22 | 50 mid-sentence interruptions; scored on graceful recovery. |
| Task outcomes | 0.22 | 60 booking calls with deliberate ambiguity, scored against the correct outcome. |
| Escalation to a human | 0.14 | Three request routes including an indirect one; all must transfer with context. |
| Cost per minute | 0.14 | All-in per-minute cost excluding telephony, which we price separately. |
AI coding agents and automated pull request tools
weights sum to 1.00| Merged without edits | 0.30 | 40 real issues across three repositories; share of pull requests merged unchanged. |
|---|---|---|
| Honest about failure | 0.24 | Counts every run claiming a green test suite that was not actually green. |
| Scope discipline | 0.18 | 20 diffs reviewed for unrelated files, formatting churn and unrequested dependencies. |
| Asks when under-specified | 0.14 | Deliberately vague issues; scored on asking versus guessing and committing. |
| Cost per merged pull request | 0.14 | All-in spend divided by pull requests actually merged. |
LLM observability and agent evaluation tools
weights sum to 1.00| Trace completeness | 0.28 | 1,000 known spans across three agents; scored on spans captured and correctly nested. |
|---|---|---|
| Ingestion overhead | 0.20 | Added request latency at 50 requests per second, median and 95th percentile. |
| Evaluation reproducibility | 0.22 | Identical evaluation runs repeated; scored on score variance. |
| Human labelling workflow | 0.14 | Time to label 200 traces, keyboard-only, including disagreement resolution. |
| Export and exit cost | 0.16 | Time and completeness of a full raw-trace export in an open format. |
Scores are not comparable across categories
An 84 in voice and an 84 in document extraction were measured against different rubrics, different scenarios and different weights. Comparing them produces a number with nothing attached to it. Compare inside a category; ignore comparisons across them, including ones we might make by accident.
Verdict
How a score becomes a verdict
| Verdict | Threshold | What it means |
|---|---|---|
| Passed | ≥ 78 and zero named failures | Completed every scenario with no failure we could not design around. |
| Passed with conditions | ≥ 60, or ≥ 78 with a named failure | Usable, with at least one named failure you have to design around. The failures are listed on the report. |
| Did not pass | < 60 | Failed a scenario in a way we could not work around. We re-test on request once a vendor tells us what changed. |
A high score with a named failure still reads “passed with conditions”. A product that silently reports success when it has failed cannot be a clean pass regardless of how well it scores elsewhere.
Community score
Four multipliers, applied to every review
1. Verification tier
How much proof sits behind the reviewer, and what that is worth.
| Tier | How it is earned | Weight |
|---|---|---|
| Registered | Email sign-up only | 0.0 held for moderation |
| Identity verified | LinkedIn or work email on a real company domain | 1.0 |
| Verified user | In-product screenshot, invoice, or vendor-confirmed customer match | 1.6 |
| Verified active customer | Signed in with the product itself, confirming an active account | 2.2 |
| Expert tester | Hands-on test by named Launch500 staff | 0.0 shown separately |
Registered-tier reviews are published in full for transparency and counted at zero. Expert-tier reviews are our own testers, shown as the editorial verdict and never blended into the community average, otherwise we would be quietly voting in our own poll.
2. Recency decay
Full weight for 90 days, then a smooth decline to a small residual by three years. This rewards review velocity over an accumulated pile of opinions about a product that no longer exists in that form, which is also why a long-established product cannot coast on 2024.
weight = 1 for d ≤ 90; otherwise 1 − 0.95 × ((d − 90) / 1005)0.75, floored at 0.05 from d = 1095
3. Sampling correction
A review invited from a random sample of a vendor’s customers is evidence. A hand-picked set is a marketing asset. They are not worth the same.
- Invited from a random sample
- ×1.15
- Written natively here
- ×1.00
- Imported by the vendor
- ×0.55
- Syndicated under licence
- ×0.40
Enforced by us, not chosen by the vendor
The default
With a consent record, spot-checked by us
Attributed, and phased out as native volume grows
4. Depth
Longer, more specific reviews carry marginally more weight, capped at ×1.20 so verbosity cannot be farmed.
depth = min(1.20, 0.85 + characters / 3000), counting the pros, cons and usage fields only
Confidence
- insufficienteffective sample < 4
- Not enough verified reviews to publish a score yet. The reviews below are shown in full.
- low< 10
- Low confidence, fewer than 10 weighted reviews. Treat this as a signal, not a verdict.
- medium< 25
- Medium confidence, enough reviews to be indicative, not enough to be precise.
- high25 or more
- High confidence. A stable sample across more than 25 weighted reviews.
Why we withhold a score below an effective sample of four
Almost all the information in reviews arrives in the first ten. Below a handful, an average is noise with a decimal point on it, and publishing one invites a reader to treat it as a measurement. We show every review we have and say plainly that there is not yet enough to score.
Launch board
How a vote is weighted
weight = tierFactor × (0.5 + 0.5 × min(1, accountAgeDays / 90)) × min(1.6, 0.6 + reputation / 1000)
- Tier factor: registered ×0.40, identity ×1.00, usage ×1.15, integration and staff-expert ×1.25.
- Weighted to zero: any vote arriving in a burst from a single referrer, and any vote from an account less than 24 hours old that first appeared on launch day.
- Vote health is the share of raw votes that survived these checks. It is printed on every launch card, whether it flatters the launch or not.
- Division seeding uses the maker’s prior launch record here and nothing else. It is not for sale and it is not adjusted by hand.
Every discarded vote will be counted in the transparency report.
Changelog
Every change to this method, and what it did to real scores
Think the maths is wrong?
Every review prints its computed weight and every test report prints its sub-scores and the multiplication. If you recompute a score and get a different answer, send it, one of our five published corrections came from exactly that, and it moved a score in the vendor’s favour.