Method
The same wall, every time, and we tell you which wall.
A marketplace can tell you what other people thought. It cannot tell you what happens when you cut an API mid-run, plant an instruction inside a customer message, or push ten thousand records through and read the invoice. That is the gap this bench exists to fill, and it only works if the procedure is public.
- 0
- Products tested
- 0
- Hours logged
- 0
- Failures named
- 68
- Scenarios in use
Rules
What we hold ourselves to
- 01
Scenarios are published before anything is tested
Every category's scenario set is on its hub page and on this one, fixed before the first product in that category goes on the bench. A test designed after seeing the results is not a test.
- 02
Every product in a category gets the identical run
Same scenarios, same volumes, same rubric, same weights. Where a product cannot attempt a scenario, that is recorded as a zero with a note, not quietly dropped from the denominator.
- 03
We test the plan a real buyer would be on
Usually the mid tier, because that is what most teams end up paying for. The plan and build number are printed at the top of every report, because a result against an enterprise tier tells a small team nothing.
- 04
Failures are named, numbered and quotable
Not softened, not folded into a sub-score, not described as 'an area for improvement'. If it broke, the report says what broke and what we were doing at the time.
- 05
Hours are logged and published
Between 12 and 41 hours per product so far. It is the single easiest thing to fake in this industry and the single easiest thing for a vendor to challenge, which is why we print it.
- 06
A test date cannot be bought, moved or refused
Order is set by search and prompt demand. A vendor cannot pay to be tested sooner, cannot pay to be tested later, and cannot decline. Declining to provide access is recorded on the report.
- 07
We re-test when a vendor says what changed
Free, and usually within a fortnight. What we will not do is remove a result because it is old, reports carry their date, and anything over nine months carries a staleness flag.
- 08
We never accept anything from a vendor
No hardware, no hospitality, no consulting work, no equity, no pre-briefing. We pay for our own accounts at list price, and the invoices are part of the test record.
The bench, category by category
Exactly what every product is put through
Customer engagement and lifecycle messaging platforms
0 tested · reviewed 28 September 2026Platforms that hold behavioural customer data and send messages across several channels off the back of it, branching on what somebody did rather than on a list they sit in. The heavier end of messaging, where the product is the segmentation as much as the sending.
Must be true to be listed
- Messages across at least three channels, typically email, push and SMS or in-app
- Journeys that branch on events and traits, not just scheduled sends
- Ingests behavioural data from a product, rather than only from a contact list
- A documented export of profiles and event history
Scenarios run on every product
Customer support and help desk software
0 tested · reviewed 28 September 2026Tools that take inbound customer contact, route it to somebody who can answer, and keep the history. Spans shared inboxes through to full help desks with knowledge bases and automated first-line handling.
Must be true to be listed
- A shared queue with assignment and status, not just a forwarding address
- Conversation history attached to a customer across channels
- Self-serve signup and published pricing, or pricing we could obtain and snapshot
- A documented export of conversations and contacts
Scenarios run on every product
Email marketing and lifecycle messaging software
0 tested · reviewed 28 September 2026Tools that hold a subscriber list, send broadcast and automated email against it, and report what happened. The category spans simple newsletter senders through to lifecycle platforms that branch on behaviour.
Must be true to be listed
- Owns the subscriber list and the sending, rather than being a template editor that hands off to somebody else
- Automated sequences that branch on subscriber behaviour, not just one-off broadcasts
- Self-serve signup and published pricing, or pricing we could obtain and snapshot
- A documented export of subscribers and their engagement history
Scenarios run on every product
CRM software and sales pipeline tools
0 tested · reviewed 16 September 2026Systems that store customer and deal records, track a pipeline through stages, and give a sales team a shared view of who said what and what happens next.
Must be true to be listed
- Contact, company and deal objects with a configurable pipeline, not just a contact list
- Self-serve signup and published pricing, or pricing we could obtain and snapshot
- An API and a documented data export covering every object
- Granular permissions, or an explicit statement that there are none
Scenarios run on every product
- 1.Import 25,000 messy contacts with duplicates, bad phone formats and three date conventions; score field fidelity and duplicate handling against a known-good set
- 2.Build a seven-stage pipeline with required fields and rotting rules, then have a non-admin try to move a deal they should not be able to
- 3.Measure list-view and report performance at 250,000 records, cold and warm
- 4.Run a full export and time how long it takes to get every object out in a usable format
Helpdesk and customer support ticketing software
0 tested · reviewed 10 September 2026Shared inbox and ticketing systems where a support team receives, assigns, tracks and answers customer conversations across email, chat and other channels.
Must be true to be listed
- Shared queue with assignment, status and a full conversation history
- At least two channels, one of which is email
- Reporting on response and resolution time that a customer can export
- Documented data retention and deletion behaviour
Scenarios run on every product
- 1.Run 2,000 conversations across email and chat through a four-person rota; measure collision rate, misassignment and lost threads
- 2.Break the email connection mid-shift and record what the tool tells the team and what happens to inbound mail
- 3.Reconcile the product's own resolution-time reporting against our timestamps on the same 2,000 conversations
- 4.Delete a customer's data on request and verify it is gone from search, exports and reporting
Accounting and invoicing software for small business
0 tested · reviewed 13 September 2026Software that records income and expenses, issues invoices, reconciles against a bank feed and produces the reports an accountant or a tax authority expects.
Must be true to be listed
- Double-entry bookkeeping, not just invoice generation
- Automated bank feeds in at least one major market, with a documented fallback
- Multi-currency support, or an explicit statement that there is none
- Export in a format an accountant can actually open
Scenarios run on every product
- 1.Reconcile 1,200 bank transactions across two currencies against 340 invoices; score auto-match accuracy and false matches against a hand-reconciled ledger
- 2.Break the bank feed for 72 hours and measure how the product recovers, and whether it double-imports on reconnection
- 3.Issue, part-pay, credit-note and void an invoice; check the audit trail survives all four
- 4.Produce a year-end pack and give it to a qualified accountant to find what is missing
Project management and team task tracking software
0 tested · reviewed 7 September 2026Tools where a team plans work, assigns it, tracks status across boards or timelines, and sees what is blocked or late.
Must be true to be listed
- Assignable work items with status, owner and due date
- At least two views of the same data, such as board and timeline
- Permissions that can restrict a project to a subset of the workspace
- Full export including comments and history
Scenarios run on every product
- 1.Model a 400-task programme with dependencies across four teams; change one date and measure what the tool correctly cascades
- 2.Invite a contractor with restricted access and attempt to reach a project they should not see, by search, by link and by notification
- 3.Measure load time on a 2,000-item board, cold, on a mid-range laptop
- 4.Export the workspace and check whether comments, attachments and history survive
HR software and employee onboarding systems
0 tested · reviewed 5 September 2026Systems of record for employees, contracts, personal data, leave, onboarding and offboarding checklists, and the reporting an HR team and a regulator need.
Must be true to be listed
- Employee record with contract and compensation history
- Leave management with an approval chain
- Role-based access that separates HR, managers and employees
- Documented data-protection posture including retention and deletion
Scenarios run on every product
- 1.Onboard 50 employees across three countries with different contract types; score how much needed manual correction
- 2.Attempt to reach another employee's compensation as a line manager, by report, by export and by search
- 3.Run a leave year-end rollover with carry-over rules and part-time pro-rating; check the arithmetic by hand
- 4.Exercise a data-deletion request end to end and verify what survives in backups and reports
Product analytics and business intelligence software
0 tested · reviewed 19 September 2026Tools that collect events or connect to a warehouse, and let a non-engineer answer questions about usage, funnels, retention and revenue without writing SQL.
Must be true to be listed
- Self-serve exploration for a non-engineer, not just a dashboard someone else built
- Funnel and retention analysis as first-class features
- Documented handling of identity resolution across anonymous and known users
- Raw data export or a warehouse connection the customer controls
Scenarios run on every product
- 1.Send a known event stream of 5 million events with deliberate duplicates and late arrivals; score counts against ground truth
- 2.Build the same three-step funnel in each product and compare the numbers they report for identical data
- 3.Test identity stitching: anonymous session, then signup, then login on a second device
- 4.Time a non-engineer building a retention chart unaided, from a written question
Workflow automation and integration software
0 tested · reviewed 15 September 2026Tools that connect applications and run multi-step processes automatically, triggered by an event or a schedule, without an engineer writing and hosting the glue code.
Must be true to be listed
- Connects at least 100 third-party applications, or exposes a generic HTTP and webhook builder
- Supports branching, error handling and retries, not just linear two-step triggers
- Publicly documented pricing, or pricing we could obtain and snapshot
- Generally available to self-serve buyers; no invite-only betas
Scenarios run on every product
- 1.Build a five-step order-to-invoice flow with a conditional branch and a failure path
- 2.Force three classes of failure (rate limit, malformed payload, expired auth) and record what the tool reports and retries
- 3.Run 10,000 records through and measure wall-clock time and the billed usage
- 4.Hand the finished flow to a non-technical teammate and time how long they take to change one step
AI agent builders and no-code agent platforms
0 tested · reviewed 18 September 2026Platforms for building software agents that use a language model to decide what to do next, call tools or APIs, and carry out multi-step tasks with limited human supervision.
Must be true to be listed
- The agent chooses its own sequence of tool calls at runtime. A fixed prompt chain is not an agent
- Ships a tool or function-calling layer with at least 20 pre-built connectors
- Offers some form of run history or trace so a human can audit what the agent did
- Available without a sales call
Scenarios run on every product
- 1.Build an agent that triages a shared inbox, reads a knowledge base, and drafts or escalates, measure correct routing over 200 held-out messages
- 2.Give the agent a deliberately ambiguous instruction and record whether it asks, guesses, or fabricates
- 3.Cut off a required API mid-run and check whether the failure is surfaced or silently swallowed
- 4.Re-run the identical task 20 times and measure output variance and cost spread
AI customer support agents and helpdesk automation
0 tested · reviewed 11 September 2026Agents that read incoming support conversations, answer from a company's own documentation and ticket history, and either resolve the conversation or hand it to a human with context attached.
Must be true to be listed
- Resolves conversations end to end, not just suggests replies to an agent
- Grounds answers in customer-supplied sources with a citation a human can check
- Integrates with at least two mainstream helpdesks
- Publishes or will disclose its deflection measurement method
Scenarios run on every product
- 1.Load a 900-article help centre and 5,000 historical tickets, then replay 300 real questions and score answer accuracy against the known resolution
- 2.Ask ten questions the documentation does not answer and count how many produce a confident wrong answer instead of a handoff
- 3.Measure handoff quality: does the human receive the conversation, the sources consulted and the reason for escalation?
- 4.Attempt a prompt injection from inside a customer message and see whether the agent leaks its instructions or takes an unauthorised action
AI SDR and sales outreach agents
0 tested · reviewed 9 September 2026Agents that research prospects, write and send outbound sequences, and book meetings with limited human involvement. The automated end of the sales development function.
Must be true to be listed
- Performs its own prospect research from public sources rather than only merging CRM fields
- Sends through the customer's own mailboxes with deliverability controls
- Provides reply classification and an auditable send log
- Has a documented policy on compliance with anti-spam law in at least the US and EU
Scenarios run on every product
- 1.Research 100 named prospects and score factual accuracy of each generated personalisation line against source pages
- 2.Run a two-week warmed sequence and measure spam placement using seed accounts across three mail providers
- 3.Test opt-out handling: reply 'unsubscribe' in six phrasings and check every one is honoured within one send cycle
- 4.Review generated copy for claims the product invented about the prospect's company
RAG platforms and AI knowledge base software
0 tested · reviewed 16 September 2026Systems that index a company's documents and data so a language model can answer from them with citations, covering ingestion, chunking, retrieval, permissions and evaluation.
Must be true to be listed
- Handles ingestion and retrieval, not just vector storage
- Enforces per-document access control that survives retrieval
- Returns citations that resolve to a specific source location
- Offers a retrieval evaluation or at least exportable retrieval traces
Scenarios run on every product
- 1.Index 12,000 mixed documents (PDF, HTML, spreadsheets, scanned contracts) and measure ingestion failures and silent truncation
- 2.Run a 250-question retrieval benchmark and score citation correctness, not just answer plausibility
- 3.Permission test: confirm a restricted document never surfaces for a user without access, including via summarisation
- 4.Update 50 source documents and measure how long until answers reflect the change
Browser automation and web agent software
0 tested · reviewed 19 September 2026Agents that drive a real browser to complete tasks on websites, logging in, navigating, filling forms and extracting data, where no API is available.
Must be true to be listed
- Drives a real browser engine, not an HTTP scraper
- Handles authenticated sessions with credential storage the customer controls
- Provides a replayable trace of every action taken
- States a policy on site terms of service and rate limiting
Scenarios run on every product
- 1.Run a 12-step authenticated task across three unfamiliar sites and measure completion rate over 50 attempts
- 2.Change a target site's layout and measure how many runs recover without a human editing the script
- 3.Measure cost and wall-clock time per completed task, including retries
- 4.Check what the product does when it encounters a bot check. We require it to stop and report, not attempt a bypass
AI meeting notetakers and call recording software
0 tested · reviewed 8 September 2026Tools that join calls, transcribe them, and produce summaries, action items and searchable records, usually syncing the result into a CRM or workspace.
Must be true to be listed
- Joins or records at least two of Zoom, Meet and Teams
- Produces structured output (actions, decisions) not just a transcript
- Documents its consent and recording-notification behaviour
- Offers per-user data deletion
Scenarios run on every product
- 1.Transcribe 20 recorded calls including two with heavy accents and one with three people talking over each other; score word error rate against a human transcript
- 2.Score action-item extraction against a human-labelled list, precision and recall reported separately
- 3.Check the consent flow: does every participant get notified, and what happens if one declines?
- 4.Delete a meeting and verify the transcript is gone from search, exports and the CRM sync
AI document processing and data extraction software
0 tested · reviewed 12 September 2026Software that reads unstructured documents, invoices, contracts, forms, statements, and outputs structured fields with a confidence score and a path for human review.
Must be true to be listed
- Outputs structured fields with per-field confidence, not just document-level text
- Provides a human-in-the-loop review queue
- Handles scanned and photographed documents, not only digital PDFs
- Supports field-level accuracy measurement against ground truth
Scenarios run on every product
- 1.Process 500 invoices across 60 vendor layouts, including 80 phone photographs, and score field-level accuracy against a hand-keyed ground truth
- 2.Measure calibration: when the product reports 95 percent confidence, how often is the field actually right?
- 3.Time the human review loop end to end for 100 documents
- 4.Introduce a new unseen layout and measure documents needed before accuracy stabilises
Voice AI agents and automated phone answering
0 tested · reviewed 17 September 2026Agents that hold a real-time spoken conversation over the phone or in an app, answering, qualifying, booking and escalating, using speech recognition, a language model and synthesised speech.
Must be true to be listed
- Handles live, interruptible two-way conversation, not just IVR menus or voicemail
- Reports end-to-end latency
- Can transfer to a human mid-call with context
- Provides call recordings and transcripts the customer controls
Scenarios run on every product
- 1.Measure end-to-end response latency across 200 turns on a standard line, reported at median and 95th percentile
- 2.Interrupt the agent mid-sentence 50 times and measure how often it recovers gracefully
- 3.Run 60 booking calls with deliberate ambiguity ('sometime next week, not Tuesday') and score correct outcomes
- 4.Request a human three ways, including indirectly, and check every route transfers with context
AI coding agents and automated pull request tools
0 tested · reviewed 20 September 2026Agents that take a written task or an issue, change code across a repository, run tests, and open a pull request for human review.
Must be true to be listed
- Operates across a repository, not a single file or an editor autocomplete
- Runs the project's own tests before proposing a change
- Opens a reviewable diff rather than writing directly to a default branch
- Documents what repository access it requires
Scenarios run on every product
- 1.Assign 40 real issues from three open-source repositories of different sizes and measure the share of pull requests merged without human edits
- 2.Measure how often the agent claims tests pass when they do not
- 3.Review 20 diffs for scope creep, unrelated files touched, formatting churn, dependency additions
- 4.Test behaviour on an under-specified issue: does it ask, or does it guess and commit?
LLM observability and agent evaluation tools
0 tested · reviewed 21 September 2026Tooling that records what an agent or language-model application did, prompts, tool calls, costs, latencies, and scores output quality against test sets so regressions are caught before users find them.
Must be true to be listed
- Captures full execution traces including tool calls and token costs
- Supports running an evaluation set against a candidate change
- Allows human labelling of traces
- Exports raw traces in an open format
Scenarios run on every product
- 1.Instrument a three-agent application and measure trace completeness against a known ground truth of 1,000 spans
- 2.Measure ingestion overhead added to request latency at 50 requests per second
- 3.Run an evaluation suite and check that scores are reproducible across identical runs
- 4.Attempt a full data export and time how long it takes to get usable traces out
Productivity and document collaboration software
0 tested · reviewed 23 September 2026Tools for writing, organising and sharing the documents, notes and wikis a team works from, usually combining a text editor with some structure such as databases, tasks or search.
Must be true to be listed
- Used by a team rather than a single person, with shared workspaces and permissions
- Documents are the primary object, not a side feature of something else
- Publicly documented pricing, or pricing we could obtain and snapshot
Scenarios run on every product
Finance, payments and banking software for business
0 tested · reviewed 23 September 2026Software that moves, holds or reconciles company money: taking payments, paying people, managing spend, and converting between currencies.
Must be true to be listed
- Handles real money movement, not only reporting on it
- Regulated or partnered with a regulated institution in at least one of our markets
- Publicly documented pricing, including the FX margin where one applies
Scenarios run on every product
AI assistants and general-purpose AI tools
0 tested · reviewed 23 September 2026General-purpose AI products used across a company rather than for one job: chat assistants, writing and research tools, and the model providers behind them.
Must be true to be listed
- General-purpose rather than built for a single workflow, which has its own category
- Available directly to a business, not only as an API for developers
- States which models it runs and whether your inputs train them
Scenarios run on every product
SEO tools and web analytics software
0 tested · reviewed 23 September 2026Tools that measure how people find and use a website: search rankings, backlinks, keyword research, and traffic and behaviour analytics.
Must be true to be listed
- Measures a public website or its search performance, not only in-product events
- Own data collection or a disclosed data source, not a reseller of someone else's index
- Publicly documented pricing, including limits on rows, keywords or events
Scenarios run on every product
Website builders and no-code app platforms
0 tested · reviewed 23 September 2026Platforms for building a website, store or internal application without writing the hosting and deployment yourself, whether by visual editor or by configuration.
Must be true to be listed
- A non-engineer can ship something usable without writing application code
- Hosting is included rather than left to the customer
- Publicly documented pricing, including what happens to your site if you stop paying
Scenarios run on every product
Developer tools and software engineering platforms
0 tested · reviewed 23 September 2026Tools engineering teams use to write, review, ship and run software: repositories, CI, error tracking, observability and the platforms around them.
Must be true to be listed
- Bought by or for an engineering team as part of shipping software
- Integrates with mainstream version control or deployment targets
- Publicly documented pricing, including per-seat and usage components
Scenarios run on every product
Founder essentials, incorporation and company admin
0 tested · reviewed 23 September 2026The administrative software a new company needs to exist and stay compliant: incorporation, cap tables, equity, contracts, filings and company secretarial work.
Must be true to be listed
- Serves the company itself rather than a department inside it
- Covers a legal or statutory obligation in at least one of our markets
- States clearly which jurisdictions it actually supports
Scenarios run on every product
Team communication and meeting software
0 tested · reviewed 23 September 2026Where a team talks: chat, video calls, async video and the scheduling around them.
Must be true to be listed
- Real-time or async communication between people is the primary function
- Supports a whole team or company, not only one-to-one use
- States its recording, retention and consent defaults
Scenarios run on every product
Marketing automation and email software
0 tested · reviewed 23 September 2026Tools for reaching an audience you already have: email campaigns, newsletters, lifecycle automation and the lists behind them.
Must be true to be listed
- Sends on your behalf to a list you own, with a working unsubscribe
- Handles consent and suppression rather than leaving it to you
- Publicly documented pricing, including the cost of list growth
Scenarios run on every product
Design tools and visual collaboration software
0 tested · reviewed 23 September 2026Tools for designing interfaces and brand assets, and for the whiteboarding and review that happens around them.
Must be true to be listed
- Produces or reviews visual work as its primary output
- Multiple people can work in the same file or board
- States what happens to your files if you stop paying
Scenarios run on every product
Cloud hosting and infrastructure platforms
0 tested · reviewed 23 September 2026Where software actually runs: compute, databases, storage, edge networks and the platforms that sit on top of them.
Must be true to be listed
- Runs customer workloads rather than only managing them
- Publishes its own pricing per unit of compute, storage or bandwidth
- States its regions and what data residency it can guarantee
Scenarios run on every product
What a Crash Test cannot tell you
It is one team, one week, one version. It will not catch a pricing change in March, a support team hollowed out in June, or the thing that only breaks at ten times our test volume. That is what the verified reviews beside it are for, and it is why we never average the two into a single number.
Want a scenario added?
The permission-integrity probe that two retrieval products currently fail exists because a member started a thread about it. If there is a failure mode we are not testing for, say so, proposals that get adopted are credited in the methodology changelog.