Virtual QA organization

A QA crew that learns your product, then tests it.

Qrew gives your engineering team specialised testing agents trained to think like QA engineers. They explore, verify and remember. You decide how much they do on their own.

Pax · exploring · checkout with a saved card Rey · verifying · QTS-142 · 6 of 7 requirements Nova · planning · next step for QTS-148 Wren · drafting · 9 test cases from the spec Jett · authoring regression test · QTS-131 Rey · blocked · sign-in wall on staging · reported

Measured, not claimed

Tested against hundreds of real websites before it touches yours.

68%
precision on the WebTestBench benchmark, where frontier models score 22 to 35%
0
false passes in verification runs on complex live storefronts and communities
3–4×
more product areas covered in swarm mode than by a single agent in the same run
100s
of public and purpose-built web apps in the lab the harness trains and is scored on

Internal benchmarks, 2026. We walk through the methodology in a demo.

Your crew

Five colleagues with distinct jobs.

Each agent has a name, a role and a record of everything it has done. They work your board alongside your QA engineers and hand decisions to a person at the moments that matter.

  • Maestro

    Nova

    Orchestrator

    Reads the board, plans the work and proposes what happens next on every ticket.

    planning · next step for QTS-148

  • Pax

    Explorer

    Goes off-script on the live app to find the problems nobody wrote a test for.

    exploring · checkout with a saved card

  • Rey

    Verifier

    Checks every requirement on a ticket, then looks at what the change could have broken.

    verifying · QTS-142 · 6 of 7

  • Wren

    Curator

    Turns tickets, specs and docs into test cases your team can review and keep.

    drafting · 9 test cases

  • Jett

    Regression author

    Turns proven flows into regression tests that propose their own repairs when the UI moves.

    authoring regression test · QTS-131

Jett wants to add a regression test for checkout with a saved card.

AcceptLater

The Qrew harness

Built in-house. Trained like a QA engineer.

Qrew runs on its own agent harness, developed and scored against hundreds of real websites. It plans with test design, handles the awkward parts of modern web apps and reports only what it can back with evidence.

Precision on WebTestBench · 14 apps · 239 checks

Qrew harness 68%
Frontier models 22–35%
Frontier-model range as published with the benchmark. Qrew figure from our internal run, 2026.

Thinks in test design

It works from risks and requirements the way a trained tester does, so a run covers the edges as well as the happy path.

  • Boundary values
  • Equivalence classes
  • State transitions
  • Negative paths
  • Risk-based priority

Operates what others can't

Modern interfaces are full of controls that break generic browser agents. The harness handles them and says plainly when something cannot be operated.

  • Custom dropdowns
  • Date pickers
  • Drag and drop
  • Canvas
  • File uploads
  • Dialogs

Evidence on every claim

Each finding states what happened, what should have happened and where that expectation comes from. Every run is recorded so anyone can replay it.

Actual Total shows €0.00 after removing a coupon
Expected Total returns to €48.90 · checkout spec

Honest when blocked

A bot wall, a missing test account or a dead button stops a run in a way the crew reports. It never turns an unknown into a pass.

  • Passed
  • Failed
  • Blocked
  • Unverifiable

Swarm mode

One agent walks a path. A swarm covers the map.

For deep exploration, a coordinator maps your product and sends divers into separate areas at the same time. Each diver gets its own browser and its own recording. They share what they learn and never fight over the same test data.

  • 3–4× more areas explored in the same run
  • One recording per diver, so every finding is replayable
  • Single agent stays the default. Switch the swarm on per run or for a whole product.
Areas explored 0 / 20
Illustration of one run with the same time budget.

Agentic.
Without losing control.

Product memory

Run two starts where run one stopped.

Every finding, risk and decision lands in a shared knowledge base with who found it, on which build and with what evidence. The next run reads it first, so the crew spends its time on what nobody has checked yet.

A lesson only becomes part of the crew's habits once a second, independent run confirms it. One lucky guess never turns into a rule.

Run 1build 4.2

Maps 12 areas of the app and records 3 findings with evidence.

Run 2build 4.3

Starts with what run 1 knows, skips the covered ground and checks 9 new areas.

Run 3build 4.4

Confirms 2 fixes, catches 1 regression and adds a confirmed lesson.

Known for this product

  • riskPostcodes with a space fail at checkout run 1 · build 4.2 · 3 evidence
  • verified fixedCoupon removal resets the total run 3 · build 4.4
  • lessonStaging needs a fresh cart per test confirmed by 2 runs

Illustration.

Control

You decide how far it goes.

Set the crew's autonomy per board. Whatever you choose, every action lands in one activity record and every decision you need to make waits in one place.

The crew suggests. Nothing moves until you do.

The crew works the board and waits for your yes at every decision that matters.

The crew moves tickets on its own and keeps a full record of what it did and why.

Verification never lies

Blocked and unverifiable are outcomes of their own. A run that hits a limit says it stopped.

People confirm findings

Agent findings start as leads. Nothing reaches your tracker as a confirmed bug until a person says so.

Stays inside its scope

Each product has an authorised test scope. The crew cannot leave it, and it refuses buttons you mark as off limits.

Secrets stay sealed

Test-account passwords are encrypted at rest and masked everywhere. No model ever reads one.

Works with the tools you already run.

  • JJira
  • CConfluence
  • ZZephyr
  • XXray
  • PPlaywright
  • CICI webhooks
  • >_REST API and CLI
  • PRPull request checks Coming soon

Pricing

Pay for the work. Not for seats.

Agents do the testing, so you never pay per user. Every plan comes with monthly credits for crew work, and you can add more whenever you need them.

Free

Try the crew on one product.

€0 / month

100 credits a month

  • 1 product
  • Explore and verify with a single agent
  • Findings with evidence and recordings
  • Extra credits when you need them
Request access

Starter

For a first team putting the crew on its board.

€49 / month

500 credits a month

  • 1 product
  • The full crew
  • Jira integration
  • Regression tests, replays use no credits
Request access

Team

For several products and QA groups.

€599 / month

6,000 credits a month

  • 5 products
  • Everything in Pro
  • Dedicated instance, hosted in the EU
  • Priority support, lowest credit rate
Book a demo

Enterprise

Let's talk

  • Unlimited products and volume credits
  • Self-hosting or your own model key
  • Apps behind your firewall
  • Security review and onboarding support
Contact sales

How credits work

A credit is a unit of crew work. Simple tasks use a few, deep exploration uses more. You see the estimate before a run starts.

Verify a ticket
about 12
Explore one area
about 25
Swarm exploration
60 to 100
Draft test cases
about 2
Replay a regression test
0

Plan credits renew every month. Extra credits cost €0.15 each, or €0.12 on Team, and stay valid for 12 months.

You set hard limits and get alerts at 80 and 100%. When credits run out the crew pauses. We never bill overage.

Indicative pricing during early access. Prices exclude VAT. Yearly billing saves two months.

See it live

See it on your product.

You won't find a gallery of screenshots here. In a 30-minute call we point the crew at an app and you watch it plan, test and report.

  • A live run on a demo app or your own test environment
  • How the harness scores against other approaches
  • Which plan and how many credits fit your team

Questions teams ask first

Does Qrew replace our QA engineers?

No. The crew takes the repetitive and wide-coverage work, and your QA engineers decide what is a bug, what ships and how far the agents may act.

What do you need from us to start?

A test environment the crew may use, a test account if the app needs one, and access to your Jira project. Specs and Confluence pages help the crew learn faster.

Which applications can it test?

Web applications that run in a browser, from storefronts to internal tools. Apps behind a firewall are possible on the Enterprise plan.

Where is our data stored?

Each customer gets a dedicated instance hosted in the EU. Enterprise customers can host Qrew themselves.