StackConnect World Tour Zurich · 6 Oct 2026

The Agentic QE Scorecard

Is the agent improving your quality, or improving your dashboard? Four questions, scored 0 to 3, that a release manager can ask without reading a single test.

Companion to the talk "When AI Makes Every Dashboard Green" by Ludovico Besana, QA Lead @ Doctolib.

Activity is not evidence

Keep your activity metrics: they tell you about efficiency. Pair each one with an evidence metric that tells you whether you can ship.

Measures agent activityMeasures release confidence
Tests generatedCritical risks covered by tests you have seen fail
Coverage %Injected faults the suite actually catches
Flaky failures healedHeals reviewed and classified by a human
Triage timeRoot causes verified by reproduction
Pipeline pass rateResidual risk stated, with a named owner

Four dimensions, four questions

01

Risk coverage

"Are we testing what would hurt, or what was easy to generate?"

Risk-weighted coveragecritical risks with ≥ 1 test seen failing ÷ critical risks identified with product
02

Evidence quality

"Would this evidence have turned red if we were wrong?"

Fault detection rateinjected faults caught ÷ faults injected. For triage: agent root causes verified by reproduction or by a validated fix ÷ accepted.
03

Agent intervention

"What did the agent change, decide or hide, and who saw it?"

Reviewed intervention rateheals, retries, quarantines and reclassified failures reviewed by a human ÷ total. Plus: reverted agent decisions (your agent's error rate).
04

Residual-risk ownership

"What is still uncovered, and who signed for it?"

Owned residual riskreleases with a residual-risk statement signed by a named person ÷ releases

Score your team

Pick the highest level you can prove today, not the one you're planning for. Each level includes the ones before it.

Level 0ActivityLevel 1SignalsLevel 2EvidenceLevel 3Owned

Your scorecard

Score all four dimensionsThe verdict appears here.

When to scale an agentic workflow

✓
No dimension is below 2Evidence exists for every dimension, not just activity.
✓
Residual risk has a named ownerSomeone signs, every release.
✓
It holds over timeSeveral releases in a row, not one good demo week.

Not there yet? Keep the agent, reduce its autonomy. Let it suggest instead of merge. Let it flag instead of heal on critical journeys. Give autonomy back as the evidence comes in.

Three things to do next week

1
Pair every activity metric with an evidence metricNext to "tests generated", add "critical risks covered by tests seen failing".
2
Plant five faults that would hurt your usersInject them one at a time. Count how many your suite catches. That's your fault detection rate.
3
Add three lines to your next release noteWhat's not covered, why you accept it, and who owns it. A name.

Templates

Fault-injection worksheet

Risk (from your top-10 list):
Fault injected (flag / mock / config):
Expected: which test should turn red?
Result:  [ ] caught  [ ] missed
If missed: owner + date for a new test

Fault detection rate = caught / injected

Application-level faults via feature flags or mocked responses are enough to start. Tools such as PIT, Stryker or mutmut automate the code-level version.

Residual-risk statement

Release: ____________________

Not covered:
Why we accept it:
Owner (a name, not a team):

Agent interventions this release:
  reviewed __ / total __   reverted __

Example: Not covered: rescheduling an appointment across a daylight-saving change. Why: twice a year, low volume, alert on booking-time mismatches.

Why these practices

Sources