The Agentic QE Scorecard
Is the agent improving your quality, or improving your dashboard? Four questions, scored 0 to 3, that a release manager can ask without reading a single test.
Activity is not evidence
Keep your activity metrics: they tell you about efficiency. Pair each one with an evidence metric that tells you whether you can ship.
| Measures agent activity | Measures release confidence |
|---|---|
| Tests generated | Critical risks covered by tests you have seen fail |
| Coverage % | Injected faults the suite actually catches |
| Flaky failures healed | Heals reviewed and classified by a human |
| Triage time | Root causes verified by reproduction |
| Pipeline pass rate | Residual risk stated, with a named owner |
Four dimensions, four questions
Risk coverage
"Are we testing what would hurt, or what was easy to generate?"
Evidence quality
"Would this evidence have turned red if we were wrong?"
Agent intervention
"What did the agent change, decide or hide, and who saw it?"
Residual-risk ownership
"What is still uncovered, and who signed for it?"
Score your team
Pick the highest level you can prove today, not the one you're planning for. Each level includes the ones before it.
Your scorecard
When to scale an agentic workflow
Not there yet? Keep the agent, reduce its autonomy. Let it suggest instead of merge. Let it flag instead of heal on critical journeys. Give autonomy back as the evidence comes in.
Three things to do next week
Templates
Fault-injection worksheet
Risk (from your top-10 list): Fault injected (flag / mock / config): Expected: which test should turn red? Result: [ ] caught [ ] missed If missed: owner + date for a new test Fault detection rate = caught / injected
Application-level faults via feature flags or mocked responses are enough to start. Tools such as PIT, Stryker or mutmut automate the code-level version.
Residual-risk statement
Release: ____________________ Not covered: Why we accept it: Owner (a name, not a team): Agent interventions this release: reviewed __ / total __ reverted __
Example: Not covered: rescheduling an appointment across a daylight-saving change. Why: twice a year, low volume, alert on booking-time mismatches.
Why these practices
- Coverage is a weak proxy. Across 31,000 test suites for five large Java systems, coverage showed only a low to moderate correlation with effectiveness once suite size was controlled (Inozemtseva & Holmes, 2014).
- Injected faults are a valid stand-in for real ones. Mutant detection correlates with real-fault detection independently of coverage (Just et al., 2014). At Google, developers exposed to mutants wrote more and better tests, and mutants were coupled with real faults (Petrović et al., 2021).
- Generated tests need filters. Meta shipped LLM-generated tests only after they cleared filters proving measurable improvement: 75% built, 57% passed reliably, 25% increased coverage (Alshahwan et al., 2024).
- Confident is not correct. RCACopilot reached a micro-F1 of 0.766 on 653 real incidents: excellent, and still about one category in four wrong (Chen et al., 2024).
- Review must be designed in. Automation bias affects novices and experts and isn't prevented by training or instructions alone (Parasuraman & Manzey, 2010).
- Accountability needs names. The NIST AI RMF (GOVERN) asks for documented roles and accountable individuals for AI risk decisions.
- AI amplifies what's there. DORA 2024: a 25% increase in AI adoption was associated with an estimated 7.2% drop in delivery stability. DORA 2025: throughput now improves, instability still rises.
Sources
- DORA, Accelerate State of DevOps Report 2024: cloud.google.com
- DORA, 2025 State of AI-assisted Software Development: PDF
- M. Strathern (1997), paraphrasing C. Goodhart (1975): Goodhart's law
- L. Inozemtseva, R. Holmes, "Coverage Is Not Strongly Correlated with Test Suite Effectiveness", ICSE 2014: PDF · ICSE Most Influential Paper 2024
- R. Just et al., "Are Mutants a Valid Substitute for Real Faults in Software Testing?", FSE 2014: abstract
- G. Petrović et al., "Does Mutation Testing Improve Testing Practices?", ICSE 2021: arXiv
- N. Alshahwan et al., "Automated Unit Test Improvement using Large Language Models at Meta", FSE 2024: arXiv
- Y. Chen et al., "Automatic Root Cause Analysis via Large Language Models for Cloud Incidents", EuroSys 2024: arXiv
- R. Parasuraman, D. Manzey, "Complacency and Bias in Human Use of Automation", Human Factors 52(3), 2010: record
- NIST, AI Risk Management Framework 1.0, 2023: nist.gov