Estimating QA Effort for AI Features You Can't Predict
AI features break test-count estimates because output changes every run. Here is a method built on behaviours instead, with a worked example.

A QA lead gets asked the same question every sprint planning: how many days will testing take? For a login form or a checkout flow, that question has an answer, because the team can count screens, count states, and count test cases against them. Estimating QA effort for AI features is a different problem, and it gets asked in the same confident tone. The honest answer is that nobody in the room actually knows yet, because the thing being tested does not produce the same output twice. A three day sprint estimate gets typed into Jira anyway, and two weeks later the team is still arguing about whether a reworded but correct answer counts as a pass. This post lays out a method that avoids that fight, and shows the arithmetic behind a real estimate.
Why counting test cases breaks down for non-deterministic output
A traditional test case assumes a fixed expected result. Click this button, see this screen, get this total. That assumption is what makes a test count meaningful. Forty test cases roughly means forty units of verification work, and a QA lead can multiply by an average minutes-per-case figure to get a schedule.
An AI feature does not hold still. Ask a support summarizer the same ticket twice and the wording shifts even when the underlying content is right. Ask a recommendation feature for the same user twice and the ranked list can reorder without being wrong.
If a QA lead tries to write "exact output equals X" test cases against that kind of feature, most of them fail on a rerun. The failures have nothing to do with a real bug. The team either drowns in false failures and starts ignoring red builds, or gives up on automated checks entirely and falls back to spot checking by hand before every release. Neither outcome helps anyone estimate effort, and neither survives a second sprint.
The unit that actually works: behaviours you are willing to assert on
The fix is to stop counting test cases and start counting behaviours, see what to automate first for more on scoping decisions like this. A behaviour is a specific thing the feature must or must not do, stated narrowly enough that two engineers would agree whether it happened, regardless of the exact wording the model produced.
Each behaviour then gets assigned one of two assertion types:
- An exact assertion, for anything with a fixed correct answer regardless of wording, such as the summarizer must not invent a ticket number that was not in the input
- A scored rubric, for anything where correctness is a matter of degree, such as the summary must capture the customer's actual complaint, scored on a 1 to 5 scale by a reviewer or a grading prompt
This reframes the estimate. Instead of "how many tests," the question becomes "how many behaviours are we willing to name and stand behind." That is a question a team can actually answer in a room. It is a scoping conversation, not a forecasting one, and it produces a list the whole team signed off on rather than a guess one person defends alone.
A worked estimate, from behaviour list to effort
Take a real shape of feature: an AI assistant that drafts replies to support tickets before a human sends them. A QA lead sits down with the product manager and the tech lead for 45 minutes and lists the behaviours that matter for launch.
- The draft must not fabricate an order number, refund amount, or policy that was not in the ticket or the knowledge base. This gets 6 test scenarios covering different fabrication risks, and it is an exact assertion
- The draft must match the tone setting the agent selected, formal or casual. This gets 2 scenarios, also exact
- The draft must address the customer's actual question, not a related but different one. This is a scored rubric, with 10 scenarios pulled from real past tickets, each graded 1 to 5 by a reviewer
- The draft must not recommend a competitor product. Exact assertion, 3 scenarios
- The draft must degrade safely when the knowledge base has no relevant article, by saying so rather than guessing. Exact assertion, 4 scenarios
That is 25 scenarios across 5 behaviours.
Exact assertions run fast. Writing and wiring one up, including edge cases, takes roughly 20 minutes. The 15 exact scenarios cost about 5 hours in total.
Scored rubric scenarios take longer. Someone has to define the rubric and calibrate it against 2 or 3 example gradings before it is trustworthy. Call it 45 minutes each for the 10 rubric scenarios, which adds another 7.5 hours.
Add half a day, about 4 hours, for a reviewer to run an initial calibration pass and confirm the rubric produces consistent scores across two graders. The total lands close to 16.5 hours, roughly two working days. That number came from a list the team can point to line by line, not a multiplier pulled from memory or an industry benchmark nobody can trace.
What to tell a product manager who wants a number today
Sometimes the behaviour list does not exist yet and the product manager wants an estimate anyway. This usually happens before the meeting where the behaviours would even get discussed, in a roadmap review where every feature needs a number attached.
The honest move is not to invent a placeholder number that will anchor expectations wrongly. It is to say what is actually true. The estimate depends on how many behaviours the team is willing to name, and that conversation takes about an hour, not a guess typed into a spreadsheet cell.
Offer a range instead of silence, grounded in the shape of the feature rather than an industry figure.
- A small AI feature with a narrow job, like classifying ticket urgency into three buckets, usually settles around 8 to 12 behaviours
- A broad feature like the reply drafter above tends to land at 20 to 30 behaviours
- A feature that touches money or safety, like an assistant that can issue refunds, tends to need more exact assertions relative to scored rubrics, which pushes effort up even at a similar behaviour count
That range is not a benchmark borrowed from a report. It is a pattern from having run the behaviour-listing exercise before, and it should be stated as exactly that: a pattern, not a promise.
A second worked case, to show the method holds up
A smaller example makes the pattern easier to trust. Consider a feature that auto-tags incoming bug reports with a severity level, so triage does not wait on a human every time.
The team lists four behaviours for this feature. The tag must never mark a security related report as low severity. The tag must match the severity a human triager would assign, most of the time. The feature must flag reports it is unsure about instead of guessing. And the feature must not process reports written in a language it was not trained to handle.
The security behaviour is an exact assertion with 5 scenarios, because there is no ambiguity about whether a report mentioning a credential leak got marked low. The severity matching behaviour is a scored rubric with 15 scenarios pulled from a backlog of real past reports, each compared against how a human actually triaged it.
The uncertainty flag is an exact assertion with 3 scenarios. The unsupported language case is an exact assertion with 2 scenarios.
That is 25 scenarios again, but weighted differently. There is less exact assertion work and more rubric-heavy work, because severity matching is inherently graded. The rubric scenarios alone run about 11 hours once calibration is included, more than double the reply drafter example's rubric cost, even though the total scenario count is identical. The lesson is that behaviour count alone is not the whole estimate. The mix between exact and scored assertions moves the number too.
Where this connects to the rest of your QA process
This method does not replace the rest of QA effort estimation. It fixes the one part that breaks for non-deterministic features. The team still needs to decide what to automate first once the behaviour list exists, since not every behaviour needs the same rerun frequency on every build. Some, like the fabrication check, are worth running on every commit. Others, like a broad rubric-scored tone check, might only need a run once a week.
Once testing is underway, the QA metrics that actually mean something will tell you whether the behaviour list was the right scope, or whether production is surfacing failures the list missed entirely. A behaviour list is a starting estimate, not a permanent contract, and it should get revised the first time a real user finds something the list did not anticipate.
Questions people ask
Can this method work for features that are not AI based?
Yes, though it earns its keep most clearly on non-deterministic features. For deterministic features, plain test case counting still works fine.
How do we keep a scored rubric from becoming subjective?
Write the rubric with 2 or 3 concrete example outputs and their score before anyone grades a real result. Then have a second reviewer check a sample to confirm agreement.
What if the team cannot agree on which behaviours matter?
That disagreement is useful information on its own. It usually means the feature's scope is not settled yet, and testing effort cannot be estimated until scope is.
Does every behaviour need its own automated check?
No. Some behaviours are cheap to check by hand on every release and expensive to automate reliably, especially scored ones. Decide that per behaviour, not as a blanket policy.
How often should the behaviour list be revisited?
Whenever the feature's prompt, model, or scope changes meaningfully. A stale behaviour list is a common reason teams stop trusting their own test results.
Try Tesbo, or get the next useful idea
Start building your testing workflow now, or get one practical email a month.
Start freeOne email a month
What we shipped, what we learned, and the occasional infographic worth pinning. Unsubscribe in one click.


