All insights
Test management

Test Suite Storage Sizing: How Big Does a Suite Get on Disk?

Test case text barely takes any space. Attachments do, and they grow with every run. Here is the arithmetic to size storage before the volume fills up.

Oct 1, 20266 min read
Test Suite Storage Sizing: How Big Does a Suite Get on Disk? — Tesbo

Nobody puts test suite storage sizing on a sprint plan. It is the kind of infrastructure task that only gets attention once CI starts failing to post results because the disk is full. That usually happens on a Friday afternoon, right before someone wanted to ship. Picture the team that owns a 900 case regression suite: the pipeline goes red at 4pm, not because a test failed, but because the volume has no room left for the log it tried to attach. That team spends the evening deleting old attachments under pressure. This post gives you the arithmetic so you can set a retention policy months earlier instead.

Two things get stored, and only one of them grows on its own

A test management system holds two kinds of data that behave completely differently.

The first is case text: the steps, the expected results, the preconditions. This is bounded by how much your team writes. A suite of 1,400 cases with a few hundred words each might total a few megabytes across the whole project. It grows slowly, when someone adds or edits a case. It plateaus once your product's test coverage stabilizes.

The second is execution records and their attachments. Every time someone, or a CI job, runs a case and logs a result, that creates a row. Often it also creates a screenshot, a log file, or a video. This does not plateau. It grows with every run, forever, for as long as the suite exists. A suite that never adds a single new case still accumulates storage every week it keeps running.

Confusing these two is the root of most surprise capacity problems. Teams budget for the first and get blindsided by the second.

A worked estimate: what a year of one suite actually costs

Assumptions first. These are inputs you should swap for your own numbers, not measured data from any real system.

Take a suite of 1,400 cases, run once a week as a full regression pass. Assume 15% of executions fail or need investigation, and each of those attaches one screenshot averaging 400KB.

  • Executions per run: 1,400
  • Failures needing a screenshot: roughly 210, which is 15% of 1,400
  • Screenshot storage per run: 210 times 400KB, about 84MB
  • Runs per year: 52
  • Screenshot storage per year: 84MB times 52, about 4.4GB

That is one suite, one weekly cadence, a conservative failure rate. Execution record text itself (pass or fail, timestamp, who ran it, a short note) is small. It is often under a kilobyte per execution. So 1,400 times 52 executions a year adds maybe 70MB of text. The attachments dwarf the records that reference them.

Now scale that up. Three product teams, each with a suite this size, each running weekly, puts you past 13GB a year. That is before anyone adds mobile screenshots, video capture on failure, or a second suite for a different platform.

Why CI makes this worse in a way manual testing never did

A manual tester runs a case once, logs a result, and moves on. An automated suite wired into CI can run on every merge to a shared branch. On an active repository that might mean a dozen runs a day instead of one a week. If that pipeline is configured to attach a log or screenshot to every execution, not just failures, the volume math above gets an order of magnitude worse almost immediately.

This is not a flaw in automation. It is a consequence of automation doing exactly what it is supposed to do: run constantly and leave evidence. The evidence is the point. But nobody sized the disk for a pipeline that runs on every merge when they sized it for a QA team that runs regression on Fridays.

Retention is a decision, not an accident

Before you can set a sensible policy, separate two categories of reasons for keeping data.

The first is a real requirement: a regulated industry, a customer contract, or an internal audit standard that specifies how long test evidence has to be retrievable. If that applies to you, the retention period is not a technical choice. It is a compliance one, and it should be documented as such.

The second is habit. Nobody deleted anything because nobody decided to. Now the export folder from three product versions ago is still sitting in the primary volume because deleting it feels risky, even though nobody has opened it in a year.

Most teams, once they actually look, find they are treating the second category like the first. They keep everything indefinitely by default, without a stated reason. That is expensive and gives no one confidence that the right things are protected.

The compromise most teams land on

Once teams do this exercise, they tend to converge on a similar split.

  • Keep execution records (the pass or fail result, the timestamp, the author) indefinitely, because they are small and are the actual audit trail
  • Expire large attachments (screenshots, videos, logs) on a schedule, commonly 90 days to a year depending on the requirement that applies
  • Keep an exception path for attachments tied to a specific incident or a compliance flagged run, so those survive the normal expiry

This keeps the searchable history intact. Who ran what, when, and did it pass, all stay on record, while the part of storage that actually grows unbounded gets capped. It is a reasonable default if nothing in your industry says otherwise, and a starting point to adjust if something does.

Monitor the trend, not the alarm

The failure mode is not usually a surprise that the volume could fill up. It is finding out only once it already had. A threshold alert at 70% or 80% of capacity, checked weekly against the trend from the past month, gives a team weeks of runway to act instead of an hour of scrambling. The useful signal is the slope of the graph, not a single reading.

If you run CI heavily, track attachment growth separately from case text growth. They move at different speeds. A policy that only watches total volume will miss the fact that one part is flat and the other is climbing, the same way that 900 case suite's screenshot folder can double in a quarter while its case text barely moves.

Questions people ask

Does case text really stay small forever?

It grows slowly and predictably as your suite gets bigger. It does not multiply with every run the way attachments do. A 1,400 case suite's text total is roughly the same whether you ran it once this year or fifty times.

What attachment size should I assume if I do not have real numbers yet?

Use a conservative estimate from a similar suite you already run. Or start with something like 300 to 500KB per failure screenshot and revise once you have a few weeks of real data.

Should we delete old execution records too, not just attachments?

Most teams don't, because the records themselves are small and are the part people actually search later. The savings from deleting records are minor compared to attachments.

What if our industry requires us to keep everything for seven years?

Then that is your retention period for whatever is in scope. It becomes a compliance requirement to plan storage against rather than a discretionary choice.

How often should we check the storage trend?

Weekly is enough for most teams. The point is catching a slope change, such as a new CI job attaching video, before it becomes an emergency.

Keep going

Try Tesbo, or get the next useful idea

Start building your testing workflow now, or get one practical email a month.

Start free

One email a month

What we shipped, what we learned, and the occasional infographic worth pinning. Unsubscribe in one click.