# The Real Cost of Test Infrastructure - Playwright and more

Written by Dominik Szahidewicz
Reviewed by Mariusz Wójcik
Published: 2026-09-29
Updated: 2026-09-30

\[TABLE\_OF\_CONTENTS]

Fourteen failures. Same error, every time: TimeoutError: locator.click: waiting for selector "#submit". Your suite was green ten minutes ago, on your machine, watching it run. You push, GitHub Actions picks it up, and now half your regression suite is red.

That’s a test infrastructure problem as much as a test problem: the CI environment, runner resources, worker concurrency, Docker setup, and other execution details change how automated tests behave, how reliable they are, and how easy they are to debug. If you run Playwright in CI/CD and keep seeing tests pass locally but fail in GitHub Actions, this is for the engineers and QA teams trying to find the real cause instead of calling it “just flaky.”

Your first move is probably to hit rerun. It passes. You shrug and move on. That instinct is the expensive part. Not because rerunning is wrong — it’s a reasonable first check — but because the label “[flaky](https://bugbug.io/blog/software-testing/flaky-test/)” is where most teams stop looking. And a test that passes locally and fails on Playwright + GitHub Actions almost never got there by chance. It’s telling you something true about your setup that your laptop was hiding.

So this guide breaks down why CI fails when local runs don’t, how to debug those failures systematically, when to tune worker counts or parallel execution, what Docker changes in the browser environment, how much flaky CI actually costs, and where cloud execution with BugBug can remove part of the infrastructure burden. Finding the cause usually takes longer than the fix itself, and the teams that understand their test infrastructure waste less time, ship faster, and spend less on false failures.

## It's Not Flaky. It's Telling You the Truth.

Your local machine is lying to you, gently. It's warmed up. You're probably already logged in. You're running one spec file, not the full suite. Your development machine likely has more CPU and memory headroom than the runner GitHub hands your workflow.

CI removes every one of those comforts at once, as part of the broader tools and resources used for software testing. Playwright runs your tests in parallel workers by default, and CI usually pushes that parallelism harder than local development ever does. That's when the assumptions your test was quietly making start to show up as failures:

* Two tests sharing the same seeded account, updating the same row at the same time
* A database fixture one worker relies on getting overwritten by another
* An API you call hitting a rate limit only once real concurrency shows up
* Test order mattering — when it was never supposed to

None of this is random. It's test infrastructure exposing hidden assumptions in software development. A CI runner with less CPU, less memory, and noisier disk I/O is often the first environment that tells you the truth about a test that was already fragile. Google's own internal data backs this up at scale: the vast majority of their CI pass-to-fail transitions turn out to be flaky tests exposing real timing assumptions, not actual product bugs.

The practical way to catch this before it ever reaches CI: run the suspect test dozens or hundreds of times in a row locally with --repeat-each=100. If it's going to fail intermittently, this will surface it in minutes instead of letting CI surface it — one confusing, unreproducible failure at a time — over the next two weeks, and better infrastructure improves software quality and speed while increasing confidence in code changes before they reach CI.

## The Two-Hour Debug That's Actually a Two-Week Debug

Here's the honest version of the playbook, because it's genuinely correct — it's just not fast.

**Start with the trace.** [Playwright ](https://bugbug.io/blog/test-automation-tools/playwright-alternative/)generates a trace on the first retry of a failed test in CI. Pull the trace.zip artifact and open it in Trace Viewer. You get a full timeline: DOM snapshots, network calls, console logs, exactly what the page looked like at the moment of failure. The execution engine is what execute tests and report test results through CI artifacts. This alone resolves a large share of CI-only failures without you writing a single line of debug code.

**Reproduce it under load, not in isolation.** A test that fails in a 20-worker CI run but passes when you run it alone locally isn't lying — it's telling you the failure only exists under concurrency. Run your full suite locally with a matching worker count before you conclude anything.

**Use --debug to watch it happen.** The Playwright Inspector lets you step through the test line by line, live, against the actual page. It's slow motion for a problem that only happens at full speed.

**Check the artifacts, not just the log.** CI logs tell you what Playwright thinks happened. In GitHub Actions, which automates testing on every code change, check error messages alongside screenshots, video, and logs. A modal that renders in CI but not locally, an animation that hasn't finished, a layout shift that pushes your target element three pixels to the left — these show up in the video, not the stack trace.

Every step here is correct, documented, and worth doing. The problem isn't the playbook. It's that each round trip through this loop — reproduce, trace, fix, push, wait for CI, confirm - costs a full pipeline run and a context switch, and most fixes need more than one round trip. That's how a two-hour debugging session on paper turns into two weeks of calendar time, spread across a sprint, competing with everything else on your plate. Monitoring test health over time is how teams catch recurring flakiness and preserve trust in the results.

## Docker Was Supposed to Fix Test Environments. It Moved the Problem.

The standard advice for environment parity is to run Playwright inside Docker, using the official Microsoft image, so your local environment and your CI runner use the identical browser build and OS libraries. Docker also helps by virtualizing the environment without needing physical devices, which makes it easier to keep test environments close to production environments. This is good advice. It's also not free.

The most common failure signature: your package.json pins one Playwright version, your Docker image tag pins another, and the browser binary path Playwright expects simply isn't there. The tempting fix — running playwright install inside the container to patch over the mismatch - doesn't resolve the drift. It hides it, and usually doubles your image size in the process.

The second common trap is a minimal base image. Alpine-based Node images are missing the glibc-dependent shared libraries Chromium needs to launch at all — libnss3, libatk1.0-0, libgbm1, and a long tail of others. The symptom isn't a test failure. It's the browser process dying on launch with an error naming a missing .so file, which tells you nothing about your app and everything about your base image choice.

Getting this right the first time is a solved problem: use the official mcr.microsoft.com/playwright image, keep the image tag and the npm package version in lockstep, and pin both deliberately instead of letting one drift on a dependency bot update, with that setup maintained as infrastructure as code so configuration changes are tracked. Staying right is the part nobody schedules. If you already have an existing workflow, check the github actions workflow file before changing versions so the container, CLI, and runner stay aligned. Falling more than a couple of minor versions behind reintroduces the exact drift you built Docker to prevent - and most teams budget zero ongoing time for this, then spend roughly half a day a month on it, indefinitely. Cloud-based testing environments can be provisioned quickly, but they still need the same version discipline.

## How Many Workers Should You Actually Use to Run Playwright Tests?

This is the question that gets guessed at more than it gets calculated, so here's a starting point that accounts for actual CI resource classes:

* **Under 50 tests, any runner:** leave workers unset. The default is fine.
* **50–200 tests on a 2 vCPU runner:** 2 workers.
* **50–200 tests on a 4 vCPU runner:** 3–4 workers.
* **Over 200 tests:** tune worker count before reaching for sharding.
* **New flakiness right after raising worker count:** drop it back by 1–2 and recheck before assuming it's a test problem.

The mechanical reason this matters: a 2 vCPU runner asked to run 4 workers is launching four separate Chrome instances on two cores. That's not a Playwright bug or a flaky test — it's four browsers fighting over CPU time, and the symptom is exactly the kind of intermittent timeout that gets misdiagnosed as test flakiness. Parallel runs also work best when you isolate test environments so workers do not share state.

There's a second, less obvious lever: fullyParallel. By default, Playwright runs every test inside one spec file on the same worker, one after another - even on a machine with plenty of workers to spare. If your files each contain ten tests, that's ten tests queuing behind each other on a single worker while your other three workers sit idle. Setting fullyParallel: true changes the unit of scheduling from "file" to "individual test," which is usually the difference between a suite that scales with your runner size and one that doesn't, delivering faster execution while helping surface potential failures under various load conditions. In cloud infrastructure testing, tuning concurrency is also part of identifying failure points before they turn into recurring CI issues across different load conditions.

Getting this dialed in for one repo is a Tuesday afternoon. Getting it right for every repo, every time your test count grows past the next threshold, every time you change runner size - that's the part that turns into a recurring line item nobody assigned to anyone.

## What This Actually Costs, in Money

Here's where "it's just CI flakiness" stops being a shrug and starts being a budget line, because regular infrastructure testing can also create **cost savings** by preventing failures.

| Team size     | Est. flaky/CI-debug cost per year | Where it comes from                                                                         |
| ------------- | --------------------------------- | ------------------------------------------------------------------------------------------- |
| 20 engineers  | \~$235,000                        | 15% flake rate — reruns, investigation, deploy delays                                       |
| 50 engineers  | $200,000–$400,000                 | 5–10 hrs/week per engineer in investigation + context switching                             |
| 100 engineers | \~$2.5M                           | 6–8 hrs/week per engineer dealing with flaky failures (Spotify Engineering survey baseline) |

The developer-time component is what dominates every version of this math, not the CI compute bill. The median time to triage a single flaky failure is around 28 minutes. Add another 15–25 minutes for the engineer to get back to the depth of focus they were at before the interruption, and one "quick CI check" costs closer to an hour than five minutes - repeated every time it happens, for every engineer it happens to. That is exactly where **proper test infrastructure** starts to **improve efficiency**: it improves system reliability and performance, while a **robust test infrastructure** removes avoidable interruptions before they turn into delivery bottlenecks.

None of this shows up as a line item anywhere. It's not in your GitHub Actions bill. It's not in your Playwright license, because there isn't one. It's just quietly gone from your sprint, a little at a time, until someone adds it up, though cloud test infrastructure can also reduce waste by allocating resources more efficiently.

## The GitHub Actions Bill Doesn't Sit Still Either

Even the infrastructure math keeps moving. GitHub cut hosted-runner prices by up to 39% starting January 2026 - and in the same announcement, added a new $0.002-per-minute charge for self-hosted runner usage on private repos, effective that March. The backlash was immediate enough that GitHub paused the self-hosted charge within about a week of announcing it. It's not cancelled — GitHub says they're re-evaluating it with more input first — just not active right now. Those choices also shape your compliance requirements and exposure to security concerns, not just spend.

That's the part worth sitting with: the team that built their own runner fleet specifically to control costs found out their pricing model could change with a single blog post, twice in one quarter. Owning the infrastructure didn't remove the variable. It just meant the variable now shows up on your invoice instead of someone else's. That also affects how well teams keep up with industry regulations.

The compute reality underneath adds its own quiet tax. A single Playwright container idles at 500MB–1GB of memory and can climb past 2GB under load. Ten parallel browsers in CI is a completely normal setup - and it's also usually the first time a team discovers their runners are undersized, because the failure mode is an OOM kill, not a helpful error message. In practice, infrastructure testing helps surface security vulnerabilities and other weak points before they reach live systems, while improving adherence to compliance requirements.

## Own the Coverage. Not the Testing Infrastructure.

None of the above is a reason to avoid Playwright. It's an excellent framework, and if your team has the engineering bandwidth to own worker tuning, Docker parity, and a runner bill that changes on GitHub's schedule instead of yours - along with deployment and configuration automation - the playbook in this article is the right one: follow it.

Most teams in the 20–150 person range don't have that bandwidth sitting around unclaimed, though. That's the gap BugBug's cloud execution is built for. Test runs get up to 32x parallel without anyone on your team tuning worker counts per repo, choosing a runner size, or figuring out which vCPU tier stops timing out. There's no Docker image to keep in version lockstep with your test runner, because there's no Docker image to manage at all. You still decide what to test and how the flow should behave, including realistic test data setup through test data management that generates realistic databases for testing purposes - you just stop being the one who finds out at 2am that a runner OOM'd during a critical release.

Worth saying plainly: BugBug runs on Chromium-based browsers only, so if your team needs cross-browser coverage across Firefox or Safari as part of this workflow, that's a real limitation, not a footnote.

**If your team wants to keep building and owning this layer**, the guidance above - trace-first debugging, matched worker counts, a pinned Docker image - will get you there. Teams that use Playwright with GitHub Actions to run secure workflows from a github repository usually lean on the CLI introduced in Playwright version 1.8.0 and GitHub Secrets to protect API keys. In practice, that often means using npm init to bootstrap setup, installing playwright browsers, validating a playwright test locally before you run playwright tests in automation, and wiring checks to each pull request for continuous integration and continuous delivery. It's real work, and it's worth budgeting for honestly instead of discovering it a sprint at a time.

**If nobody on your team signed up to own CI infrastructure**, that's exactly the job BugBug's free plan is built to take off your hands - no credit card, first test running in under ten minutes.

For teams weighing this against a code-based framework more broadly, our breakdown of [automation testing for startups](https://bugbug.io/blog/test-automation/automation-testing-guide-for-startups-level-1/) covers the tradeoff in more depth, and if Cypress is the other framework on your shortlist, our [Cypress alternatives guide](https://bugbug.io/blog/test-automation-tools/cypress-alternatives/) puts it side by side with Playwright and a handful of others.

`<custom-cta-card><json>{"name":"cta-card","children":[{"name":"cta-card-title","children":[{"name":"paragraph","children":[{"attributes":{"bold":true},"data":"Automate your tests for free"}]},{"name":"paragraph"}]},{"name":"cta-card-details","children":[{"name":"paragraph","children":[{"data":"Test easier than ever with BugBug test recorder. Faster than coding. Free forever."}]},{"name":"paragraph"}]},{"name":"cta-card-cta","children":[{"name":"paragraph","children":[{"attributes":{"bold":true},"data":"Get started"}]},{"name":"paragraph"}]}]}</json>  <cta-card-title-md>  **Automate your tests for free**  </cta-card-title-md>  <cta-card-details-md>  Test easier than ever with BugBug test recorder. Faster than coding. Free forever.  </cta-card-details-md>  <cta-card-cta-md>  **Get started**  </cta-card-cta-md>  </custom-cta-card>`

Happy (automated) testing!
