Anastasiia Sokolinska

Written by: Chief Operating Officer

Anastasiia Sokolinska

Posted: 08.10.2026

15 min read

Run a sharded Playwright suite in CI, upload playwright-report as an artifact, and the published report will describe one shard out of four. The rest never left their runners.

That failure survives for weeks in most teams, because a Playwright test report sitting inside a zip file behind a CI login is a report nobody opens.

Reports are cheap to generate and expensive to distribute. Here is how to configure reporters that hold up under sharding and retries, write a custom reporter that does not die silently, and publish results without exporting your staging environment to the public internet.

What a Playwright report contains, and what it forgets

A report is a per-run artifact. It carries pass, fail, skip and timeout statuses per test, a separate result object per attempt, step-level timings, and whatever the run attached along the way: screenshots, video, and traces you can open in Trace Viewer. The reporter reference covers the full surface.

The part that shapes every decision below: the built-in HTML report is stateless. It describes one run and knows nothing about the previous fifty. Flake rate, duration trend, "this test has failed on main four times this month" — none of that exists inside the artifact. Every signal a team actually manages the suite with has to come from somewhere you built or bought.

Book a call for trusted Playwright test automation expertise

Pick reporters by consumer, not by feature list

Playwright ships eight built-in reporters, and the useful way to choose is to ask who reads the output rather than which one has the most features.

Reporter
Consumer
Use when
Avoid when

list

Engineer watching the run

Local development, streaming CI logs

You need machine-readable output

line

Engineer running a large suite

Hundreds of tests locally, compact output

You want per-test visibility

dot

Nobody, it is a log-size optimization

CI logs that would otherwise scroll for pages

Debugging a specific failure

html

The person debugging a failure

Local runs, merged output after sharding

Running inside shards, output collides

json

Anything the team writes itself

Feeding a script, dashboard, or database

A human needs to read it directly

junit

The CI gate and test-management tooling

Jenkins, GitLab, TestRail-style ingestion

You need attachments or step detail

blob

The merge-reports command

Any sharded run

Single-worker runs, adds a merge step

github

The developer who broke the build

GitHub Actions, inline PR annotations

Any other CI platform

Production suites run at least two: one a human reads, one a machine reads. And the reporter array branches on process.env.CI, because the reporter that helps you locally is the one that hangs the pipeline.

Reporter config that does not break in CI

Four failure modes cost teams a pipeline run, and all four live in configuration rather than in test code.

open: 'always' on the HTML reporter tries to spawn a browser on a headless runner. JUnit output paths differ per CI platform, so drive the path from an environment variable instead of hardcoding it. Artifact upload steps need to run even when the test command exits non-zero, which means if: ${{ !cancelled() }} on every upload.

The fourth one is less obvious. Reporters run in the same Node process as the test run, so any blocking I/O inside a reporter adds directly to wall-clock CI time. A reporter that writes to a database synchronously per test is a tax on every run.

// playwright.config.ts
import { defineConfig, devices } from '@playwright/test';

const isCI = Boolean(process.env.CI);

export default defineConfig({
  testDir: './tests',
  fullyParallel: true,
  forbidOnly: isCI,
  retries: isCI ? 2 : 0,
  workers: isCI ? 4 : undefined,
  reporter: isCI
    ? [
        ['blob'],
        ['junit', { outputFile: process.env.JUNIT_OUTPUT_PATH ?? 'results/junit.xml' }],
        ['github'],
      ]
    : [
        ['list'],
        ['html', { open: 'never' }],
      ],
  use: {
    baseURL: process.env.BASE_URL ?? 'http://localhost:3000',
    trace: 'on-first-retry',
    screenshot: 'only-on-failure',
    video: 'retain-on-failure',
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
  ],
});

One detail that catches teams later: under sharding, the JUnit reporter writes one file per shard, not one file for the run. Name them per shard through the env var and let the ingesting system merge, or the last upload overwrites the rest.

Sharding breaks the report, and blob plus merge-reports fixes it

Here is the mechanism. With a matrix of four shards, each shard runs in its own isolated runner and writes its own playwright-report directory. There is no single report at the end of that. There are four partial ones, and a naive artifact upload either collides on the artifact name or leaves you with four unrelated reports describing a quarter of the suite each.

The blob reporter exists for exactly this. It writes a zipped intermediate format per shard, and merge-reports reassembles them into a single report afterwards. Blob output carries the attachments too, so the merged HTML report is the full report with traces and screenshots intact, not a summary of counts. Playwright's sharding documentation covers the reference workflow.

The sequence:

1.

Switch the reporter to blob when process.env.CI is set.

2.

Run each shard with --shard=N/4 inside a matrix job.

3.

Upload each shard's blob-report directory as a uniquely named artifact.

4.

Add a dependent merge job with needs: [test] and if: !cancelled().

5.

Download every shard artifact into one directory, for example all-blob-reports.

6.

Run npx playwright merge-reports --reporter html ./all-blob-reports, then upload the output.

The merge has to be a separate job because shards run in parallel in isolated environments, and it has to run after all of them including the ones that failed. fail-fast: false on the matrix and if: ${{ !cancelled() }} on the merge job are what make that true. Skip either and your report goes missing on exactly the runs where you needed it.

# .github/workflows/e2e.yml
name: e2e

on:
  push:
    branches: [main]
  pull_request:

jobs:
  test:
    runs-on: ubuntu-latest
    timeout-minutes: 30
    strategy:
      fail-fast: false
      matrix:
        shard: [1, 2, 3, 4]
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - name: Run shard ${{ matrix.shard }}
        run: npx playwright test --shard=${{ matrix.shard }}/4
        env:
          JUNIT_OUTPUT_PATH: results/junit-${{ matrix.shard }}.xml
      - name: Upload blob report
        if: ${{ !cancelled() }}
        uses: actions/upload-artifact@v4
        with:
          name: blob-report-${{ matrix.shard }}
          path: blob-report
          retention-days: 1

  merge-report:
    needs: [test]
    if: ${{ !cancelled() }}
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
          cache: npm
      - run: npm ci
      - name: Download blob reports
        uses: actions/download-artifact@v4
        with:
          path: all-blob-reports
          pattern: blob-report-*
          merge-multiple: true
      - name: Merge into a single HTML report
        run: npx playwright merge-reports --reporter html ./all-blob-reports
      - name: Upload merged report
        if: ${{ !cancelled() }}
        uses: actions/upload-artifact@v4
        with:
          name: playwright-report
          path: playwright-report
          retention-days: 14

Shard-level blob artifacts get a one-day retention because they are disposable the moment the merge succeeds. How you split work across shards and workers is a separate decision with its own trade-offs covered in sharded across parallel workers.

Book a call to tap into proven Playwright automation expertise

Write a custom reporter that survives production

The Reporter interface is small, and most of the useful work happens in three hooks: onBegin for run metadata, onTestEnd for per-attempt results, onEnd for the summary and any network call. onStepBegin and onStepEnd fire constantly, and implementing them is the fastest way to make a reporter expensive.

Four things separate a reporter that works from one that looks like it works.

onTestEnd fires once per attempt, not once per test. A counter that ignores result.retry reports inflated failure numbers on any suite with retries enabled. Flaky means a test that failed an earlier attempt and passed a later one, and test.outcome() already computes that for you. Getting this wrong is how a suite ends up with a failure rate nobody trusts distinguishing flaky from broken.

Playwright swallows exceptions thrown inside reporter methods. A reporter that dies on a network timeout looks identical to one that works: the run passes, the pipeline goes green, and the dashboard stops updating. Wrap every I/O call in try/catch and write failures to stderr, or you will find out weeks later.

printsToStdio() returning false tells Playwright your reporter is not the terminal output, so it attaches a default terminal reporter alongside. Return true by accident and engineers watching CI logs see nothing.

onEnd can be async, which is where a webhook or metrics post belongs. Give it a timeout. A reporter waiting on an unresponsive internal endpoint holds the whole run open.

// reporters/suite-health-reporter.ts
import type {
  FullConfig, FullResult, Reporter, Suite, TestCase, TestResult,
} from '@playwright/test/reporter';

interface RunSummary {
  runId: string;
  status: FullResult['status'];
  durationMs: number;
  expected: number;
  unexpected: number;
  flaky: number;
  skipped: number;
  failures: string[];
}

export default class SuiteHealthReporter implements Reporter {
  private rootSuite!: Suite;
  private startedAt = 0;
  private readonly failures: string[] = [];
  private readonly endpoint = process.env.QA_METRICS_URL;

  // Keep Playwright's terminal reporter attached.
  printsToStdio(): boolean {
    return false;
  }

  onBegin(_config: FullConfig, suite: Suite): void {
    this.rootSuite = suite;
    this.startedAt = Date.now();
  }

  onTestEnd(test: TestCase, result: TestResult): void {
    // Fires per attempt. Ignore everything except the final one.
    const isLastAttempt = result.retry === test.retries || result.status === 'passed';
    if (!isLastAttempt) return;
    if (test.outcome() === 'unexpected') {
      this.failures.push(test.titlePath().slice(1).join(' > '));
    }
  }

  async onEnd(result: FullResult): Promise<void> {
    const tests = this.rootSuite.allTests();
    const summary: RunSummary = {
      runId: process.env.GITHUB_RUN_ID ?? 'local',
      status: result.status,
      durationMs: Date.now() - this.startedAt,
      expected: tests.filter((t) => t.outcome() === 'expected').length,
      unexpected: tests.filter((t) => t.outcome() === 'unexpected').length,
      flaky: tests.filter((t) => t.outcome() === 'flaky').length,
      skipped: tests.filter((t) => t.outcome() === 'skipped').length,
      failures: this.failures.slice(0, 25),
    };

    if (!this.endpoint) return;

    try {
      await fetch(this.endpoint, {
        method: 'POST',
        headers: { 'content-type': 'application/json' },
        body: JSON.stringify(summary),
        signal: AbortSignal.timeout(5_000),
      });
    } catch (error) {
      // Playwright discards reporter exceptions. Surface it or lose it.
      process.stderr.write(`[suite-health] metrics post failed: ${String(error)}\n`);
    }
  }
}

One more option most guides miss: merge-reports accepts a reporter path, so npx playwright merge-reports --reporter ./reporters/suite-health-reporter.ts ./all-blob-reports applies a custom reporter to merged data after the fact. That keeps the reporter out of every shard's hot path and gives it a complete picture instead of a quarter of one.

The organizational point matters more than the code. A custom reporter is the right tool when results need to reach a system your team already runs: an internal dashboard, a Postgres table, a Slack channel. It is the wrong tool for building a reporting UI from scratch, which is a multi-year maintenance commitment sold to a lead as a weekend project.

Third-party reporters and where trend data has to live

Two open-source options are worth naming. Allure via allure-playwright gives you severity labels, test-case metadata, and trend charts. Monocart gives a single-file HTML report with code coverage and a trend callback.

Compare them on state rather than on feature checklists. Allure's trend charts require the history folder from previous runs to be restored before the run and persisted after it. Monocart takes a trend callback that reads previous report data from wherever you decided to keep it. Both are the same admission: history is an external state, and someone on your team has to own the bucket it lives in.

There is an operational cost worth checking before adoption. Several established reporters in this space expect a Java or Docker runtime for report generation, which is a real addition to a Node-only pipeline that currently installs one thing.

That leaves three honest paths for cross-run analytics. Persist history alongside an open-source reporter, and accept that the persistence layer is yours to maintain. Build storage plus a consumer around the JSON or blob output, and accept the schema drift when Playwright changes. Or buy a platform, and accept that your results now live somewhere you do not control. Option (a) is the cheapest, and it is not free.

Publishing from CI: access, retention, and what it costs

Most guides end at "upload as an artifact" or "push to GitHub Pages." Both deserve more scrutiny than they get.

Artifact-only is the default and it works for small teams. The friction is that downloading a zip and running npx playwright show-report is enough of a barrier that nobody outside the QA team ever looks at a report. Two clicks of friction is the difference between a report that informs releases and one that exists.

GitHub Pages is public by default. For a fintech or healthcare product, publishing there is a data-exposure decision, not a convenience decision. Treat it as one.

Private object storage is where most mid-market teams land: an S3-compatible bucket with public access blocked, a per-run prefix so runs do not overwrite each other, and a lifecycle rule that expires objects on a schedule.

- name: Publish report to a private bucket
  if: ${{ !cancelled() }}
  env:
    AWS_REGION: eu-central-1
    REPORT_PREFIX: >-
      playwright/${{ github.repository }}/${{ github.run_id }}/${{ github.run_attempt }}
  run: |
    aws s3 sync ./playwright-report "s3://qa-reports-private/${REPORT_PREFIX}" \
      --only-show-errors \
      --cache-control "public, max-age=300"
    echo "Report: https://reports.internal.example.com/${REPORT_PREFIX}/index.html" \
      >> "$GITHUB_STEP_SUMMARY"

The prefix includes run_attempt, so a re-run does not silently overwrite the evidence from the failure that triggered it. The link goes into the step summary because a URL in the CI output is the only thing anyone reliably clicks.

What a published report actually exposes

Playwright's CI documentation warns that traces and reports capture credentials, tokens, request headers, DOM snapshots, and application source. A published report is an exported copy of what your staging environment was serving at the moment the test ran.

Retention has to be deliberate on both sides. Report retention and raw-results retention must match, or you end up with a live report whose trace links point at artifacts that expired a month earlier.

Storage cost, roughly

A trace for a multi-step test commonly lands in the 10 to 50 MB range. Set trace: 'on' across a few hundred tests and a single run produces gigabytes, at which point upload time becomes a measurable share of pipeline duration rather than a rounding error.

retain-on-failure and on-first-retry remove that cost from passing tests entirely, which on a healthy suite is most of them. Seven to 14 days of retention covers the window in which anyone actually investigates a failure. Anything past 30 days needs a compliance reason, not a habit. The same arithmetic applies to runner minutes cost of running the suite.

If your suite runs green and your reporting pipeline still consists of an artifact nobody downloads, that gap is usually a CI ownership problem rather than a testing problem.

Route results to the people who act on them

Three consumers, three formats, and no reason for them to share one artifact.

The engineer who broke the build needs failing test names and a link, inside the pull request. The github reporter writes inline annotations; anything more specific comes from JSON plus a comment step.

The person debugging needs the merged HTML report with traces attached, which is the only artifact where the full detail lives.

The lead tracking suite health needs flake rate and duration trend from persisted history, reviewed weekly. Not per run. Per-run trend data is noise.

Teams that publish one report for everyone end up with a report nobody reads. The four-shards-one-report problem at the top of this article is that same failure with a technical cause attached: when nobody has a reason to open the report, nobody notices it has been wrong since March.

Reporting mistakes that cost a debugging cycle

  • No blob reporter under sharding. Your published report covers one shard and understates coverage by 75%.

  • open: 'always' in CI. The HTML reporter tries to spawn a browser on a headless runner and hangs the job.

  • Mismatched JUnit paths across CI platforms. The gate reads a file that is not there and passes a broken build.

  • A reporter that throws. Playwright swallows it, the run stays green, and your dashboard quietly stops updating.

  • Retry-blind counters. Every retry counts as a failure, flake rate inflates, and the suite loses credibility.

  • trace: 'on' everywhere. Gigabytes per run, longer uploads, and traces for tests that passed.

  • Report retention outlasting artifact retention. The report loads, every trace link 404s.

  • Publishing to a public host from a regulated product. Tokens, headers, and DOM snapshots on the open internet.

  • Uploads without if: ${{ !cancelled() }}. The report disappears on exactly the runs that failed.

FAQ

How do I view an HTML report generated in CI?

Download the playwright-report artifact and run npx playwright show-report against the extracted folder. Opening index.html directly from the filesystem does not work, because the report fetches its data over HTTP. For anything past a two-person team, publish to private object storage and share a link instead.

What is the blob reporter for?

It writes a zipped intermediate format that merge-reports reassembles into a single report. It exists for sharded and multi-machine runs, where each shard would otherwise produce its own partial report. Blob output includes attachments, so merged HTML retains traces and screenshots.

Can I run multiple reporters at once?

Yes, and you should. Pass an array to reporter in playwright.config.ts, with tuple entries for reporters that take options. The standard split is one human-readable reporter and one machine-readable reporter, branching on process.env.CI.

How do I merge reports from sharded runs?

Set reporter: 'blob' on CI, upload each shard's blob-report directory under a unique artifact name, then add a dependent job that downloads all of them into one folder and runs npx playwright merge-reports --reporter html ./all-blob-reports. Guard the merge job with if: ${{ !cancelled() }}.

The report is the easy part

Generating a Playwright report takes one line of config. Deciding who reads it, where it lives, who can reach it, and how long it survives is the work, and it is the part that determines whether the suite influences a single release decision.

Clean test suite, but reports nobody opens? DeviQA reviews your Playwright config, sharding setup, and CI reporting pipeline — book a consultation for a written assessment of what breaks and where to fix it.

Book a strategic QA consultation

Anastasiia Sokolinska

About the author

Anastasiia Sokolinska

Chief Operating Officer

Anastasiia Sokolinska is the Chief Operating Officer at DeviQA, responsible for operational strategy, delivery performance, and scaling QA services for complex software products.