← All observations

Field notes / 04 ·

The tests passed.
Which tests?

We had useful checks. We also had a command that ran only one part of them. The next requirement was to make the evidence as understandable as the code it describes.

Three stations on a miniature paper workbench inspect a component, an assembled website and its desktop and mobile views.
A component, an assembly, a visitor’s view. Each inspection answers a different question. A conceptual illustration, not a diagram of the test runner.

01 / The missing promise

A familiar command with a narrow meaning.

In our server field note, five HTTP checks helped expose a cache still running old instructions. Those tests remained useful. But npm test invoked that particular file, while SEO and monitoring checks lived behind separate commands.

A successful run therefore meant something narrower than the command suggested. It also depended on a running proxy and cache. Someone changing a small URL helper could not use the ordinary test command for quick, isolated feedback.

A test result needs a scope and an environment before it becomes evidence.

The immediate requirement was simple: discover all unit tests with one command. Meeting it properly also meant separating checks that need built files, local servers or a real browser. Renaming the scripts alone would have preserved the confusion.

02 / Boundaries before tools

What does this check need to run?

We took the command layout of a related project as a starting point, then kept only what this website needs. There is no test database to reset here. Copying its database setup command would have added an obligation without a capability we use.

Unit

A function and its inputs. No build, listening server or Docker stack.

Integration

Real pieces working together: generated HTML, configuration files or GraphQL and metrics listeners.

Browser

A built website opened and navigated through Playwright.

Vitest now discovers the unit and integration suites independently. Playwright owns browser scenarios. These are development dependencies, pinned in the package manifest and lockfile; they are not a new runtime layer for visitors.

03 / Fast feedback

The default command can stay small.

npm test
npm run test:watch
npm run test:coverage

Unit tests live in tests/unit/, with discovery also configured for colocated tests in application and server source. The environment is Node.js. Adding a matching test file does not require adding another package script.

The first cases cover canonical URL normalization, rejection of foreign origins and safe JSON-LD serialization. Query strings, fragments and trailing slashes should not create different canonical pages. Text containing a closing script tag must survive serialization without becoming executable markup.

Coverage is available in the terminal and as HTML and LCOV reports under coverage/. Its configured scope is SEO and server TypeScript. This is unit coverage; it does not combine the integration or browser runs, and a percentage does not tell us whether the right behavior was checked.

04 / Real combinations

Some tests need more than a function.

npm run test:integration
npm run test:integration:watch

The first command builds the website before running tests/integration/. SEO checks inspect actual generated HTML: titles, canonical URLs, descriptions, images, structured data, author references and the fixed revisions attached to articles. The unknown-route check exercises the built React Router handler and expects 404 and noindex.

The monitoring configuration test runs the generator in a temporary directory. It checks that two sites receive separate probes and labels, and that a router cannot belong to both. That is an integration with files and a process, even though it requires no deployed services.

The GraphQL test starts local API and metrics listeners on operating-system-assigned ports. A resolver deliberately fails. The HTTP response is still 200, so the test also checks that the failure appears in the server-error metric. Cleanup is registered as resources are acquired, and temporary environment values are restored.

05 / The visitor’s path

A browser can catch what HTML inspection cannot.

npx playwright install chromium webkit
npm run e2e
npm run e2e:webkit
npm run e2e:ui
npm run e2e:report

By default, Playwright builds the site and starts its production Node.js server on port 4317. It refuses to reuse an existing listener there. Linux hosts may need browser system dependencies as well as the downloaded binaries.

Three scenarios run in desktop Chromium, desktop WebKit and mobile Chromium. They cover navigation with browser history and metadata updates, opening and refreshing a blog deep link, and unknown pages and missing assets returning 404.

The navigation scenario also checks heading focus and watches for browser errors. A marker placed on the window must survive the transition and back/forward navigation: a full document reload would erase it. These are concrete observations about the tested path, not a general proof that every component hydrates correctly.

Failures retain traces and screenshots; an HTML report helps inspect the result. To use an existing environment, set PLAYWRIGHT_BASE_URL. In that mode Playwright neither builds nor starts the application, so preparing the correct revision becomes the caller’s responsibility.

06 / The infrastructure boundary

A local server cannot stand in for Varnish.

TEST_URL=http://haih.localhost \
MONITORING_TEST_URL=http://127.0.0.1:18080 \
GRAFANA_TEST_URL=http://127.0.0.1:13000 \
npm run test:integration:stack

The separate stack suite expects an already running production preview. The site URL must pass through Traefik and Varnish; the monitoring checks also need provisioned services and the local Grafana password file. The addresses above are an example, not automatic provisioning.

This suite preserves checks for cache hits and TTL limits, uncached failures, HTTP methods, API availability, probes, metrics, log filtering and the dashboard. Missing infrastructure is a failure, not a silently skipped success.

Failure injection and notification fixtures stay behind explicit commands. One pauses the preview application; the other uses local SMTP and Telegram API fixtures. Ordinary test discovery does not start those operational exercises.

07 / Evidence, with its limits

A smaller claim we can repeat.

  • 9 unit cases passed without a running application.
  • 10 local integration checks passed against the generated build and isolated fixtures.
  • 3 browser scenarios passed in each of 3 projects: 9 Playwright executions.
  • Type checking, lint and the unit coverage command completed successfully.

Those counts describe the testing checkpoint linked above, before this article was added. The Docker stack was stopped during that verification, so the migrated Varnish and Grafana checks were not rerun. The local browser run does not establish cache behavior, development HMR, complete accessibility or performance. Lazy-chunk failures and broader error-recovery paths still need focused coverage.

The versioned tests and commands and prerequisites make the result inspectable. They also expose the cost: a small unit suite is cheap to run, while browser binaries and a provisioned monitoring stack require additional resources.

For an AI-assisted workflow, that distinction matters. “Tests passed” is easy to report. Naming the behavior, revision and environment makes the statement useful. The next requirement can now bring its own check into a known place, with a command that actually finds it.

Written against 349ed44. The illustration is conceptual; technical claims refer to the linked project snapshot.

← Back to the field notes