Why We Built Our Own Testing Platform
We kept losing a quarter to procurement.
The engineering case for visual regression and Lighthouse gating on a storefront takes an afternoon to make. The off-the-shelf tools are good and we would recommend several of them. Then the request goes into a vendor queue at a large client and sits there, because adding a SaaS product means a security questionnaire, a legal review, a question about where screenshots of unreleased pages are stored and who can see them, a new contract, a new billing relationship, and a new set of accounts to offboard when people leave. None of that is unreasonable. It is just slow.
Meanwhile the client already has an AWS account, an approved AWS relationship, and a team who operates it every day. Screenshots of a pre-release page are much easier to keep inside that boundary than to argue about moving out of it.
So we built the tests we wanted to run as a service that deploys into the client's own account.
What it is
A small serverless platform with two planes and a dashboard.
The visual plane takes screenshots that CI captured on the deploy preview and pixel-diffs each one against the approved baseline for that branch. The performance plane accepts uploads from the stock Lighthouse CI client, evaluates them against server-side budgets, and records per-page trends. The dashboard is where a reviewer looks at diffs, approves or rejects them, manages budgets, chooses which pages CI should check, and sees who did what.
Both planes post a commit status on the pull request, each linking to the build in the dashboard. A team can make either one a required check, or leave it advisory while they tune it.
How it finds issues
Visually, each uploaded screenshot is compared to the baseline for the same page. The diff engine does very little on purpose. It compares pixels, ignores anti-aliasing so sub-pixel font rendering does not register as a change, and reports the changed bands of the page so a reviewer can jump to the affected region. A page whose dimensions changed (the classic "a section disappeared" symptom) is treated as changed outright, not as an error.
Every snapshot ends up as one of unchanged, changed, new, or removed. The build passes only if every snapshot is unchanged. Anything else puts it into changes_pending and the PR check goes red until a human signs off. Approval promotes the new screenshots to baseline. Rejection keeps the old one and leaves the check red.
Two rules keep that useful rather than noisy. Decisions are made on integer pixel counts, never on floating-point ratios, so a verdict cannot flip on rounding. And removed is a real result: if a page quietly drops out of the capture set, the platform says so and asks someone to confirm, instead of silently shrinking the baseline.
A config change on a different theme moved this storefront's header navigation. Nobody had touched this theme, so nobody would have thought to check the header. Onion-skinning the baseline over the current capture shows the shift in a second, and side by side confirms nothing else moved.
Performance wise, the perf plane accepts uploads from the standard Lighthouse CI client, so there is nothing bespoke to install in a client's pipeline. Budgets live on the server: minimum scores per category and ceilings on the handful of metrics that matter. A team can wire up CI first and switch gating on once they have real numbers to set them from.
Those budgets need headroom. Lighthouse in CI is a lab number from one machine, as the docs say, and a budget pinned to the typical result will flake.
Trends are tracked per page and per device, not per deploy preview. Previews get a new hostname every commit, and if that hostname were the key every build would start a fresh line. Tracking /collections/shop-all as one line across builds means a slow regression shows up as drift over weeks instead of hiding in the noise.
Never leave a PR stuck
A gating tool that occasionally leaves a PR pending forever is worse than no tool, so every build resolves to a status, including the ones whose worker died halfway through.
If the platform itself is down or slow, CI notices, falls back to how it worked before, and carries on. The platform can fail. The client's pipeline should not.
It runs in the client's AWS account
This is the part that solved the original problem. The whole thing is an SST app and deploys with one command into whichever client account the operator is signed into.
Everything it creates is a managed AWS service the client's team already runs. Lambda for the API and the workers, S3 for screenshots and reports, DynamoDB for the records, CloudFront in front of the dashboard, Cognito for sign-in. Users are admin-created only. There is no self-signup.
No servers to patch. No always-on cost. No account IDs anywhere in the repo, so the same code deploys into any client's account under their own credentials. Screenshots of unreleased work never leave that account, access follows the client's existing controls, and every reviewer action lands in an audit trail.
From a procurement point of view, nothing was bought. It is code running on infrastructure the client already pays for and has already approved.
What it cost us
We own the maintenance now, and the runbooks, and keeping up with the Lighthouse CI client. We pinned the scope hard to keep that bearable: two planes, one dashboard, and an infrastructure diff reviewed before every deploy. It is not a feature-for-feature replacement for a mature SaaS product.
For the clients it was built for, though, the realistic alternative was no tool for a quarter. Being able to say "it runs in your AWS account, here is the infrastructure diff, and nothing leaves" moved the conversation from procurement to engineering. That is the conversation we are good at.