An accessibility analytics vendor can look inexpensive until the product team has to explain what its numbers miss and maintain the data feeding them. I would buy only after a short, instrumented pilot against real user journeys. A dashboard demonstration is a poor basis for a delivery estimate because it reveals little about authenticated screens, manual testing, or the work required to keep results trustworthy.
The buying decision should start with journeys, not dashboards
Define the product surfaces before inviting proposals. Pick 8 journeys as a pilot scope to tune: include account creation, a failed form submission, a keyboard-only task, and the highest-value authenticated workflow. Record the route, user state, device width, and expected outcome for each one. That gives vendors the same test boundary; without it, one supplier may scan public pages while another prices the application your users actually operate.
Require each candidate to show how it identifies a result under WCAG 2.2 AA, preserves the affected URL or application state, and points a developer to reproducible evidence. Ask it to distinguish a rule failure from a suspected issue requiring human review. Automated checks cannot establish full WCAG conformance because tasks such as judging whether alternative text conveys an image’s purpose depend on context. A vendor that presents a single “accessibility score” without showing its denominator is therefore difficult to use for release decisions.
Cheap accessibility analytics will mislead your tech lead decisions is right to question dashboard confidence, but I would make reporting speed a secondary criterion because a fast chart cannot compensate for journeys the scanner never reaches. Ask each vendor to demonstrate the same login, validation error, and modal interaction live. Then have a tester repeat those tasks with a keyboard and a screen reader, noting which findings the product captured, missed, or described without enough context to fix.
Make ownership visible in the estimate. A product manager should budget time for preparing test accounts, stable fixtures, consent-safe analytics events, issue triage, and retesting fixes. Assign one person to decide whether two reports describe the same defect; otherwise, ticket counts can rise after a tool change even when the user experience has not worsened. The useful output of the pilot is a list of reproducible issues and uncovered paths, not a prettier trend line.
Open-source checks and a managed platform win at different stages
Compare axe-core with Playwright against the Level Access Platform rather than treating them as interchangeable purchases. axe-core with Playwright wins when the team owns its test pipeline and needs repeatable checks at specific interaction states, because developers can run the same assertions during pull-request testing. Its cost is engineering time: someone must write journeys, maintain selectors and fixtures, review results, and provide reporting outside the test runner.
The Level Access Platform is the stronger candidate when several teams need a shared review and remediation process, because a managed system can give product and accessibility specialists a place to coordinate findings beyond a single repository. Its cost includes subscription terms, onboarding, data-access review, and the work of mapping its findings back to the team’s release process. Ask for a quote against the same 8 journeys and named user seats; a platform price without that scope does not support a meaningful comparison.
Run a small baseline with @axe-core/playwright before evaluating claims about incremental coverage. After installing playwright and @axe-core/playwright with npm, run npx playwright install chromium, save this as audit.cjs, and execute node audit.cjs. Replace the example URL with a representative route once authentication and test data are ready.
const { chromium } = require('playwright');
const AxeBuilder = require('@axe-core/playwright').default;
(async () => {
const browser = await chromium.launch();
const page = await browser.newPage();
await page.goto(process.argv[2] || 'https://example.com/');
const result = await new AxeBuilder({ page }).analyze();
console.log(JSON.stringify(result.violations.map(v => ({
rule: v.id, impact: v.impact, instances: v.nodes.length
})), null, 2));
await browser.close();
})().catch(error => { console.error(error); process.exitCode = 1; });
This script produces rule-level evidence, not an accessibility verdict, because it audits the state reached after page load rather than every interaction or interpretation. Extend it only after deciding which states matter. Pa11y and Lighthouse CI can provide additional automated checks, but I would not add both to the pilot by default: overlapping reports create triage work unless the team can name a decision each extra result will change.
Integration cost appears before a team reaches scale
Cut-cost back ends cost more when your team hits 20 puts team size at the center of the cost discussion, but I would price accessibility-tool integration before headcount because authentication, test data, and finding ownership can consume a release budget on day one. A vendor trial against public pages conceals that work. Ask the supplier to demonstrate one authenticated journey using the same identity flow and deployment environment the product will support.
Scope the connection points explicitly. If the product uses OpenID Connect, determine whether automated tests can use dedicated accounts without weakening production sign-in rules. If a vendor collects interaction data, identify the fields sent, where they are stored, who can see them, and how deletion works. A report containing a URL with a customer identifier can expose information even when the vendor says it does not collect form contents. Have the security reviewer check the contract and the actual payload rather than relying on a feature sheet.
Require a portable path from finding to fix. A result should retain its WCAG criterion, route or state, reproduction steps, evidence, and owner when exported to the team’s issue tracker. If a supplier offers an API, request its OpenAPI 3.1 description or equivalent documentation and test pagination, authentication, and deletion in the pilot. If it offers SARIF for GitHub code scanning, verify that repeated runs update the expected finding rather than creating a new ticket each time. These checks expose maintenance effort that a dashboard cannot show.
Set a 30-day retention assumption to verify for pilot evidence and ask whether the vendor can enforce it; the right production period depends on contractual and debugging needs. Ask for the vendor’s availability commitment in writing, but treat a quoted 99.9% uptime figure as vendor-published until the agreement defines exclusions and remedies. Neither number proves accessibility coverage. They belong in the scope because the team will otherwise have to resolve unclear data and service expectations after signing.
A fixed pilot produces a defensible delivery estimate
Give each option the same 10-working-day planning allowance and require a written accounting of setup, test authoring, manual review, triage, and reporting. That is a budget to adjust, not a claim that every product can finish in 10 days. Reserve part of it for a keyboard and screen-reader review of the chosen journeys, because automated rule detection and task completion answer different questions. Have the reviewer record the browser, assistive technology, and result so the team can reproduce the observation.
Define acceptance before the trial starts. For each journey, the candidate should reach the intended state, produce a finding that another person can reproduce, or state clearly that the state was not tested. During the pilot, measure the minutes from a new finding to an actionable ticket; that observed interval is more useful for staffing than a vendor’s count of detected issues. Also record the proportion of findings closed as duplicates or non-issues, with the reviewer’s reason, so apparent coverage gains do not hide triage cost.
Test one change deliberately: introduce a known accessibility defect in a staging branch, then fix it and rerun the check. Confirm that the tool detects the defect at the relevant state, preserves enough evidence to locate it, and clears or updates the finding after the fix. Use WAI-ARIA 1.2 only where native HTML cannot express the control correctly; adding ARIA to make a test pass can create a misleading result if the control still fails during keyboard use.
I would not commit to a multiyear platform contract on the basis of automated issue totals, because those totals depend on which states were scanned and how repeated instances were counted. The decision record should instead show uncovered journeys, confirmed defects, manual-review effort, integration tasks, and the remaining uncertainty. If the managed option reduces coordination work enough to justify its subscription and onboarding cost, buy it. If the team can maintain the journeys and needs only reproducible checks, keep the Playwright baseline and fund the human review it cannot replace.
The first purchase task is to prepare a comparable test
Tomorrow, choose one authenticated journey and write down its start state, successful outcome, and most likely failure state. Ask both candidates to test those same states, then have a reviewer reproduce their findings. That small exercise will reveal missing access, weak evidence, and manual work early enough to change the scope—before a contract turns an optimistic demonstration into your delivery plan.



