Skip to main content

Accuracy benchmark

How accurate are Lumtera's checks?

On the 21 W3C ACT rules its checks cover, Lumtera reported 130 of 132 known failures (98%) at any severity: 95 as errors and 35 for review, with 4 false alarms in 222 clean examples. Here is the method, the full results next to axe-core's, and what the numbers do not show.

Run on September 27, 2026 with the fixes released in Lumtera 1.0.0.2. Our own test, not an independent audit.

A second test, on October 2, 2026, measured false alarms on 529 accessible and real-world pages. See the false-alarm results

The numbers

  • 98% of failing examples found at any severity, on the 21 W3C ACT rules Lumtera's checks cover.
  • 72% of those failing examples reported as an error. The rest were marked Needs review, for a person to decide.
  • 1.8% of passing or inapplicable examples reported as an error (4 of 222).

Results, next to axe-core

We ran axe-core, the open-source engine behind many accessibility tools, on the same pages as a second opinion. It is a strong, well-established checker, and on several measures it does better than Lumtera:

Lumtera and axe-core on the W3C ACT test cases
Measure Lumtera axe-core 4.x
Failing examples found at any severity (21 covered rules, 132 examples) 130 (98%) 105 (80%)violations only; "incomplete" not counted
Failing examples found as an error (like for like) 95 (72%) 105 (80%)axe violations are its errors
False alarms reported as errors (of 222 clean examples) 4 2
All 34 rules benchmarked (183 failing examples), at any severity 133 (73%) 139 (76%)

Read the first row with care. The 21 rules were chosen because Lumtera covers them, so a high score there says Lumtera does what it claims on those rules — not that it finds more than axe. Counted as errors, axe found more; across every rule we benchmarked, axe found more; and axe raised half as many false alarms. Lumtera marks anything it cannot decide as Needs review rather than Error, which is why its "any severity" figure is higher than its "error" figure.

False alarms on accessible pages

On October 2, 2026 we ran every Lumtera engine on 529 pages that meet WCAG 2.2 AA or nearly so, and on all the W3C ACT test cases, then checked every error by hand. The fixes released in Lumtera 1.2.2 took the errors we judged to be false alarms from 194 to 0 in review mode, without losing any known failure in the ACT test cases. An error on a page that meets WCAG counts as a false alarm; Needs review does not, because it asks a person to decide.

Errors and review items before and after the October 2026 fixes
Measure Before (Lumtera 1.2.1) After (Lumtera 1.2.2)
Errors on the 529 accessible and real-world pages, every engine 384 190
… of which false alarms, after checking each by hand 194 0 in review mode; 8 in the editor's check of saved content
Needs review, whole-page checks 1,275 592
ACT passing or inapplicable examples reported as an error (of 496) 7 1
ACT known failures found at any severity (of 259) 177 177

Each of the errors left is a real problem in the page as published: broken in-page links in the GOV.UK component examples, untitled frames in the W3C pattern examples, and two core output problems in Twenty Twenty-Four. The eight remaining false alarms are in two W3C pattern examples whose script changes the page after it loads; the editor's check reads the saved markup and cannot see that, while review mode checks the page as the browser shows it and reports nothing there.

The pages: the 1,209 ACT test cases, 76 WAI-ARIA Authoring Practices examples, 248 GOV.UK Design System component examples, WordPress 7.1 with Twenty Twenty-Five and Twenty Twenty-Four (every pattern included), WooCommerce pages and nine hand-written accessible pages. The fixes lower a check's certainty where the page cannot decide: content in noscript and template tags is skipped, focus guards, symbol icons, unmuted autoplaying video, empty alt text on Image blocks and unnamed dialogs became Needs review, more than one H1 became a tip, and the whole-page checks now read modern CSS colors, overlays and custom focus styles. These figures use a different corpus and scoring from the September test, so they are not added to it.

How we tested the ACT results

  • The corpus: 572 test cases for 34 ACT rules from the W3C ACT Rules repository — small example pages, each marked as passing or failing a rule — plus the W3C "Before and After" demo pages.
  • Every case was checked twice: by Lumtera's content checks (the code behind wp lumtera check), and by review mode's whole-page checks in headless Chrome, with default settings. On-demand tests such as the keyboard walk and text spacing were not run.
  • Each ACT rule was mapped to the Lumtera checks that cover it. A failing example counts as found when one of those checks reports it; a passing or inapplicable example counts as a false alarm when one reports it as an error. Needs review and tips are counted separately.
  • axe-core ran in the same browser page. Its mapped rule's violations count as found; its "incomplete" results do not.
  • On the W3C "Before and After" demo, each inaccessible page had between 33 and 45 Lumtera errors, all real barriers, and the four accessible pages had none.

What these numbers do not show

  • Test cases are small, focused examples, not real sites.
  • Some of the ACT rules we ran are outside Lumtera's checks, among them ARIA attribute values, decorative elements and some table header rules. Lumtera found 3 of their 51 failing examples; axe found 34. Those gaps are listed in full in the benchmark.
  • Many WCAG criteria need a person: whether alt text is right, whether captions match the audio, whether reading and focus order make sense, whether error messages help. No automated checker can decide these.
  • Known limits remain: contrast inside shadow DOM is not measured yet, and the editor's check of saved content cannot see changes a script makes after the page loads. Focus guards that a script moves on and silent autoplaying videos are now Needs review rather than errors, because only running the page, or hearing it, can settle them.

Reproduce it

  1. Fetch testcases.json and the test case files from the w3c/wcag-act-rules repository, and the W3C "Before and After" demo pages.
  2. Serve the files locally.
  3. Run wp lumtera check on each file, then load each page in a browser with review mode's page checks, and run axe-core on the same page.
  4. Score each ACT rule with the rule-to-check mapping in the check library.

The regression cases the benchmark produced ship with the plugin's own tests, so the fixes it led to stay fixed. The October false-alarm test can be repeated with the scripts in the plugin's tests/corpus folder, which fetch the third-party pages rather than shipping them.

Sources

See what it finds on your site

Every check is in the free plugin: 69 content checks and 33 whole-page checks, with anything uncertain marked Needs review.

Start with the free plugin today

All 69 content checks and 33 whole-page checks, fixes you approve with undo, and the site report — free, with no account and no page limits.