Key data points

  • Across 5,733 individual test findings, comprehension friction, content that testers could technically access but couldn’t understand, was flagged more often than any other category, at 29.4%, ahead of “content not found” (21.6%) and accessibility-specific friction (19.4%).
  • The most common WCAG criteria underlying comprehension findings are 3.3.2 (Labels or Instructions) and 3.1.3 (Unusual Words), tied at 99 occurrences each, followed by 2.4.6 (Headings and Labels, 66) and 3.1.5 (Reading Level, 62).
  • Comprehension friction is not confined to any one disability group: it is the single most-cited friction category for testers with limited English proficiency (181 findings), and the second most-cited for blind testers (163) and low-vision testers (104), cohorts more commonly associated with access barriers than language barriers.
  • When findings are synthesised into higher-level patterns rather than counted individually, accessibility-specific issues account for a larger share (35.9%) than comprehension (29.1%), because accessibility complaints from the same tester tend to consolidate into a smaller number of large, high-severity patterns, while comprehension issues are more numerous but more scattered.
  • Accessibility-tagged patterns are rated high or severe severity 74.3% of the time, compared to 37.7% for comprehension-tagged patterns, meaning comprehension issues are reported more often, but accessibility issues are more likely to be serious when they occur.

Two different failure modes, both labelled “accessibility” by most teams

Most product teams treat “accessibility” as a single category: does the product work with assistive technology. This dataset (5,733 individual findings across 15 independently tested products) separates that into two measurably different failure modes. One is access: can the tester’s assistive technology or interaction method operate the interface at all. The other is comprehension: once the content is reached, does the tester actually understand what it says. These are tracked as separate friction types in this data, and the comprehension category is, by simple count, the largest single failure mode measured, ahead of pure access failures.

It’s worth being precise here about where comprehension sits in how a product is scored, because a category called “comprehension” could easily be mistaken for a usability concern rather than an accessibility one. On this platform it’s scored as the latter, with a deliberately narrow test behind it: not whether a tester read and understood every word on a page, but whether they understood the content that was critical to making an informed decision about the product or service under test. The clearest illustration in this dataset is insurance: the accessibility question isn’t whether a tester read every clause of a policy document, it’s whether they understood enough of what a policy did and didn’t cover to judge its value and whether it suited them, without that judgement being blocked by content they couldn’t access or interpret. Whether they also read the fine print further down the page is a separate, usability-level question, with a different bar and a different owner.

Comprehension is a plain-language problem more than a translation problem

Looking at which WCAG criteria are actually tagged against comprehension-labelled findings clarifies what’s failing:

WCAG criterionDescriptionOccurrences in comprehension findings
3.3.2Labels or Instructions99
3.1.3Unusual Words99
2.4.6Headings and Labels66
1.3.1Info and Relationships65
3.1.5Reading Level62
3.1.4Abbreviations34
3.3.1 / 3.3.3Error Identification / Suggestion18 each

Unusual Words (3.1.3) and Reading Level (3.1.5) together account for a substantial share of comprehension findings, and both concern plain-language writing: unexplained jargon, technical terminology, and sentence complexity above what the content actually requires. This is a content design and UX writing problem, not primarily a translation, captioning, or code-level accessibility problem, and it shows up across content that is already technically “accessible” by conformance standards.

Why 3.1.3 (Unusual Words) is easy to miss in a conformance report

One criterion sits underneath a large share of the comprehension findings above without getting talked about much on its own: WCAG Success Criterion 3.1.3, Unusual Words. It’s a Level AAA criterion, and what it actually asks for is narrower than the name suggests: a mechanism must exist for identifying the specific meaning of words or phrases used in an unusual or restricted way, including idioms and jargon, not that a page avoid unusual language altogether.

That narrowness matters, because it helps explain why 3.1.3 is rarely flagged in a typical WCAG conformance report, even on content that testers plainly struggle to understand. Deciding what counts as “unusual” is, in practice, a judgement call made by whoever is doing the assessment, and it’s a call that tends to run in one direction: a term that reads as ordinary trade language to a specialist auditor is often exactly the jargon that trips up a tester with limited English proficiency, or a screen reader user meeting it without the surrounding visual layout to lean on. Because the threshold isn’t fixed, 3.1.3 is one of the criteria most likely to be under-flagged relative to how often testers actually report being confused by wording, and its scarcity in conformance reports says more about how subjectively it gets applied than about how rare the underlying problem is.

There’s a second, more structural problem with 3.1.3 as it’s conventionally tested, and a pass/fail check for “does a mechanism exist” tends to hide it. The criterion is satisfied by the presence of a glossary, a tooltip, or an inline definition, checked as a yes-or-no. But the reason unexplained jargon is a comprehension barrier in the first place is the cognitive load of not knowing what a term means while trying to act on it, and sending a reader away from the sentence they’re on to a glossary entry and back reintroduces a version of that same load, just relocated rather than removed. A page can carry a technically compliant glossary and still fail the outcome 3.1.3 exists to protect, because whether the fix worked can’t be measured by whether a definition exists somewhere on the page; it depends on whether the tester actually got the meaning without losing their place in what they were doing.

Comprehension friction cuts across cohorts that don’t share a language or sensory barrier

The cohort breakdown for comprehension findings is the clearest evidence that this isn’t a niche language-access issue: limited-English-proficiency testers generate the most comprehension findings (181), unsurprising, but blind testers (163) and low-vision testers (104) are the second and third largest sources, followed closely by deaf testers (101), over-65 testers (90), and neurodivergent testers (86). These six cohorts share almost nothing in terms of assistive technology or sensory need, yet all six generate substantial comprehension friction. The common thread is not disability type; it’s that unclear labelling, undefined jargon, and dense sentence structure create friction for any tester, regardless of how they’re accessing the content.

Why the ranking flips at the pattern-synthesis level

A close look at this dataset’s two levels of analysis explains an apparent contradiction that’s worth being transparent about. At the level of individual, unsynthesised findings, comprehension is the largest category (29.4% vs. 19.4% for accessibility). But once findings are synthesised into consolidated patterns (grouping multiple individual observations from the same root cause into one pattern-level insight), accessibility overtakes comprehension (35.9% vs. 29.1%). The explanation is consistent with how accessibility failures tend to behave: a single structural access barrier (for example, a control with no accessible name) often generates many individual findings from the same tester across a session, which consolidate into one large, high-severity pattern. Comprehension issues are more evenly distributed (many separate wording, labelling, and structure problems scattered across different parts of a product), so they generate more individual findings but fewer giant consolidated patterns. Both views are accurate; they answer different questions. If the question is “what do testers run into most often,” comprehension leads. If the question is “what’s driving the largest, most severe consolidated problem areas,” accessibility leads.

What this means for content and product teams

The practical implication is that a plain-language pass is not a “nice to have” layered on top of accessibility work: in this dataset, it addresses the single most frequently occurring friction category, ahead of pure access failures, and it benefits cohorts far beyond the audience most content teams assume they’re writing for. Two specific, low-effort interventions map directly onto the largest comprehension sub-categories: defining or removing unusual/technical terms on first use (addressing 3.1.3, tied for the top comprehension-linked criterion) and auditing form labels and instructions for specificity (addressing 3.3.2, the other top criterion): together these two fixes touch roughly 43% of all comprehension findings with a WCAG tag in this dataset.

Frequently asked questions

What is the most common type of usability friction found in real testing, comprehension or accessibility?

It depends on the level of analysis. Counting individual findings directly, comprehension friction is the largest category at 29.4%, ahead of accessibility-specific friction at 19.4%. Counting consolidated, synthesised patterns, accessibility edges ahead at 35.9% versus 29.1% for comprehension, because accessibility issues tend to generate fewer but larger, more severe consolidated patterns.

Which WCAG criteria are most associated with comprehension failures rather than access failures?

Labels or Instructions (3.3.2) and Unusual Words (3.1.3) are tied as the most common, each appearing in 99 comprehension-tagged findings in this dataset, followed by Headings and Labels (2.4.6) and Reading Level (3.1.5).

Is content comprehension only a problem for non-native speakers?

No. While limited-English-proficiency testers generated the most comprehension findings in this dataset, blind testers were the second-largest source (163 findings) and low-vision testers were third (104), followed by deaf, over-65, and neurodivergent testers: comprehension friction affects cohorts with no shared language or sensory barrier.

About this analysis

Figures in this article are drawn from an anonymised aggregation of 15 independent usability testing projects conducted by See Me Please between late 2025 and mid-2026. Friction-type tagging, cohort assignment, and WCAG mapping were applied during test synthesis by trained reviewers based on task-based sessions. Individual-level figures reflect unsynthesised test findings; pattern-level figures reflect the same findings after consolidation into higher-level insights. All client and participant identities have been removed.

Why this matters commercially, not just ethically

What makes this dataset unusual is not that any single cohort struggled with content: it is that six cohorts with almost nothing in common, testers with limited English proficiency, blind testers, low-vision testers, deaf testers, testers over 65, and neurodivergent testers, independently converged on the same complaint. When that many structurally different groups flag the same friction, it stops being a niche access issue and becomes a textbook case of universal usability: a problem large enough to affect a substantial share of the general population, not an edge case that only shows up for a narrow assistive-technology audience. That distinction is exactly why mature digital organisations invest in user testing in the first place. If a customer can genuinely self-serve online, that is not just more convenient for them: it is significantly cheaper to deliver than routing the same transaction through an assisted channel, and the cost gap between digital and assisted channels is large and consistently documented across jurisdictions.

Deloitte’s Digital Government Transformation analysis estimated the average cost of an Australian government customer transaction as 94% cheaper than servicing customers over the phone, and 98% cheaper than in-person delivery. The Western Australian Parliament’s Public Accounts Committee reproduced these figures in its inquiry into public-sector ICT: $16.90 face-to-face, $6.60 telephone, $0.40 online. The WA Auditor General separately described this as face-to-face transactions costing 42 times more than online.

The pattern is not unique to Australia. The UK Government Digital Strategy cited a 2012 study of 120 local councils finding average contact costs of £8.62 face-to-face, £2.83 by phone, £0.15 via the web, concluding central-government digital transactions could be almost 20 times cheaper than telephone and up to 50 times cheaper than face-to-face. Separately, the UK Government’s more recent 2026 Digital and Data Benefits Framework takes a different approach: rather than asserting one universal ratio, it recommends agencies calculate channel-specific cost-to-serve, and its worked example uses £0.25 online versus £4 by phone (a 16 times differential, a 93.75% reduction in unit cost). GDS estimates greater digitisation across 7,000+ government services could deliver approximately £1.5 billion in savings.

In Ireland, the Comptroller and Auditor General’s examination of motor-tax administration found processing a payment in a motor tax office cost approximately €10, compared with €5 online, around 50% less. That differential is worth reading with a little caution when comparing percentages internationally: the €5 online figure appears to include more of the underlying processing cost than some UK/Australian definitions of a marginal digital interaction, which helps explain why this differential lands at 2 times rather than the 15 to 50 times seen elsewhere.

The United States shows the same pattern at a more granular level. The U.S. Government Accountability Office (GAO), using IRS data, reported the following costs for taxpayer authentication in FY2017:

ChannelCost per authentication
OnlineUS$0.20
Automated telephoneUS$0.60
Telephone with a customer service representativeUS$54

An online transaction was approximately 99.6% cheaper than a representative-assisted telephone interaction in this case: the human-assisted telephone interaction cost roughly 270 times as much as online.

Most mature organisations already understand this commercial case for digital self-service in the abstract. But mainstream user testing programs still tend to focus on general sentiment and ease of navigation for a broad, assumed-typical user, while testing with older users, testers who speak English as a second language, and people living with disability gets bucketed separately as “accessibility” testing or treated as an edge case to be addressed later, if budget allows. The data in this article suggests that framing is backwards. If comprehension friction, the single most frequently flagged usability failure in this dataset, shows up hardest in exactly the cohorts that mainstream testing programs sideline, shouldn’t those cohorts actually be the priority, not the afterthought? Testing with them isn’t a narrower, more specialised exercise than testing with a general audience: on this evidence, it is a more powerful and more indicative way to find out what creates a seamless experience for everyone.