This piece proposes a hypothesis about who self-report bias affects most in accessibility research, tests it against cohort-level data, and considers explanations drawn from existing disability research rather than inventing a new one from scratch.

The hypothesis

A separate analysis of this same testing program found that self-reported satisfaction scores run substantially higher, on average, than scores derived from independent behavioural review of the same sessions: people say it went better than a trained reviewer, watching the same footage, would rate it. That raises an obvious next question: is that gap the same size for everyone, or does it vary systematically by who’s doing the rating? If it varies, the follow-up question is more uncomfortable: does the gap happen to be largest for the group with the worst underlying experience, meaning the people most affected by a product’s failures are also the least likely to tell you, in their own words, how bad it actually was?

What we measured

Across 1,500 tester-task sessions with both a raw self-reported rating and a calculated accessibility score, we normalized the self-report to the same 0-100 scale and computed the gap: self-reported score minus calculated score. A positive gap means the tester rated their experience more favourably than the calculated score would suggest; a negative gap means the reverse.

Overall, the average gap is +16.8 points, with a median of +12.5, confirming the general inflation pattern. The more revealing result is what happens when this is broken down by cohort:

Cohort n Mean gap (self-report − calculated)
Indigenous 56 +0.6
Deaf 190 +10.1
Physical disability 38 +11.3
Low vision 271 +13.7
Over-65 123 +14.2
Limited English proficiency 242 +15.7
Neurodivergent 207 +19.1
Blind 373 +27.0

Blind testers, the cohort with, by a wide margin, the lowest average calculated accessibility score in this entire program (29.0 out of 100, reported separately), show the largest self-report inflation gap of any cohort measured, at +27.0 points, on the largest sample size in this table (n=373). Indigenous testers, by contrast, show almost no gap at all (+0.6): their self-reported ratings track the calculated score about as closely as any cohort in this dataset.

Put plainly: the group experiencing the worst outcomes, by the most objective measure available in this data, is also the group whose own account of that experience is the least reliable indicator of how bad it actually was, and it isn’t close. The gap for blind testers is roughly 1.6 times the overall average and roughly 1.4 times the gap size of the next-largest well-sampled cohort (neurodivergent, at +19.1).

Why we don’t think this is a data artifact

Before treating this as a real finding, it’s worth ruling out the boring explanations. The sample size for blind testers (373) is the largest of any cohort in this table, so it isn’t a small-n fluke. The overall correlation between self-report and calculated score across the full dataset is a moderate r = 0.568, meaning self-report is still tracking real variation in experience, just from a systematically shifted baseline for this cohort, rather than being random noise. And the pattern holds in the direction we’d actually worry about: it isn’t that blind testers rate everything a flat 100 regardless of experience (which would produce no correlation at all), it’s that their ratings move in the right direction but land consistently higher than the independently observed reality would justify.

Competing explanations

We think three explanations are worth taking seriously, and they aren’t mutually exclusive.

Adaptive expectations. Disability researchers have long documented what’s sometimes called the disability paradox or adaptive preference, the pattern where people who navigate chronic barriers calibrate their expectations to the reality they actually encounter, rather than to some external ideal. A blind participant who has spent years working around inaccessible websites may reasonably rate a session a “4 out of 5” not because the session was objectively good, but because it was better than their typical experience, or because they did eventually succeed despite the barriers, a comparison point a sighted, first-time reviewer applying a consistent behavioural rubric would not use. If this is the mechanism, the gap isn’t really “bias” in the sense of an error, it’s a genuine difference in reference frame that self-report surveys, by design, can’t distinguish from a report about the product itself.

Relief and gratitude effects. A related but distinct possibility is that when a screen reader user succeeds at a task that could easily have failed, the emotional relief of success gets folded into the satisfaction rating, independent of how many workarounds or how much extra effort the success required. This would predict exactly what we see: high self-report specifically on sessions with meaningful friction, since the friction is precisely what makes eventual success feel notable.

Instrument mismatch. It’s also possible the self-report question itself (“how was that?”) and the calculated accessibility score are measuring different things for this cohort specifically. If the calculated score weights strict conformance criteria (labelling, structure, keyboard operability) more heavily for blind testers than the self-report question implicitly does, the two numbers could diverge for measurement reasons rather than psychological ones. We can’t fully separate this from the two explanations above with the data available here.

Why the Indigenous cohort is the most interesting control group in this table

If self-report were simply unreliable across the board, we’d expect every cohort’s gap to cluster somewhere well above zero, with only the size varying. Indigenous testers break that pattern entirely, sitting at +0.6, essentially perfect agreement between what they said and what was calculated. That’s useful because it demonstrates the gap isn’t a universal property of self-report surveys or of testing methodology; it’s specific to certain groups’ relationship with certain products or barriers. Combined with a separate finding from this program (that Indigenous testers’ dominant friction pattern is about content findability and trust rather than assistive-technology operability), one plausible reading is that self-report tracks reality closely when the barrier is “I couldn’t find what I needed” (a concrete, easily narrated failure) and diverges most when the barrier is a diffuse pattern of small operability frictions accumulated over an entire session (harder to narrate, easier to round up to “fine, I got there”).

Why this matters beyond research methodology

If organisations rely on customer satisfaction surveys, app store reviews, or “was this helpful” widgets as a proxy for accessibility quality (which most do, because it’s cheap and scalable), this finding suggests those mechanisms will most severely understate the real experience of exactly the population most affected by accessibility failures. A satisfaction dashboard built entirely on self-report would show blind users as only modestly less happy than everyone else, while the underlying, independently observed reality in this dataset shows them as by far the most affected group. That’s not a small measurement error; it’s a systematic one that runs in the direction of hiding the problem it would be most important to surface.

Limitations

This is a single testing program, and we have not been able to independently verify which of the three explanations above (adaptive expectations, relief effects, or instrument mismatch) is doing the most work, or whether it’s some combination that varies session to session. The physical disability (n=38) and Indigenous (n=56) cohorts, while large enough to report, are meaningfully smaller than the blind (n=373) or neurodivergent (n=207) samples and should be weighted accordingly. We also have not been able to test whether the gap is stable over repeated sessions with the same individual, which would help distinguish a stable trait (adaptive expectations) from a session-specific state (relief at a particular success).

Open question

If adaptive expectations are part of what’s happening, the practical implication is uncomfortable: asking users to self-report on accessibility may be least trustworthy precisely for the population a team most needs accurate signal from. That suggests behavioural observation shouldn’t be treated as a nice-to-have supplement to self-report for accessibility testing: for at least this cohort, it may be closer to a necessity. Whether that holds outside this specific dataset, and whether it’s possible to design a self-report instrument that corrects for it (for example, by asking about specific task steps rather than overall satisfaction), is the natural next study.

About this analysis

Figures in this piece are drawn from an anonymised aggregation of scoring data across 15 independent usability testing projects conducted by See Me Please between late 2025 and mid-2026. Self-report gap is calculated as a tester’s own post-task rating (normalized to a 0-100 scale) minus the calculated accessibility score for the same session. All client and participant identities have been removed.