AI, Accessibility, and Learning
Large language models sit in an unusually double-edged position for disabled learners. The same technology that can turn a dense PDF into plain language on request, or generate live text suggestions for someone who communicates through an augmentative and alternative communication (AAC) device, is trained on data that measurably encodes bias against disability, and answers confidently whether or not it is right. Evaluating what these tools are actually good for — as opposed to what the marketing around them claims — means holding both of those facts at once, rather than picking the optimistic story or the cautionary one and stopping there.
Genuine Assistive-Technology Uses
The clearest, most directly studied case is AAC. A 2023 study tested live large-language-model text suggestions with twelve AAC users across three realistic communication scenarios — extending short replies, answering biographical questions, and making requests — and found the models could meaningfully improve the speed and variety of what users were able to say, while also identifying real failure modes: suggestions that didn't match a user's authentic voice, or that constrained rather than expanded what they could communicate [1]. That combination — genuine benefit alongside specific, documented ways it can go wrong — is a more honest picture than either "AI transforms AAC" or "AI can't be trusted here" on its own.
A larger-scale study of visually impaired university students (n=385) found that trust in LLM tools was a significant predictor of whether students used them, and that usage was associated with an improvement in self-reported quality of life; usage did not translate directly into better academic outcomes on its own, but its effect on academic success operated indirectly, through that improvement in quality of life [2]. That's a more modest and more specific claim than "AI improves outcomes for disabled students" — it says something about the mechanism, and about where the benefit does and doesn't show up.
Beyond these two directly-studied cases, the plausible everyday uses are the ones anyone with a screen reader, a reading difficulty, or a processing difference already recognises: converting dense text to a simpler register on request, describing an image in more or less detail depending on what's asked for, and drafting a first pass at something a person then edits rather than accepts outright. The evidence base for these everyday uses specifically, as opposed to the two studied cases above, is thinner than the volume of enthusiastic commentary about them would suggest — which is itself worth knowing before treating any of it as settled.
The Bias Problem Is Measured, Not Hypothetical
It would be easy to treat "AI models might be biased against disability" as a plausible-sounding worry rather than a demonstrated finding. It is demonstrated. A 2020 study from Google evaluated toxicity-prediction and sentiment-analysis models and found they systematically rated sentences that merely mentioned disability — "I am a person with mental illness," with no other content — as more negative or more toxic than equivalent sentences without a disability mention, and traced part of the cause to the training data itself: text about mental illness, for instance, disproportionately co-occurs with topics like gun violence and homelessness, and the model absorbs that association as if it were a property of the topic rather than an artefact of what gets written about it [3]. A downstream content-moderation system built on a model like this doesn't just produce an occasional bad output — it can systematically suppress disabled people's own descriptions of their own experience, on the grounds that the topic itself reads as "toxic." That is a concrete, testable, already-tested failure mode, not a hypothetical one, and it means bias-testing an LLM-based tool against disability-related language specifically — not just testing it in aggregate — is a real requirement, not a nicety.
Fallibility Is the Same Tool, Not a Different One
The tendency of LLMs to produce fluent, confident, and sometimes entirely fabricated output — usually called hallucination — is well documented and remains an open research problem rather than a solved one; a recent comprehensive survey catalogues the causes across the whole model lifecycle, from stale or incorrect training data through to the way models are aligned and fine-tuned [4]. This matters more, not less, in an accessibility context: a student using an LLM to simplify a dense passage because reading the original is itself the barrier is often not well placed to independently notice when the simplification has silently dropped or distorted a fact, precisely because reducing that verification burden was the point of using the tool. The tools that reduce a barrier and the tools that introduce a new, harder-to-detect one are, here, literally the same tool — which argues for treating LLM output in these contexts as a draft that still needs an independent check, not for avoiding the tool altogether.
What Fallibility Does to Critical Thinking
A 2025 survey of 319 knowledge workers by researchers at Carnegie Mellon University and Microsoft Research, examining 936 real examples of how people actually used generative AI at work, found that higher confidence in the AI's output was associated with less critical engagement with it, while higher confidence in one's own judgement was associated with more — and that heavier AI use was associated with reduced critical thinking, an effect the study attributes to cognitive offloading: handing the AI the evaluative work as well as the generative work, not just the first draft [5]. A separate analysis in AI & Society makes a related, more structural argument: unrestricted access to an LLM is in tension with learning specifically because the model does too much of the work for the learner by default, and getting genuine learning value out of it requires the learner to already have enough domain knowledge to evaluate what comes back — the same critical capacity the tool, used carelessly, erodes [6].
That combination is a genuine bind for accessibility use in particular. The whole appeal of an LLM-based tool for many disabled learners is that it lowers the effort needed to access material in the first place; but the research above says the same lowered effort is exactly the mechanism through which critical evaluation of the tool's own output gets skipped. This is the general mechanism described in Learning as a Feedback Loop — a loop that "spins inside the model, not inside the learner" when the tool supplies the output, the explanation, and the assessment of correctness all at once — playing out in a context where the barrier the tool is removing is real, which makes the temptation to skip the evaluation step correspondingly stronger, not weaker. The distinction drawn in Scaffolding, the Zone of Proximal Development, and Using GenAI Well — support that is contingent, temporary, and transfers capability, versus a tool that just produces the artefact — applies directly: an LLM that simplifies a passage and is then checked against the original is scaffolding; one that is trusted outright because checking it defeats the purpose of using it is not.
Practical Implications
- Test accessibility tools against disability-related language specifically. Aggregate accuracy or aggregate toxicity scores can look fine while the tool is measurably worse for the population it's meant to serve — the Hutchinson et al. finding is exactly this pattern.
- Treat simplified or generated output as a draft, not a fact. Where the LLM is removing a reading or processing barrier, build in an independent check — a second source, a human reviewer, a targeted fact-check of the specific claim that matters — rather than trusting fluency as a proxy for accuracy.
- Notice when convenience becomes an excuse to skip evaluation. The mechanism identified in the critical-thinking research isn't "AI makes people worse at thinking" in the abstract — it's confidence in the tool substituting for one's own judgement, which is a choice made per use, not an inevitable side effect.
- Involve disabled users in evaluating accessibility tools before deployment, not after a complaint — the AAC study's value came specifically from testing with real AAC users across realistic scenarios, not from evaluating the model in isolation.
Related Topics
- Learning as a Feedback Loop (and the AI Partner) — the cybernetic model of why handing evaluation, not just generation, to an AI tool is what breaks learning.
- Scaffolding, the Zone of Proximal Development, and Using GenAI Well — the distinction between AI as scaffold and AI as answer machine, applied to reading, writing and design.
References
- Valencia, S., Cave, R., Kallarackal, K., Seaver, K., Terry, M., & Kane, S. K. (2023). "The less I type, the better": How AI Language Models can Enhance or Impede Communication for AAC Users. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. https://dl.acm.org/doi/10.1145/3544548.3581560
- Elshaer, I. A., AlNajdi, S. M., & Salem, M. A. (2025). Measuring the Impact of Large Language Models on Academic Success and Quality of Life Among Students with Visual Disability: An Assistive Technology Perspective. Bioengineering, 12(10), 1056. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12562211/
- Hutchinson, B., Prabhakaran, V., Denton, E., Webster, K., Zhong, Y., & Denuyl, S. (2020). Social Biases in NLP Models as Barriers for Persons with Disabilities. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5491–5501. https://aclanthology.org/2020.acl-main.487/
- Huang, L. et al. (2025). A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems. https://dl.acm.org/doi/10.1145/3703155
- Lee, H.-P. (H.), Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI '25). https://www.microsoft.com/en-us/research/publication/the-impact-of-generative-ai-on-critical-thinking-self-reported-reductions-in-cognitive-effort-and-confidence-effects-from-a-survey-of-knowledge-workers/
- Rus, V., & Kendeou, P. (2025). Are LLMs actually good for learning? AI & SOCIETY. https://link.springer.com/article/10.1007/s00146-025-02323-9