Accessibility and inclusion in AI products
Be able to design AI features that work for people with disabilities and different language backgrounds.
Prerequisites
- EUX for AI featuresrequired
Intuition
WCAG 2.2 rests on four principles — content should be perceivable, operable, understandable and robust enough to work with assistive technology.
For the public sector in Sweden, level AA is a legal requirement (the Act on the Accessibility of Digital Public Services), and the Accessibility Directive extends the requirements to private actors from 2025.
The basics apply to AI interfaces too, of course:
| Requirement | In practice |
|---|---|
| A contrast of ≥ 4.5:1 | body text against the background |
| Keyboard navigation | everything has to be reachable with Tab |
| A visible focus indicator | you have to be able to see where you are |
| Alt text on images | or alt="" if decorative |
| Labels on form fields | connected with for/id |
| Click targets ≥ 24×24 px | especially on mobile |
But AI interfaces have problems of their own that WCAG was not written for.
Formal
Five AI-specific accessibility problems:
| Problem | Why it arises | The solution |
|---|---|---|
| Streaming answers | a screen reader re-reads the text when it changes | aria-live="polite" on a stable container, update in blocks |
| Generated images with no alt text | there is no author who wrote one | generate a description with a VLM, let the user correct it |
| No time indication while waiting | it is unclear whether the system has hung | a status message in a live region, not just a spinner |
| Long answers | heavy to navigate without structure | headings, lists, skip links |
| A voice-only interface | excludes deafness and speech difficulties | always a text alternative |
Streaming answers are the hardest and the most common failure. A naive implementation updates the same element token by token; a screen reader interprets that as the content changing hundreds of times and re-reads from the beginning. The solution is to buffer to sentence level and use aria-live="polite" so that the reading does not interrupt itself.
Linguistic inclusion is the other half:
| Group | The need |
|---|---|
| Swedish as a second language | simpler language on request, not just one complicated answer |
| Dialect and spoken forms | the model has to understand input that is not written standard |
| Sign language | a text alternative, and video where possible |
| Reading difficulties and dyslexia | read-aloud, shorter paragraphs, a clear structure |
| Cognitive disability | one step at a time, no unexpected changes |
The model's own bias is also an accessibility question. A speech recogniser that works worse for a dialect, an accent or a speech difficulty excludes users in practice, even if the interface is perfect. Measure the performance per group, not just overall.
Testing, at three levels:
| Level | What | Catches |
|---|---|---|
| Automatic (axe, Lighthouse) | contrast, labels, roles | ~30 % of the problems |
| Manual | keyboard, screen reader, 200 % zoom | much more |
| With real users | actual use with assistive technology | what no other method finds |
The first level is cheap and belongs in CI — AI-grafen runs npm run a11y with axe-core over every page on each build. But it only catches a third, and the last level is the only one that finds the real problems.
Code
<!-- A streaming AI answer that works with screen readers -->
<div role="region" aria-label="The tutor's answer">
<div id="status" aria-live="polite" class="sr-only">The tutor is writing …</div>
<div id="answer" aria-live="polite" aria-atomic="false"></div>
<button id="cancel">Cancel</button>
</div>
// Buffer to sentence level — otherwise the screen reader re-reads the text on every token
let buffer = "";
const answer = document.getElementById("answer");
const status = document.getElementById("status");
function receive(token) {
buffer += token;
if (/[.!?]\s$/.test(buffer)) { // a whole sentence is ready
const p = document.createElement("p");
p.textContent = buffer.trim();
answer.appendChild(p); // APPEND, do not overwrite
buffer = "";
}
}
function done() {
if (buffer.trim()) {
const p = document.createElement("p");
p.textContent = buffer.trim();
answer.appendChild(p);
}
status.textContent = "The answer is complete.";
}
# Alt text for generated images — with the ability to correct it
def alt_text(image, context: str) -> dict:
description = vlm(
"Describe the image in at most 125 characters for somebody who cannot see it. "
"Mention only what is visible. Write 'not clear' if something is unclear.\n"
f"Context: {context}", image)
return {"alt": description[:125], "generated": True, "correctable": True}
# Measure the model performance PER GROUP — not just overall
def wer_per_group(asr, testset):
import jiwer
from collections import defaultdict
per = defaultdict(list)
for case in testset:
w = jiwer.wer(case["reference"], asr(case["audio"]))
per[case["group"]].append(w) # dialect, accent, age, speech difficulty
return {g: round(sum(v) / len(v), 3) for g, v in sorted(per.items())}
# {'accent': 0.31, 'dialect_norrland': 0.19, 'standard_swedish': 0.08, 'speech_difficulty': 0.44}
# ↑ the overall figure would have hidden that the system is unusable for one of the groups
The last output is the point: an average WER of 0.12 looks fine and hides the fact that the system in practice excludes users with speech difficulties.
Mastery means
- Applies WCAG to AI interfaces
- Handles the AI-specific accessibility problems
- Tests with assistive technology
Sign in to do the exercises and build your mastery up.
Sources
- W3C — WCAG 2.2 — W3C Document License
- DIGG — Guidance on web development and accessibility (in Swedish) — myndighetsmaterial
- MDN — ARIA live regions — CC BY-SA 2.5