Personalized speech recognition for transcription, restatement, and text entry
People with non-standard speech need not rely only on listener familiarity, repetition, or text entry.
Customer service · Multi-country Android beta; Ghana user study
Executive brief
The operating-model shift, in one view.
Assistive AI is not useful merely because it recognizes more speech; onboarding burden, device access, changing voices, and user control determine whether it increases autonomy.
AI value · Feasibility of personalized ASR use in Ghana
Verified
Twenty people with communication difficulties completed a Ghana user study; the study identified feasibility but also access, adaptation, support, and accuracy limitations
The study is a small feasibility evaluation, not proof of universal accuracy or communication benefit; Google currently accepts no new users.
Before
Speak, repeat, gesture, or switch to typing when a listener or device does not understand. → Interpret, ask for clarification, or fail to complete the interaction.
After
Records at least 500 phrases to create a personalized recognition model. → Transcribes speech, restates it in a synthesized voice, or sends text through a keyboard. → Use the output with active listening, gestures, and clarification.
Human boundary
The model proposes a transcript or synthesized restatement; the user decides whether it expresses the intended message.
Why it matters
People with non-standard speech need not rely only on listener familiarity, repetition, or text entry.
How the work changed
Before
How the work ran before the change.
Step 1 of 2
Person with non-standard speech
Speak, repeat, gesture, or switch to typing when a listener or device does not understand.
ControlHuman accommodation
Step 2 of 2
Listener
Interpret, ask for clarification, or fail to complete the interaction.
ControlListener understanding
What changed
People with non-standard speech need not rely only on listener familiarity, repetition, or text entry.
Decision rightHuman moves from creator to judge
After
How the same work runs now.
Step 1 of 3
User
Records at least 500 phrases to create a personalized recognition model.
ControlAndroid, connectivity, and voice stability
Step 2 of 3
Project Relate
Transcribes speech, restates it in a synthesized voice, or sends text through a keyboard.
ControlPersonalized model
Step 3 of 3
User and listener
Use the output with active listening, gestures, and clarification.
ControlHumans retain meaning and communication control
Process model built from the published workflow evidence for Google Research Project Relate. Every step, actor, and control appears in full below.Every step, actor, and control
Exception path
Users repeat, type, gesture, or ask a listener for clarification when recognition is low; severe or rapidly changing speech can make the model unsuitable.
Decision authority
The model proposes a transcript or synthesized restatement; the user decides whether it expresses the intended message.
Before
#
Actor
Action
Control
01
Person with non-standard speech
Speak, repeat, gesture, or switch to typing when a listener or device does not understand.
Human accommodation
02
Listener
Interpret, ask for clarification, or fail to complete the interaction.
Listener understanding
After
#
Actor
Action
Control
01
User
Records at least 500 phrases to create a personalized recognition model.
Android, connectivity, and voice stability
02
Project Relate
Transcribes speech, restates it in a synthesized voice, or sends text through a keyboard.
Personalized model
03
User and listener
Use the output with active listening, gestures, and clarification.
Humans retain meaning and communication control
Work that left the path
Some repeated speech repair
Some manual typing for supported users
Human role before
Users repeatedly repaired communication and listeners inferred meaning without personalized recognition support.
Human role after
Users train and choose when to use the tool; listeners combine its output with active clarification.
AI role
Personalized automatic speech recognition trained on an individual's non-standard speech.
Outcomes
Feasibility of personalized ASR use in Ghana
Verified
General ASR can have 78-89% word-error rates for severely dysarthric speech, with 0-1.2% correctly transcribed sentences in cited prior work→Twenty people with communication difficulties completed a Ghana user study; the study identified feasibility but also access, adaptation, support, and accuracy limitations
Study published at CHI 2024 · 20 users plus locally trained speech and language therapists in Ghana
The study is a small feasibility evaluation, not proof of universal accuracy or communication benefit; Google currently accepts no new users.
What leaders can reuse
Anti-pattern
Treating a transcript as the user's intended message or presenting a closed beta as broadly available.
Questions
01Can the user easily reject or correct the output?
02What happens as the user's voice changes?
Portability conditions
User-specific training data
Stable enough speech pattern
Fallback communication modes
Reputation risk
high
Evidence and authority
What the public record supports.
Watch · updated
Watch status: verify the cited source and deployment condition before reusing this case.
1 peer reviewed, 2 primary; publication outcomes are verified.
Bundle 1.0.0 · reviewed 2026-09-06 · stable ID 33e321501f83de63