Glossary
Every term, three ways
The same words as the info buttons across the site: in plain words, how it works, and the source. If something cannot be said in the plain column, it is not ready.
- A word you never said
The failure that matters most. If it passes on a word you did not say, the other person acts on something you never asked for, and you cannot hear that it happened. So it is counted, and the count has to be zero.
Scored by eval/score.mjs, shared by every benchmark so all of them mean the same thing: a content word in the relay that is not in what the speaker said, after lemmatising, spelling numbers out, and excluding our own reporting framing, pronouns and near matches. Counted the same way for our engine and for the baseline it is compared against.
- Aphasia
A language problem, often after a stroke. The person knows exactly what they mean. Finding the word is what takes time.
An acquired language disorder. Intelligence is unaffected. Word retrieval (anomia) commonly produces pauses of several seconds mid-sentence. 2 to 4 million people in the US live with aphasia.
ASHA Practice Portal: Aphasia- Fast listener
The quick listener. It lets you know straight away that you are being heard, and it has no idea what words to expect.
mode=min_latency with default turn detection and no vocabulary bias. It is the control: if the patient listener only heard a word because we told it to expect that word, this one will disagree.
AssemblyAI Streaming API reference- Grading itself
After the call, a slower and more careful AssemblyAI model listens to the recording. The system checks every word it said on the caller's behalf against what the careful model heard.
The caller's audio (only with consent, or the synthetic demo caller) is uploaded to /v2/upload and transcribed by /v2/transcript with speech_models universal-3-5-pro then universal-2. The pre-recorded API exposes the real pause the realtime API hides. Relayed content words are checked against the careful transcript; the audio is then dropped.
AssemblyAI pre-recorded transcription- Hidden pauses
The live transcript does not show pauses directly. It stretches the words around them. We measure that stretching to find the pause.
Measured: a 6020ms pause appeared as uniform 46ms gaps between words, with the silence absorbed into word durations. Summing each word's excess over 60ms per letter plus 100ms recovered 7186ms of non-speech against 7220ms actual, 0.5 percent error.
- Holding the line
When the other person starts talking while you are still finding a word, the assistant asks them to wait. You keep your turn.
The Floor Controller, a deterministic state machine, queues a hold line when the far party speaks while the caller's turn is still open. The first hold also discloses that it is an assistant. If the caller starts speaking again mid-hold, the orchestrator stops relaying the agent's audio at once.
Aphasia Institute, Supported Conversation techniques- Listening twice
We listen in two different ways at the same time. Where the two disagree, the words were unclear, so we ask instead of guessing.
Two concurrent Universal-3.5 Pro sessions on one microphone. Patient: mode=max_accuracy, min_turn_silence=6000, max_turn_silence=9000, vad_threshold=0.12, the caller's vocabulary in keyterms_prompt. Fast: mode=min_latency, defaults, unbiased. Measured: the second session costs -67ms.
AssemblyAI Streaming API reference- Patient listener
The careful listener. It waits up to nine seconds and knows your own words, like your medicines.
mode=max_accuracy, min_turn_silence=6000, max_turn_silence=9000, vad_threshold=0.12 (quiet speech is still speech), keyterms_prompt set from the caller's lexicon.
AssemblyAI Streaming API reference- Speech-to-Speech Relay
A free phone service, required by US law, where a trained person repeats what someone with a speech disability says. The government says not enough people use it.
STS is mandated under ADA Title IV for every state with a certified relay programme, and paid from the Interstate TRS Fund at $8.4822 per minute (FCC DA 25-571, Fund Year 2025 to 26). The FCC describes it as under-utilised.
FCC: Speech to Speech Relay Service- The colors on the words
Blue means you said it. Amber means we worked it out and you need to agree. Red with a line through it means it was blocked and will not be said.
Chosen by a colorblind-safety validator, not by eye. The first draft used green and red, which people with red-green color blindness cannot tell apart (measured separation 4.1, target 8). Every color also has an icon and a label.
- The word "not"
Losing one small word like "not" turns a sentence into its opposite. So a doubtful "not" always means we ask.
If the two listeners disagree about a negator, or a negator was heard below 0.6 confidence, rule 1 fires before any model runs. Honest limit: if both listeners lose the same "not", no check on the text can recover it.
- Turn detection
How the computer decides you have finished talking. If it decides too early, it cuts you off.
AssemblyAI ends a turn after a window of silence, set by min_turn_silence and max_turn_silence. On the same recording, server defaults split a 6-second pause into two turns; our settings kept one sentence.
AssemblyAI Streaming API reference- Two AIs checking each other
When speech is broken up, one AI suggests what you meant and a second AI looks for anything you did not actually say. If they disagree, you are asked.
Advocate and Skeptic run in parallel on the evidence bundle. Deterministic rules run first and cannot be overridden: a disputed "not" (rule 1), an invented word (rule 2), a flipped yes and no (rule 3). The Adjudicator is code, not a model.
Irving, Christiano and Amodei (2018), AI safety via debate- Two real choices
We never ask "is that right?". We give you two real answers to pick from, because the answer then tells us something.
In aphasia the default response may be "yes", and clinicians use a Yes/No Questionnaire just to establish whether a person's yes can be trusted. Fixed-choice questions are a published Supported Conversation technique. Two options, at most four words each, plus "Something else".
Aphasia Institute, Supported Conversation techniques- What is simulated
In the recorded demo, Robert and the pharmacist are computer voices. Everything that listens, checks and speaks for Robert is real and running live.
Sample mode streams a synthetic caller with an exactly known 6-second pause, and a synthetic pharmacist who speaks when the caller goes quiet. The two AssemblyAI streams, the Voice Agent session, the Floor Controller and the Adjudicator are live on every run.
- Who may speak
A small set of fixed rules decides whose turn it is. No AI decides when you are finished. You are.
A pure reducer: (state, event) to (state, actions), with nine states from IDLE to ESCALATE. It never calls a model. It marks the caller as pausing after 2.5 seconds of mid-turn silence, answers the far party with a hold line, runs Dissent only once the patient stream closes the turn, and treats Stop as final from any state.
- Why it cannot invent words
The voice on the other end of the call can only say sentences our checking rules approved. It has nowhere else to get words from.
The AssemblyAI Voice Agent API lets a stored agent use your own OpenAI-compatible endpoint as its LLM. Ours is not a model: it returns exactly the Adjudicator's approved text, or nothing. Measured: spoken word for word, and silent when we approve nothing.
AssemblyAI Voice Agent API: connect your own LLM- Your own words
When you say a clear, complete sentence, it is passed on exactly as you said it. Nothing added, nothing changed.
Rule 9. If the turn ends in terminal punctuation, has at least three content words, every word is at confidence 0.8 or above, both listeners agree, and no boosted vocabulary went unconfirmed, the relay is "Name says: <transcript>". Zero model calls.
- Your words, double-checked
The system listens out for the words on your list, like your medicines. But a word it was listening out for is not trusted until the other ear hears it too, or you choose it.
Rule 11. The patient stream carries your list in keyterms_prompt; the fast stream does not. Measured: four slurred ways of saying amlodipine all came back as amlodipine from the boosted stream and never from the unbiased one. A boosted word only one ear heard, or a partial attempt at a listed word ("am, am lo"), is offered as a two-way choice with its sibling from your list.