Skip to content
One Moment

Every voice agent is built to stop listening. This one is built to wait.

One Moment waits as long as a stroke survivor needs, offers two choices when the word will not come, asks the pharmacist to hold the line, and speaks only the words he actually said.

Built on AssemblyAI Universal-Streaming and the Voice Agent API. Open source, MIT.

The same recording, with a 6.0-second pause in the middle of the sentence

AssemblyAI default settings: 2 turns

  1. “I need to refill my...”
  2. “amlodipine prescription please.”

The agent treats him trailing off as a finished turn, and starts talking.

One Moment's patient ear: 1 turn

“I need to refill my amlodipine prescription, please.”

He keeps his turn. The sentence arrives whole.

Measured, on speech we did not record

The number that matters is a zero.

Speaking for someone is not like answering them. One invented word and the relay is worse than useless, because the caller cannot hear what was said in their name. So the first thing we measured is what it refuses to say.

0

Words put in the caller's mouth

Across 48 relays of degraded speech. A single model asked what the person meant added them in 11.

0

Meanings flipped

Of 30 relayed sentences carrying a “not”. The same model flipped 15, every one of them a sentence where a word had been lost.

0

Unnecessary questions on clear speech

All 16 of 16 clear sentences went straight through in the speaker's own words, with no model call. Caution that got in the way would be its own failure.

0

False alarms from the self-audit

On 46 relays that were in fact correct, on real disordered speech. What it missed is published too.

On the hard cases it buys those zeros by asking instead of speaking: it asked on 63 of 120 real dysarthric sentences, and spoke on 57. Every figure here is derived from the committed run files, and the failures are published beside the wins: the evidence.

Understand this in 60 seconds

The word is not lost. It is late.

The problem

Robert had a stroke. His thinking is fine. His words are not. When he speaks, the word he wants goes missing for six seconds, and then it arrives.

Two to four million people in the US live with aphasia. ASHA Practice Portal

Why nobody has solved it

It is not his voice. His words are clear when they come.

It is the clock. In conversation the usual gap between turns is about 208 milliseconds, and long gaps read as trouble. Stivers et al. 2009 Voice agents are built for that clock. On a recording of Robert's sentence, the default settings ended his turn in the middle of it.

What we built

One Moment joins the call. It waits as long as he needs. And when the pharmacist starts to fill the silence, it speaks to her:

“One moment please, he's still with you.”

She waits. He finishes. He made the call.

The best-evidenced way to help someone with aphasia talk is to train the other person in the conversation. You cannot train the pharmacist. So we put a trained conversation partner inside the call.

Two systematic reviews of communication partner training, 56 studies between them, all report positive outcomes.Simmons-Mackie et al. 2010Simmons-Mackie et al. 2016 No study has tested an automated partner. This is an untested application of a well-evidenced principle, and we say so.

Watch it happen

A real call, recorded, and scrubbable.

This is recorded data, not a video. The audio is one real call through the live system, and everything on screen is replayed from every event that call produced. Drag the timeline, or step moment by moment with the arrow keys. The caller and the pharmacist are computer voices; everything that listens, decides and speaks for Robert was running live.

Press play to hear the call

About 32 seconds. Sound on. Or step through it with the arrows.

0:00.0 / 0:31.7

Robert's screen

Ready when you are.

Who may speakCurrent state: Waiting

What the other person heard

The call has not started.

The record every call leaves behind: each line approved for the agent to say, whether it was spoken or cut off, every question the caller was asked and what went out as a result, and the post-call audit's verdict on each relay. It re-checks the guarantee that the voice said no word the rules did not approve.

Want the full engine view, with both listening streams, the pause chart and the decision rules? Run a call yourself.

How it works

Six things a good conversation partner does.

  1. 1. It listens twice

    Two AssemblyAI streams on one microphone: a patient ear that waits up to nine seconds, and a fast ear set like any other agent. Where they disagree, the audio was unclear.

  2. 2. It waits, and knows when to stop

    An unfinished sentence gets the full wait. A finished one, agreed by both ears, ends after 1.5 seconds of silence, so the other person is not left in dead air.

  3. 3. It holds the line

    If the other person starts talking into the pause, it asks them to wait. The moment the caller speaks again, it stops, mid-word if it has to.

  4. 4. It uses his own words

    A clear, complete sentence is passed on exactly as he said it. No model is involved, so no model can add a word.

  5. 5. It argues with itself

    When speech is broken up, one AI proposes a meaning and a second looks for anything he did not say. A fixed rule decides. When in doubt, it asks him, with two real choices.

  6. 6. It grades itself

    After the call, a slower, more careful AssemblyAI model transcribes the recording. Every word said on his behalf is checked against it, and the pause the live stream could not see is finally measured.

Why the voice cannot invent words

The voice the pharmacist hears is an AssemblyAI Voice Agent. Its “language model” is our endpoint, and our endpoint is not a model. It returns exactly the line our rules approved, or nothing. We tested this against the live API: the agent spoke our text word for word, and stayed silent when we approved nothing.AssemblyAI Voice Agent API

Why it never asks “is that right?”

In aphasia the default answer is often yes, whether or not it is right. So it never asks for a yes. It offers two real choices, at most four words each, and the answer tells it something either way. Fixed-choice questions are a published Supported Conversation technique. Aphasia Institute, SCA

Measured, not claimed

Every number here came from the live API.

2 vs 1

Turns for one sentence

Default settings split a 6.0s pause into two turns. Our patient ear kept one.

46ms

What a 6-second pause looks like live

The realtime transcript hides the pause inside the words around it, so we measure how stretched the words are.

0.5%

Error recovering the hidden pause

7186ms of silence recovered from word timings, against 7220ms actual.

-67ms

Cost of listening twice

A second, differently configured stream added no measurable latency to the first result.

0

Model calls to relay his sentence

On the recorded call, a clear and complete sentence was relayed in his own words by rule 9, decided in under a millisecond.

1.5s

Wait after a finished sentence

Both ears must agree it is finished. An unfinished one still gets the full 6 seconds and more.

4 of 4

Slurred medicine names the boosted ear 'heard' as amlodipine

The unbiased ear heard it 0 times. So a word only the listening-for-it ear heard is asked about, never trusted.

6.3s

The pause, measured after the call

AssemblyAI's pre-recorded model heard it; the live stream showed 45ms. Every relayed word was confirmed.

0

Words the agent may choose itself

The Voice Agent's only source of words is the line our rules approved. Checked live, every call, against AssemblyAI's own transcript of the agent.

On real disordered speech

120 recorded sentences from 8 speakers with dysarthria (TORGO), through both live ears. An ordinary agent would have relayed words the speaker never said, or flipped the meaning, in 65 of 120 sentences. One Moment did in 11 of the 57 it spoke, and asked the caller on 63.

How it was measured, and every sentence

Measured 17 and 18 September 2026 on synthetic speech with an exactly known pause. The scripts are in the repository and re-run against your own key. Synthetic speech is clean; real disordered speech will score lower on recognition, which is why nothing here depends on recognition being right.

Who pays for this

A service already exists, is required by law, and is under-used.

2 to 4M

People in the US with aphasia

Plus about 7.5 million with trouble using their voices.

$8.4822

Per minute, today

What the federal relay fund pays for Speech-to-Speech Relay, a trained human repeating the caller's words.

Required

In every state

Speech-to-Speech Relay is mandated under ADA Title IV, reachable by dialling 711.

Under-used

The regulator's own words

The FCC describes the service as under-utilised despite outreach.

Sources: ASHA Practice PortalNIDCDFCC DA 25-571FCC, Speech to Speech Relay One reason the service is under-used: you wait for a trained stranger before you can make a call, and that stranger hears your medical conversation. One Moment costs two streaming sessions and one voice agent session per minute. Most turns need no model call at all. We are not proposing to replace the human. Automate the routine calls, and a trained human is free for the ones that need one.

What this is not

Honest limits, stated as measurements.

  • Not tested with people with aphasia.

    Every number on this page comes from synthetic speech, published text corpora, and 120 recorded sentences from speakers with dysarthria (a motor speech disorder, not aphasia), with ground truth we did not write. It is a working prototype, not a clinical study.

  • Not a medical device.

    It does not diagnose or treat anything. It is a conversation partner on a phone call.

  • Not better recognition.

    We do not improve speech recognition on disordered speech. The design assumes recognition is often wrong, and refuses to speak when it is unsure.

  • The demo far party is simulated.

    The pharmacist in the demo is a computer voice. The system has not been tested against a live pharmacy.

If your words come out distorted but your sentences are complete, tools like Voiceitt or Google's Project Relate serve that problem well and have real users. One Moment is for a different situation: the sentence stops, and the word arrives late. Every limit, in full.

Questions

The ones a sceptic asks.

Is this a real product or a hackathon demo?

A working prototype, built for the AssemblyAI Voice Agent Hackathon. The engine runs against the live AssemblyAI APIs. It has not been used by a person with aphasia, and it is not a medical device.

What happens when the speech recognition is wrong?

It is wrong often. That is the whole design problem. Nothing is said to the other person unless it traces back to words in the audio, and when it does not, the system asks the caller instead of guessing.

Could it say something I did not say?

That is the failure we designed against hardest. A clear sentence is relayed in your own words with no model involved. When a model is involved, a second model looks for words you did not say, and the decision to speak is made by fixed rules you can read, not by a model. The voice itself can only speak lines those rules approved.

Why two speech recognition streams? Is that not wasteful?

It adds one streaming session to a call that today costs $8.4822 a minute in human time. And where the two streams disagree is the only honest measure of uncertainty the system has.

Why does it never ask me yes or no?

Because in aphasia "yes" is often the default answer, whether or not it is right. Clinicians use a questionnaire just to find out whether a person's yes can be trusted. Two real choices avoid the problem.

What happens when it stops mid-sentence?

The hold line exists to protect your turn, so it must never talk over you. When you start speaking again, the orchestrator stops relaying the agent's audio at once. If that cut off the part where it said it is an assistant, it says so again before it next speaks for you.