Skip to content
One Moment

Limits

What it does not do

This system speaks on behalf of people who cannot easily correct it. So the limits are not a disclaimer at the end. They are the design.

The failure that matters most

The worst thing this product could do is say something the caller did not mean, to a stranger, in the caller's name, faster than they can correct it. Everything else is secondary. These are the controls, in order of strength.

  1. 1. The voice has no words of its own.

    The AssemblyAI Voice Agent that speaks to the far party uses our endpoint as its language model, and our endpoint is not a model. It returns a line the rules approved, or nothing. The live view checks every reply against AssemblyAI's own transcript of what the agent said.

  2. 2. His own words first.

    A clear, complete sentence is relayed exactly as he said it, with no model involved.

  3. 3. Rules, not a model, decide.

    The Adjudicator is code: a short ordered list of rules anyone can read on the technology page. A disputed "not", an invented word or a flipped yes and no each force a question.

  4. 4. Every doubt leads to asking.

    When the system is unsure, slow, rate limited or broken, it asks the caller with two real choices. It never guesses.

  5. 5. The caller can always stop it.

    A Stop button is on screen whenever it is speaking for him. And the moment he starts talking again, a hold line is cut off, mid-word if it has to be.

None of these is sufficient on its own. The Skeptic will miss things. If both listening streams lose the same “not”, the text no longer contains it and no check can recover it; the benchmark publishes how often. So this is positioned as a first line with a human behind it, never as a replacement for a trained relay assistant.

And here is one the controls above do not catch, in the third recorded call. Robert says “The water pill, the small white one”, blocks for six seconds, and stops at “I want to stop the”. What went out was “He is saying he wants to stop the water pill.” Every word of that is a word he said, both models read it the same way independently, and so rule 8 allowed it. But he never joined “stop” to “the water pill”. The models did. Checking that every word was said does not check the relations between the words, and on a blood pressure tablet that gap matters. It is the clearest argument in the whole project for the human behind the line. Hear it.

What we measured, and what we did not

Scope
ClaimStatus
Tested with people with aphasiaNo. The numbers come from synthetic speech with a known pause, published text corpora, and 120 recorded sentences from 8 speakers with dysarthria (TORGO), which is a motor speech disorder, not aphasia. It is a working prototype, not a clinical study.
A medical deviceNo. It does not diagnose or treat. It is a conversation partner on a phone call.
Better recognition of disordered speechNo. We do not improve recognition, and the design assumes it is often wrong.
Tested against a live pharmacyThe far party is a computer voice, and that is a scoped decision rather than a missing feature. Nothing the far party says is ever relayed or fed to the Adjudicator: it reads the caller's two streams and nothing else. A real pharmacist changes when the hold line fires, not what is said in the caller's name, so every number here is measured on the caller's side.
Real phone callsThe phone leg is built and proven, and not yet connected to a number. An inbound call over Twilio Media Streams is bridged into the same Call object the browser uses: 8kHz mu-law in, re-cut from the 20ms frames Twilio sends into the 50ms the listening streams want, and everything One Moment says mixed back down the line. Proven end to end against the live AssemblyAI APIs by npm run proof:phone, which speaks the Twilio side of the protocol exactly: the 6-second pause survived on telephone audio, the sentence was relayed in his own words, 7.7 seconds of audio came back, and every word spoken was approved text. What is missing is a phone number, which is an account rather than an engineering step.
Works for every kind of aphasiaUnknown. It is designed around word-finding pauses. Fluent aphasia, where speech flows but words are wrong, is a different problem.
Ready for many callers at onceNot on the shared demo key. Every call opens two streaming sessions and our account holds four, so the public demo runs two live calls at a time and refuses the third rather than failing mid-call. Bring your own key and that limit is yours, not ours. The three recorded calls are always available.

Who does the adjacent job better

Voiceitt and Google's Project Relate make distorted speech intelligible, and they have real users. If your words come out unclear but your sentences are complete, use one of those. One Moment is for a different situation: the sentence stops, and the word arrives late. Clearer audio does not help, because the audio was already clear.