Skip to content
One Moment

For judges

The three-minute path

One Moment is a voice agent whose job is to make the other person wait, so people with aphasia can make their own phone calls. Here is the fastest way to see that it works, and to check that we are not overclaiming.

Three steps

  1. 1

    Hear the three recorded calls (2 minutes)

    Real calls through the live system, replayed with their audio and every event they produced. In the first, the hold line is cut off the instant Robert finds his word. In the second, the medicine name only half comes out, and it asks. In the third his sentence never finishes, so neither shortcut applies and both models run, which is the only one of the three where you can watch the Advocate and the Skeptic read the same evidence.

  2. 2

    Run one live (1 minute)

    “Play the recorded call” drives the live engine with a recorded caller. Or use your microphone: start a sentence, stop for a few seconds, and let the simulated pharmacist try to talk over you.

  3. 3

    Check one claim (1 minute)

    Every number links to the script that measured it. The benchmark shows the cases where we fail, not only the ones where we win.

Against the four criteria

Judging criteria
CriterionWhat to look at
Application of technologyTwo concurrent Universal-Streaming sessions configured in opposite directions, their disagreement used as the uncertainty signal, and keyterms_prompt on one only, so the other can check a boosted word was really said. ForceEndpoint for semantic patience. A Voice Agent whose LLM is an endpoint that is not a model. The pre-recorded API grading every call afterwards. Technology
Business valueSpeech-to-Speech Relay is federally mandated and paid at $8.4822 a minute, and the FCC says it is under-used. One Moment costs two streaming sessions and one Voice Agent session a minute, and usually no model call. Who pays
OriginalityEvery voice agent is tuned to respond faster. This one is tuned to hold the floor for someone else, and to say nothing it was not given. The design is the published Supported Conversation technique, made into software. And it publishes its own failure: the one case no live check can catch, and how the self-audit caught it.
PresentationTwo views of one call: what the caller sees (three things on screen, no timers) and what is really happening. Light and dark, phone and desktop, an info button on every term, and our own accessibility audit.

What is real, and what is simulated

Real versus simulated
PartIn the demo
Both AssemblyAI listening streamsReal, live, every call
The AssemblyAI Voice Agent speaking to the far partyReal, live, every call
The Floor Controller, Dissent and the AdjudicatorReal, live, every call
The caller in "Play the recorded call"A computer voice with an exactly known 6-second pause
The pharmacistA computer voice that speaks when the caller goes quiet
The caller in "Use my microphone"You

The numbers, briefly

  • Same recording, 6.0-second pause: default settings made 2 turns of one sentence; the patient ear made 1.
  • The live transcript hides that pause as 46ms gaps. Word durations recover 7186ms of the 7220ms.
  • On the recorded call, the sentence was relayed in the caller's own words with zero model calls.
  • The benchmark, on the product's own engine, with the failures included: NEGBENCH.
  • On 120 real recorded sentences from 8 speakers with dysarthria (TORGO): an ordinary agent would have relayed something wrong in 65 of 120; One Moment in 11 of the 57 it spoke, asking on the rest. Ordinary settings split 48 sentences mid-way; the patient ear split 0. AUDIOBENCH.

Run it yourself

cd one-moment && npm install
cp .env.example .env     # your AssemblyAI key
npm test                 # the engine: floor, evidence, Dissent, rule 11, semantic patience
npm run orchestrator     # opens a public tunnel so the Voice Agent can reach it
npm run web              # http://localhost:3000/demo

MIT licensed. Built for the AssemblyAI Voice Agent Hackathon, September 2026. Not a medical device.