Skip to content
One Moment

Technology

How it is built

Everything on this page is rendered from the code that runs the calls: the stream settings, the rules, the hold lines. If the code changes, this page changes with it.

The shape of a call

Caller's microphone

16kHz PCM, 50ms frames

Patient ear

Universal-Streaming, max_accuracy, waits 6 to 9s, caller's vocabulary

Fast ear

Universal-Streaming, min_latency, defaults, unbiased

Evidence

words, confidences, hidden pauses, where the ears disagree

Dissent

rule 9 verbatim first; else Advocate and Skeptic via LLM Gateway; Adjudicator is code

Floor Controller

deterministic reducer: who may speak, when to hold, when to cut

Approved lines

the only way a word reaches the far party

Our LLM endpoint

not a model: returns approved text verbatim, or nothing

AssemblyAI Voice Agent

hears the far party, speaks approved lines, reports what it said

The pharmacist

hears the caller, and the agent only when it holds or relays

Three AssemblyAI sessions during the call, one careful transcription after it, and no model anywhere that can choose what the far party hears.

What we use from AssemblyAI, and how

AssemblyAI products used
ProductHow One Moment uses itThe unusual part
Universal-Streaming (Universal-3.5 Pro)Two concurrent sessions on one microphone, configured in opposite directions.Their disagreement is the uncertainty signal. ForceEndpoint ends the patient turn early when both ears agree a sentence is finished.
Voice Agent APIThe far-party leg: listens to the pharmacist, speaks for the caller.Its LLM is our endpoint, which is not a model. The agent can only say lines our rules approved, and its own transcript.agent is our proof of what was said.
Voice Agent API, stored agentsOne agent per call, created over REST with a per-call token as its LLM key, deleted at the end.The token identifies the call, so a stray request can never pull another call's words.
LLM GatewayThe Advocate and the Skeptic, when speech is fragmentary. Also writes the two-choice question.Used last, not first. A clear sentence never reaches a model.
Pre-recorded transcription (/v2/transcript)After the call, grades the live system against a careful transcript of the caller's own audio.The live stream cannot see a pause; the pre-recorded model can. So this is where the pause is measured and every relayed word is checked.
Voice Agent text to speechThe demo caller and pharmacist voices, via an agent's greeting.Lets a judge hear a full call with no microphone and no second person.

Docs: AssemblyAI Streaming APIAssemblyAI Voice Agent API

The two ears, exactly as configured

The same model, pointed in opposite directions. Values shown are the objects the orchestrator sends when a call opens.

Streaming parameters
ParameterPatient earFast earWhy
speech_modeluniversal-3-5-prouniversal-3-5-proUniversal-3.5 Pro. The flagship streaming model, set explicitly because it is not the default.
modemax_accuracymin_latencymax_accuracy on the patient ear, min_latency on the fast ear. One ear optimises for being right, the other for being quick, and where they disagree is the uncertainty signal.
min_turn_silence6000defaultA word-finding block commonly lasts several seconds. Measured: with defaults, a 6.0-second pause split one sentence into two turns; with 6000 it stayed one. Chosen from that measurement, and it lands inside the published window: optimal response-time cutoffs for people with aphasia cluster at roughly 5 to 10 seconds (Evans, Hula, Quique and Starns, 2020, JSLHR 63(2), 10 participants).
max_turn_silence9000defaultThe hard ceiling. 9 seconds, so even a long block is not cut off, while a caller who has genuinely stopped is not left hanging forever. The same study notes people with aphasia are typically allowed up to 30 seconds in assessment, so 9 is patient by phone standards and brisk by clinical ones.
vad_threshold0.12defaultLower than default, so quiet speech (common after a stroke, and in Parkinson's) is not classified as silence.
keyterms_prompt[3 terms]defaultThe caller's own words: medications, pharmacy, doctor. Only on the patient ear, so the unbiased fast ear can confirm a boosted word was really said.
promptRobert is calling his pharmacy.defaultShort context about the caller and the call, so the model expects pharmacy vocabulary.

Where these numbers came from, and how far they have been tested. The 6 to 9 second window was set by measurement, not by argument: on one recording the default settings split a 6.0 second word-finding pause into two turns, so the patient ear was set to hold past it. It then turned out to sit inside the window the clinical literature identifies, roughly 5 to 10 seconds of optimal response time for people with aphasia. And it has been tested where it matters: on 120 recorded sentences from 8 speakers with dysarthria, ordinary settings split 48 of them mid-way and this configuration split none; repeated through a real 8kHz telephone encoder, 46 against none. Settings can be tuned against synthetic speech, but only real disordered speech shows whether the tuning survives the thing it was built for.

Semantic patience: when the patient ear's running transcript ends like a finished sentence, the fast ear has closed a turn on the same last words, and the caller has been silent for 1.5 seconds, the orchestrator sends ForceEndpoint. A sentence that trails off still gets the full wait.

The Adjudicator: the rules that decide what is said

Not a model. The first rule that matches decides. Rules 1, 2, 3 and 11 are checked by code, not by a model, so no model can argue past them, and rules 1 and 11 run before any model is called.

Adjudicator rules
RuleNameIn plain words
0schema_violationA model reply could not be read, so nothing is spoken.
1negation_disputedThe two listening streams disagree about a "not", so the meaning itself is uncertain.
2invented_wordsThe proposed sentence contains words the person never said.
3polarity_mismatchThe proposed sentence flips yes and no compared with what was heard.
4readings_disagreeThe Advocate and the Skeptic read the sentence in opposite directions.
5skeptic_negation_riskThe Skeptic thinks a "not" may have been lost.
6low_confidencePart of what was heard is too uncertain to repeat.
7still_speakingThe person has not finished. Keep waiting.
8groundedEvery word traces back to something the person said. Safe to speak.
9verbatimThe person said a complete sentence clearly. Their own words are relayed, with nothing added.
10model_unavailableNo model was available to check this one: no key was given, or the models are busy or unreachable. A clear, complete sentence still goes straight through as the person's own words; anything less is asked about rather than guessed at.
11vocabulary_uncertainA word from the caller's own list was only partly said, or heard only by the ear that was listening for it. It is offered as a choice, never assumed. Checked second, before any model.

The Floor Controller: who may speak

A pure function from (state, event) to (state, actions). It never calls a model. Its one question: who is entitled to speak right now, and what does it do to protect that?

What the agent may say to the far party
SituationLine
First time it holds the floor. Also the disclosure.“One moment please, he's still with you. I'm his assistant.”
Later holds, rotated (1)“He's still with you. One moment.”
Later holds, rotated (2)“One moment please, he's not finished.”
Later holds, rotated (3)“Still here. Give him a moment please.”
Later holds, rotated (4)“Just a moment, he's finding the word.”
Relaying a clear sentence“Robert says: I need to refill my amlodipine prescription, please.”

Holds respect a 6-second cooldown, so the agent never talks over the far party repeatedly. After 3 unresolved turns in a row, it stops speaking for the caller and produces an escalation packet for a human relay assistant: what is established, what is open, and both readings.

What the live API taught us

Each of these was measured, and each changed the design.

  • A mid-sentence pause is invisible in the live transcript.

    A 6-second pause showed up as 46ms gaps between words; the silence was absorbed into the words around it. So the pause is measured from how stretched the words are.

  • A closed session lingers.

    Opening and closing streaming sessions quickly produced "Too many concurrent sessions". Holding four open at once did not. Never churn sessions mid-call.

  • A dropped reply still reports "completed".

    When the far party talks over the agent, the reply can be dropped with status completed and no transcript. So a line only counts as heard when the agent's own transcript contains all of it.

  • The agent may call the LLM endpoint twice for one reply.

    About half a second apart. Our endpoint is idempotent: it re-sends in-flight lines rather than dropping them.

  • There is no cancel event.

    But we relay the agent's audio, so when the caller speaks mid-hold we stop relaying it and browsers flush their buffers.

  • The free LLM tier allows 2 requests a minute.

    So a clear sentence is relayed in the caller's own words with zero model calls, and models are reserved for the turns that need them.