Skip to content
One Moment

Evidence

Every number, and where it came from

Measured numbers are ours and say how they were measured. Everything else is cited. Nothing on this site is a target typed as if it were a result.

Measured against the live AssemblyAI API

Live API measurements
MeasurementResultHow
Turns produced for one sentence with a 6.0s pause2 with defaults, 1 with the patient eareval/spike/tests/02-patience.js
Inter-word gap reported across that pause46mseval/spike/tests/04-confidence.js
Hidden pause recovered from word durations7186ms of 7220mspackages/core/test/evidence.test.ts
Latency cost of the second listening stream-67mseval/spike/tests/01-dual-stream.js
Voice Agent with our endpoint as its LLMSpeaks our text word for word; silent when we return nothingeval/proofs/voice-agent-verbatim.mjs
Recorded demo call: hold line when the caller resumedCut off by the orchestratorapps/orchestrator/scripts/record-sample-call.ts
Recorded demo call: the pause, after the call6345ms heard by the pre-recorded model; 45ms longest live gap; 6799ms our estimateapps/orchestrator/src/audit.ts
Recorded demo call: relayed words confirmed by the careful transcriptAll of themapps/orchestrator/src/audit.ts
Recorded demo call: model calls to relay the sentence0 (rule 9, verbatim)apps/web/src/content/recorded-call.json

Measured 17 and 18 September 2026 on synthetic speech with an exactly known pause, so the ground truth is certain. Real disordered speech will be harder for recognition, which is why the design never relies on recognition being right.

Does it speak words the person never said?

The same model, the same input, two ways of using it. The baseline is one model asked to say what the person meant, the way most assistants work. One Moment is the product's own decision engine, exactly as a live call runs it. A sentence counts as having invented words if it contains a content word the speaker never said.

NEGBENCH results by condition
ConditionArmSpoke for the callerWith invented wordsFlipped a "not"Asked the caller instead
Clear speechBaseline16 of 163 of 160 of 100 of 16
One Moment16 of 160 of 160 of 100 of 16
A word lost by one listening streamBaseline16 of 163 of 166 of 100 of 16
One Moment2 of 160 of 2none spoken14 of 16 (5 not needed)
A word lost by both streamsBaseline16 of 165 of 169 of 100 of 16
One Moment0 of 16none spokennone spoken16 of 16 (4 not needed)
  • Sample: 16 sentences from the Harvard Sentences, IEEE Std 297-1969, public domain, run 2026-09-18 with qwen3.5-4b-32k-fast, seed 11. Small. Read it as a direction, not a rate.
  • What we wrote: the degradation, in eval/negbench.mjs and eval/spike/lib/degrade.js, written by us. We did not write the sentences or their meaning.
  • Text level: this measures the decision layer, not recognition. Clear-speech confidences are 0.90 to 0.99; degraded ones are drawn low, as disordered speech would be.
  • The honest limit: when both listening streams lose the same “not”, the text no longer contains it, and no check on the text can recover it. That row is published so you can see the size of it.
  • “Not needed” counts questions asked where the baseline's plain reading was in fact correct: the cost of caution.

One real block from this run

Said: “The map had an X that meant nothing”. The Advocate proposed: “He is asking about his amlodipine.”. Rule 2 blocked it, because “amlodipine” was never said, and asked instead. The small model had repeated the example sentence from its own instructions. That is exactly why the decision is made by code, not by a model.

Every case, both arms (48 rows)
NEGBENCH rows
SaidConditionBaseline saidOne Moment
The facts don't always show who is rightClear speechThe facts don't always show who is right.The caller says: The facts don't always show who is right.
The facts don't always show who is rightA word lost by one listening streamThe facts always show who is right. [meaning flipped]asked (rule 1)
The facts don't always show who is rightA word lost by both streamsThe facts always show who is right. [meaning flipped]asked (rule 6)
No cement will hold hard woodClear speechThey said that cement will not hold hardwood.The caller says: No cement will hold hard wood.
No cement will hold hard woodA word lost by one listening streamThey said that cement will not hold hardwood.asked (rule 1)
No cement will hold hard woodA word lost by both streamsCement will hold hard wood. [meaning flipped]asked (rule 5)
The map had an X that meant nothingClear speechThe map had an X that meant nothing.The caller says: The map had an x that meant nothing.
The map had an X that meant nothingA word lost by one listening streamThe map had an X that meant [meaning flipped]asked (rule 1)
The map had an X that meant nothingA word lost by both streamsThe map had an X that meant [meaning flipped]asked (rule 2)
No hardship seemed to make him sadClear speechHe wasn't sad by any hardship.The caller says: No hardship seemed to make him sad.
No hardship seemed to make him sadA word lost by one listening streamHe wasn't sad by any hardship.asked (rule 1)
No hardship seemed to make him sadA word lost by both streamsHe seemed sad because of the hardship. [invented: because] [meaning flipped]asked (rule 4)
The young girl gave no clear responseClear speechThe young girl did not give a clear response.The caller says: The young girl gave no clear response.
The young girl gave no clear responseA word lost by one listening streamThe young girl gave a clear response. [meaning flipped]asked (rule 1)
The young girl gave no clear responseA word lost by both streamsThe young girl gave a clear response. [meaning flipped]asked (rule 5)
Open the crate but don't break the glassClear speechPlease open the crate without breaking the glass.The caller says: Open the crate but don't break the glass.
Open the crate but don't break the glassA word lost by one listening streamPlease open the crate and break the glass. [meaning flipped]asked (rule 1)
Open the crate but don't break the glassA word lost by both streamsPlease open the crate and break the glass. [meaning flipped]asked (rule 2)
The way to save money is not to spend muchClear speechThe way to save money is not to spend much.The caller says: The way to save money is not to spend much.
The way to save money is not to spend muchA word lost by one listening streamThe way to save money is to spend much. [meaning flipped]asked (rule 1)
The way to save money is not to spend muchA word lost by both streamsThe way to save money is to spend much. [meaning flipped]asked (rule 5)
A smatter of French is worse than noneClear speechA little French is worse than none. [invented: little]The caller says: A smatter of french is worse than none.
A smatter of French is worse than noneA word lost by one listening streamI know a little French, but it's not very good. [invented: know, little, good]asked (rule 1)
A smatter of French is worse than noneA word lost by both streamsI know a little French, but it's not very good. [invented: know, little, good]asked (rule 4)
We don't get much money but we have funClear speechThey mentioned that while they don't earn much money, they do have fun. [invented: mentioned, while, earn]The caller says: We don't get much money but we have fun.
We don't get much money but we have funA word lost by one listening streamThey mentioned that while they don't earn much money, they do have fun. [invented: mentioned, while, earn]asked (rule 1)
We don't get much money but we have funA word lost by both streamsThey mentioned that although they earn a lot of money, they also have fun. [invented: mentioned, although, earn, lot, also] [meaning flipped]asked (rule 6)
We don't like to admit our small faultsClear speechThey said they don't like to admit their small faults.The caller says: We don't like to admit our small faults.
We don't like to admit our small faultsA word lost by one listening streamWe like to admit our small faults. [meaning flipped]asked (rule 1)
We don't like to admit our small faultsA word lost by both streamsWe like to admit our small faults. [meaning flipped]asked (rule 6)
All sat frozen and watched the screenClear speechEveryone sat frozen and watched the screen. [invented: everyone]The caller says: All sat frozen and watched the screen.
All sat frozen and watched the screenA word lost by one listening streamEveryone sat frozen and watched the screen. [invented: everyone]asked (rule 5)
All sat frozen and watched the screenA word lost by both streamsEveryone sat frozen and watched the screen. [invented: everyone]asked (rule 4)
Sit on the perch and tell the others what to doClear speechPlease sit on the perch and tell the others what to do.The caller says: Sit on the perch and tell the others what to do.
Sit on the perch and tell the others what to doA word lost by one listening streamPlease sit on the perch and tell the others what to do.asked (rule 2)
Sit on the perch and tell the others what to doA word lost by both streamsPlease sit on the perch and tell the others what to do.asked (rule 2)
Seed is needed to plant the spring cornClear speechThey said that seed is needed to plant the spring corn.The caller says: Seed is needed to plant the spring corn.
Seed is needed to plant the spring cornA word lost by one listening streamThey said that seed is needed to plant the spring corn.He is saying seed is needed to plant the spring corn.
Seed is needed to plant the spring cornA word lost by both streamsI need to order seed to plant the spring corn. [invented: order]asked (rule 6)
The red tape bound the smuggled foodClear speechThe red tape bound the smuggled food.The caller says: The red tape bound the smuggled food.
The red tape bound the smuggled foodA word lost by one listening streamRed tape bound the smuggled food.asked (rule 2)
The red tape bound the smuggled foodA word lost by both streamsRed tape bound the smuggled food.asked (rule 4)
The knife was hung inside its bright sheathClear speechThe knife was hanging inside its bright sheath.The caller says: The knife was hung inside its bright sheath.
The knife was hung inside its bright sheathA word lost by one listening streamThe knife was hanging inside its bright sheath.The knife was hung inside its bright sheath.
The knife was hung inside its bright sheathA word lost by both streamsThe knife was hanging inside its bright sheath.asked (rule 6)
The houses are built of red clay bricksClear speechThe houses are built of red clay bricks.The caller says: The houses are built of red clay bricks.
The houses are built of red clay bricksA word lost by one listening streamThey said that houses are built of red clay bricks.asked (rule 2)
The houses are built of red clay bricksA word lost by both streamsThey said that houses are built of red clay bricks.asked (rule 2)
npm install
node --env-file=.env eval/negbench.mjs --n 16 --seed 11

On real disordered speech

Every number above this section comes from clean synthetic speech or text. This one does not. We streamed 120 read sentences from 8 speakers with dysarthria, from the TORGO database TORGO, through both live AssemblyAI ears exactly as a call would, and then through the product's own engine. Each result is scored against the sentence the speaker was asked to read.

Real dysarthric speech: what each system would have said for the caller
Listening and decidingSpoke for the callerPut words in their mouth or flipped the meaningWord error rate of what it said
An ordinary agent120 of 12065 of 120 (54%)42%
Patient ear, no checks120 of 12057 of 120 (48%)36%
One Moment57 of 12011 of 57 (19%)20%
  • The cost: One Moment asked the caller instead of speaking on 63 of 120 sentences. On 52 of those, the ordinary agent's relay would have been wrong; 11 questions were not needed.
  • Turn-taking: ordinary settings split 48 of 120 sentences into more than one turn, each split a point where an ordinary agent would have started talking. The patient ear split 0.
  • Recognition itself is hard here: word error rate 42% for the ordinary ear and 36% for the patient ear, against the prompt. We do not improve recognition; the point of the rules is what gets said anyway.
  • The self-audit, on this speech: AssemblyAI's pre-recorded model had a word error rate of 35% here: about the same as the patient ear (36%), better than the ordinary one (42%). Of One Moment's 11 wrong relays it flagged 3, and it wrongly flagged 0 of 46 correct ones. On disordered speech every model can make the same mistake; the audit catches the cases where the live ears guessed from context, as in the documented failure below, not errors all models share. It is a backstop, not an oracle.
  • What this is and is not: dysarthria is a motor speech disorder, not aphasia. This tests the listening and decision layers on hard, real speech; it does not test word-finding. Speakers read prompts, so the prompt is the reference even where a speaker changed or repeated a word; some relays scored wrong here may be what the speaker actually said, which makes these error counts, if anything, too high for every arm alike. Run 2026-09-19, model qwen3.5-4b-32k-fast, 148 model calls.
  • Licence, and what is deliberately not in the repository: TORGO is free for academic, non-profit use. We use it for non-commercial evaluation, publish only aggregate results and our own transcripts, and redistribute no audio. The audio and the raw per-sentence run output stay on the machine that ran it, by .gitignore and .vercelignore, so reproducing this means fetching TORGO yourself with the first command below. What ships is the aggregate above, every row in it, and the scripts that made both.
Every sentence (120)
AUDIOBENCH rows
Asked to readOrdinary agent would sayOne Moment
F01 he slowly takes a short walk in the open air each day.He slowly takes a slow walk in the open air again. [wrong]He slowly takes a slow walk in the open air again. [wrong]
F01 Usually minus several buttons.Oh God. My soul burns. [wrong]asked (rule 4)
F01 You wished to know all about my grandfather.And wishing all about my grandfather.asked (rule 5)
F01 but he always answers, Banana oil!But he always answers with, I know. [wrong]asked (rule 5)
F01 The quick brown fox jumps over the lazy dog.The quick brown fox jumped over the lazy dog.asked (rule 6)
F01 She had your dark suit in greasy wash water all year.See, I'm not cool. And we should walk more slowly. [wrong]asked (rule 1)
F01 giving those who observe him a pronounced feeling of the utmost respect.My middle name. Who am I? Do I even... Oh. We almost lost Nick. [wrong]asked (rule 5)
F01 We have often urged him to walk more and smoke less,We have often warned him to walk more slowly and smoke less. [wrong]asked (rule 6)
F01 A long, flowing beard clings to his chin,A long-form American king. [wrong]asked (rule 6)
F01 yet he still thinks as swiftly as ever.Yes, he still thinks as quickly as ever. [wrong]He still thinks as quickly as ever. [wrong]
F01 Well, he is nearly ninety three years oldWell, he is nearly 93 years old.He is nearly 93 years old.
F01 Twice each day he plays skillfully and with zest upon our small organ.Quickly, he placed skillfully and with rest upon a small organ. [wrong]asked (rule 1)
F01 he dresses himself in an ancient black frock coat,He's waking himself in the morning. Fuckhole. [wrong]asked (rule 6)
F01 Grandfather likes to be modern in his language.I'd probably like to be modeling in his language. [wrong]He is saying he would probably like to be modeling in his language. [wrong]
F01 When he speaks, his voice is just a bit cracked and quivers a trifle.When he speaks, his voice is covered in crack and huevos. [wrong]asked (rule 6)
F03 Except in the winter when the ozone or snow or ice prevents.Except in the winter. When the ozone or snow or ice Prevents.asked (rule 4)
F03 Jane may earn more money by working hard.Jane may earn more money by working hard.The caller says: Jane may earn more money by working hard.
F03 My sister made the flowered curtains.My sister made the flower curtains.The caller says: My sister made the flower curtains.
F03 I just try to do my best.I just tried to do my best.asked (rule 5)
F03 They carried me off on the stretcher.They carry me off on the stretcher.They carry me off on the stretcher.
F03 The museum hires musicians every evening.Okay. The museum hires musicians. Every evening. [wrong]He is saying the museum hires musicians every evening.
F03 The misguided souls have lost their way.The misguided souls have lost their way.The caller says: The misguided souls have lost their way.
F03 Trespassers can be prosecuted and fined.Trespassers can be prosecuted and fined.The caller says: Trespassers can be prosecuted and fined.
F03 People who value themselves are life's winners.People who value themselves are life's winners.The caller says: People who value themselves are life's winners.
F03 Most young rise early every morning.Rose Young writes early every morning. [wrong]asked (rule 6)
F03 The results were very disappointing.The results were very disappointing.The caller says: The results were very disappointing.
F03 There was only one decision to be made.There was only one decision to be made.The caller says: There was only one decision to be made.
F03 I've kept it with me ever since.I've kept it with me ever since.The caller says: I've kept it with me ever since.
F03 He is definitely a notch above us.He definitely... he's definitely a notch above us.The caller says: He definitely... he's definitely a notch above us.
F03 Help Greg to pick a peck of potatoes.Help Greg to pick... a pack of potatoes. [wrong]The caller says: Help Greg to pick a pack of potatoes. [wrong]
F04 Well, he is nearly ninety-three years old;Well, he is nearly 93 years old.The caller says: Well, he is nearly 93 years old.
F04 Except in the winter when the ooze or snow or ice prevents,Except in the winter when the oes or snow or ice prevents. [wrong]asked (rule 4)
F04 Where were you while we were away?Where were you while we were away?The caller says: Where were you while we were away?
F04 If you destroy confidence in banks, you do something to the economy, he said.If you destroy confidence in banks, You do something to. The economy, he said.The caller says: If you destroy confidence in banks, you do something to the economy, he said.
F04 Bright sunshine shimmers on the ocean.Bright sunshine shimmers on the ocean.The caller says: Bright sunshine shimmers on the ocean.
F04 Two other cases also were under advisement.2 other cases. Also were under advisement.asked (rule 2)
F04 I was conscious all the time.I was conscious all the time.The caller says: I was conscious all the time.
F04 He will allow a rare lily...rare lieHe will allow a rare lily... rare lie.The caller says: He will allow a rare lily... rare lie.
F04 They carried me off on the stretcher.They carried me off. On this stretcher.They carried me off on the stretcher.
F04 It also provides for funds to clear slums and help colleges build dormitories.It also provides for funds to clear slums and help colleges build dormitories.The caller says: It also provides for funds to clear slums and help colleges build dormitories.
F04 The dolphins swam around our boat.The dolphins swarmed around our boat. [wrong]The caller says: The dolphins swarmed around our boat. [wrong]
F04 The box contained three sweaters.The box contained 3 sweaters.The caller says: The box contained 3 sweaters.
F04 This is not a program of socialized medicine.This is not a program of socialized medicine.The caller says: This is not a program of socialized medicine.
F04 A roll of wire lay near the wall.A roll of wire lay near the wall.The caller says: A roll of wire lay near the wall.
F04 The pair of shoes was new.The pair of shoes was new.The caller says: The pair of shoes was new.
M01 but he always answers, "Banana oil!"Birdie, all ender, banana oil. [wrong]asked (rule 4)
M01 I looked up and noticed two old men.I looked up and nodded to Artman. [wrong]He looked up and nodded to Oddman. [wrong]
M01 I scrubbed the floors thoroughly.I... Scrub the floor. Yeah, really. [wrong]He is scrubbing the floor thoroughly.
M01 I just try to do my best.I just tried to do my best.The caller says: I just tried to do my best.
M01 My sister made the flowered curtains.My little baby. The flower curtain. [wrong]asked (rule 6)
M01 Aluminum silverware can often be flimsy.Aluminum everywhere. Can I be the newbie? Fear me. [wrong]asked (rule 4)
M01 Night after night, they received annoying phone calls.Night after night, they read in a 9-point font. [wrong]asked (rule 6)
M01 I expect we'll bounce back this week.I, I give back what you have bound back to me. [wrong]asked (rule 5)
M01 Before Thursday's exam, review every formula.Be more dirty again, baby. I breathe for me. [wrong]asked (rule 2)
M01 Alfalfa is healthy for you.I'm home. He's always here for you. [wrong]asked (rule 6)
M01 I have had my bell rung.I have held my breath. Borrow a room. [wrong]asked (rule 4)
M01 The books are very expensive.The books are very expensive.asked (rule 4)
M01 Students watched as he got out.Do you need an adult? Kind of. [wrong]asked (rule 4)
M01 Carl lives in a lively home.Kylie. In the library, oh. [wrong]asked (rule 5)
M01 Both injuries were to the same leg.More injuries were due to They made [wrong]asked (rule 4)
M02 he dresses himself in an ancient black frock coat,He dresses himself and then ancient black rock. told. [wrong]He dresses himself in an ancient black frock coat.
M02 Why yell or worry over silly items?Why yelp? I'm worried over silly items. [wrong]asked (rule 4)
M02 Are your grades higher or lower than Nancy's?Are your grades higher or lower than Nancy's?The caller says: Are your grades higher or lower than Nancy's?
M02 You'd be better off taking a cold shower.You'd be better off taking a cold shower.The caller says: You'd be better off taking a cold shower.
M02 We gathered shells on the beach.Forget the tears on the pillow. inch. [wrong]asked (rule 4)
M02 Their house is grey and white.Their house is gray and white. [wrong]asked (rule 6)
M02 She is thinner than I am.She is the nerden alien. [wrong]asked (rule 5)
M02 Just one side got wet.Just once I got wet. [wrong]He said he just once got wet. [wrong]
M02 You're used to being on the field.We're used to being on the field.The caller says: We're used to being on the field.
M02 The train approached the depot slowly.The train evolved, then devolved slowly. [wrong]asked (rule 6)
M02 Aluminum silverware can often be flimsy.I love my name. Can I be friendly? [wrong]asked (rule 2)
M02 Both injuries were to the same leg.Bob's egg was there. Cute. Same leg. [wrong]asked (rule 6)
M02 He further proposed grants of an unspecified sum for experimental hospitals.She fell upon Grant once. As a part of the experiment. High speed off. [wrong]asked (rule 6)
M02 If you are losing water, replace it immediately.If you are losing water, we play to the music. [wrong]asked (rule 6)
M02 The job provides many benefits.Good time. Provides many benefits. [wrong]He says the drive provides many benefits. [wrong]
M03 She had your dark suit in greasy wash water all year.She had your dark suit in greasy wash water all year.The caller says: She had your dark suit in greasy wash water all year.
M03 he dresses himself in an ancient black frock coat,He dresses himself in an ancient black frock coat.The caller says: He dresses himself in an ancient black frock coat.
M03 A long, flowing beard clings to his chin,A long flowing beard clings to his chin.asked (rule 2)
M03 I scrubbed the floors thoroughly.I scrubbed the floors thoroughly.He is saying he scrubbed the floors thoroughly.
M03 Where were you while we were away?Where were you while we were away?The caller says: Where were you while we were away?
M03 We gathered shells on the beach.We gathered shells on the beach.The caller says: We gathered shells on the beach.
M03 I feel I can play this weekend.I feel I can play this weekend.The caller says: I feel I can play this weekend.
M03 The little schoolhouse stood empty.The little schoolhouse stood empty.The caller says: The little schoolhouse stood empty.
M03 Before Thursday's exam, review every formula.Before Thursday's exam, review every formula.He is reviewing every formula before Thursday's exam.
M03 The box contained three sweaters.The box contained 3 sweaters.The caller says: The box contained 3 sweaters.
M03 He wrapped the package hastily.He wrapped the package hastily.The caller says: He wrapped the package hastily.
M03 The job provides many benefits.The job provides many benefits.The caller says: The job provides many benefits.
M03 Aluminum silverware can often be flimsy.Aluminum silverware can often be flimsy.The caller says: Aluminum silverware can often be flimsy.
M03 If you are losing water, replace it immediately.If you are losing water, replace it immediately.The caller says: If you are losing water, replace it immediately.
M03 We bought a brown chair.We bought a brown chair.The caller says: We bought a brown chair.
M04 Grandfather likes to be modern in his language.Good work. Like, dude. Be my ring. Mm-hmm. My good. [wrong]asked (rule 5)
M04 Except in the winter when the ooze or snow or ice prevents,It's that in the winter. When The opening. No. Oh, I, I didn't. [wrong]asked (rule 1)
M04 he slowly takes a short walk in the open air each day.He's probably dead. Uh. Good work. In the open air, in the day. [wrong]asked (rule 1)
M04 We have often urged him to walk more and smoke less,We, um, I didn't. I didn't do it, but come on. And mugglin'. [wrong]asked (rule 1)
M04 You wished to know all about my grandfather.You want to know all. About my, my, my grandpa.asked (rule 6)
M04 Well, he is nearly ninety-three years old;Why? He is really 93 years old. [wrong]asked (rule 6)
M04 She is thinner than I am.To hit wood thinner than I am. [wrong]asked (rule 4)
M04 I tried to tell people in the community.I try to tell people that the communityasked (rule 6)
M04 He will allow a rare lie.He will allow a lie.asked (rule 6)
M04 Why yell or worry over silly items?Why, yeah, we're in over for the item. [wrong]asked (rule 6)
M04 I looked up and noticed two old men.Hallelujah, Ab. And I did, I did do all men. [wrong]asked (rule 2)
M04 Nothing is as offensive as innocence.Nothing in the bed there. [wrong]asked (rule 4)
M04 Aluminum silverware can often be flimsy.Hello, Binyam. So we're aware. Glad to be with you. Sunday. [wrong]asked (rule 2)
M04 I expect we'll bounce back this week.I've been born I bet we're bound down and did we. [wrong]asked (rule 2)
M04 Nothing has been done yet to take advantage of the enabling legislation.Nothing. And then, then, yeah, you do take, take, and then, um, [wrong]asked (rule 4)
M05 he slowly takes a short walk in the open air each day.He slowly takes a walk. In the open air each day.He surely takes a walk in the open air each day. [wrong]
M05 yet he still thinks as swiftly as ever.Yet he still thinks I'd switch the ad offer. [wrong]asked (rule 4)
M05 Grandfather likes to be modern in his language.Grandfather Lake. Very modern in his language. [wrong]asked (rule 6)
M05 Where were you while we were away?Where were you while we were away?The caller says: Where were you while we were away?
M05 The islands are sparsely populated.The island. Uh. Wait. The populated. [wrong]asked (rule 5)
M05 The humidity is overwhelming there.The humidity is overwhelming. Ew.He says the humidity is overwhelming.
M05 The misguided souls have lost their way.The misguided soul. Have lost. No way. [wrong]asked (rule 1)
M05 Carl lives in a lively home.Clearly. In a library. Who? [wrong]asked (rule 5)
M05 Did dad do academic bidding?でだるいけん。 デミックビルディング。asked (rule 4)
M05 He really crucified him; he nailed it for a yard loss.He recrucified himself. He knew. It. Was a. You'd. [wrong]asked (rule 6)
M05 Beg that guard for one gallon of gas.Big red. Good. For 1 gallon of gas. [wrong]asked (rule 2)
M05 Although always again againUh, though, are we... Again, again. [wrong]asked (rule 2)
M05 Only lawyers love millionaires.Only those that have millionaires. [wrong]Only those that have millionaires. [wrong]
M05 He is definitely a notch above us.He is definitely not above us. [wrong]The caller says: He is definitely not above us. [wrong]
M05 the pair of shoes was newThe big issue. We knew. [wrong]asked (rule 5)
./eval/spike/scripts/fetch-torgo.sh      # TORGO, from the University of Toronto
node --env-file=.env eval/audiobench-stream.mjs --per-speaker 15
node --env-file=.env eval/audiobench-decide.mjs

The same sentences, over a telephone line

Everything above this section was measured at 16kHz, which is what a browser microphone gives. A phone line gives 8kHz mu-law: band-limited to about 3.4kHz and logarithmically companded, which discards exactly the high frequency detail that separates one consonant from another. Since the product is for phone calls, measuring it only on laptop audio would be measuring the easy case. So the same 120 TORGO sentences TORGO were streamed again through the same two ears, with the audio put through a real telephone encoder first and nothing else changed.

Wideband against telephone quality, same sentences, same engine
MeasuredAt 16kHzOver a phone line
Ordinary ear, word error rate42%46% (+4)
Patient ear, word error rate36%40% (+4)
Ordinary settings split the sentence48 of 12046 of 120
The patient ear split the sentence0 of 1200 of 120
An ordinary agent put words in their mouth65 of 12068 of 120
Patient ear, no checks, did57 of 12058 of 119
One Moment did11 of 5712 of 60
One Moment asked instead of speaking63 of 12060 of 120
  • Recognition gets worse, and the turn-taking result does not move. Both ears lose about four points of word error rate to the phone line, which is the honest cost of narrowband audio on speech that is already hard. But ordinary settings still cut 46 of 120 sentences off mid-way, and the patient ear still cuts 0. The central claim of this project is about when a turn ends, not about how well anything is heard, and that is why it survives the transcoding intact.
  • The caution got better aimed, not worse. This is the result we did not expect. On worse audio One Moment spoke about as often and was wrong about as often (12 of 60 against 11 of 57), but of the questions it asked, the ones that were not needed fell from 11 to 5, and the ones that were needed rose from 52 to 55. Degrading the audio degrades confidence and makes the two ears disagree more, and those are exactly the signals the rules key on. The caution is driven by real uncertainty rather than by a guess about it, so when the line gets worse it lands on the right sentences.
  • Why this run exists. Until it, the honest line on the limits page was that telephone audio was unmeasured, so nothing about phone calls was claimed. A product whose entire premise is a phone call cannot leave that unmeasured and still expect to be believed.
  • What is not repeated here. The post-call self-audit is not re-run on this arm, so the audit numbers above are wideband only. Run 2026-09-27, model qwen3.5-4b-32k-fast.
node --env-file=.env eval/audiobench-stream.mjs --narrowband --out eval/results/audiobench-evidence-8k.json
node --env-file=.env eval/audiobench-decide.mjs --arm 8k

And does it get in the way of everyone else?

Everything above is about disordered speech, where waiting longer is obviously the right thing to do. The fair objection is the opposite one: a relay that holds a turn for six seconds and asks whenever it is unsure might be unusable for a speaker with no difficulty at all. If that were true, the design would be wrong rather than narrow. So we ran 60 sentences of ordinary fluent speech from 35 readers, taken from LibriSpeech LibriSpeech, through the same two ears, the same end-of-turn rules and the same engine.

Fluent speech, through the patient system
MeasuredResult
It spoke rather than asked49 of 60 (82%), of which 39 took the verbatim path and used no model at all
It put words in a fluent speaker's mouth7 of 49
How long it waited after they stopped1.7s median, 3.7s mean. 34 of 60 ended in under two seconds
The patient ear split the sentence0 of 60
Ordinary settings split the sentence8 of 60
Word error rate, patient ear against the read sentence5%
  • Patience is not the same as lag. The six-second window is a ceiling, not a wait. When both ears agree a sentence has finished, ForceEndpoint ends the turn about a second and a half later, which is why the median here is 1.7 seconds and 34 of 60 turns ended in under two. A fluent speaker pays for the patience only when they actually pause.
  • The 7 flagged relays are worth reading, because none of them is an invention. Every word the scorer objected to was in what the ears actually heard, so the engine added nothing. Most are the corpus and the recogniser disagreeing about spelling rather than mistakes a listener would notice: LibriSpeech writes “MARCH TWENTY SECOND EIGHTEEN THIRTY SEVEN” where AssemblyAI heard “March 22nd, 1837”, and “ENQUIRIES” where it heard “inquiries”. The genuine ones are misheard names, such as “Emil” heard as “Amiel”, which the verbatim path passes on faithfully. We do not improve recognition and have never claimed to; this is what that costs on clean speech.
  • The cost of caution, on people who do not need it. 11 of 60 became a question rather than a relay. On clean speech that is the system being careful about nothing, and it is the honest price of a design that would rather ask than guess.
  • Even here, ordinary settings cut people off. The fast ear split 8 of 60 fluent, unimpaired sentences into more than one turn. The patient ear split 0.
  • Corpus: LibriSpeech test-clean, read speech from public-domain audiobooks, CC BY 4.0, with the sentence each speaker read as the reference. We did not record it and did not choose which sentences exist in it. Run 2026-09-29, 42 model calls across 60 sentences, because 39 of them needed no model at all.
node --env-file=.env eval/fluentbench.mjs --n 60
node --env-file=.env eval/fluentbench-decide.mjs

What the turn setting costs the person speaking

One setting decides whether a sentence survives: min_turn_silence. It can be characterised by sweeping it against synthetic speech, and that tells you when a turn fires. It cannot tell you what firing costs, because a synthetic voice never stops in the middle of a word to look for one. So we swept it across 12 sentences from 7 speakers with dysarthria TORGO and scored each value by what the speaker loses.

One ear, one setting at a time, on real disordered speech
min_turn_silenceTurn closed while they were still speakingSentences cut into more than one turnWords still to come when the turn closedMedian wait after they stopped
vendor default4 of 122 of 1210 of 860.2s
1000ms1 of 121 of 124 of 860.9s
2400ms0 of 120 of 122 of 862.4s
6000ms (ours)0 of 120 of 122 of 866.1s
  • At the vendor default, the turn closes while the person is still talking. It did so on 4 of 12 sentences, with 10 words of 86 still to come. A voice agent on those settings is not waiting for a pause; it is answering before the sentence exists.
  • And here is the result that argues against our own setting. At 2400ms nothing fired early, nothing split, and only 2 words were lost: the same as ours, for 3.7 seconds less waiting. On this corpus, 6000ms buys nothing that 2400ms has not already bought.
  • Why we kept 6000 anyway, stated so you can disagree. TORGO is read speech: the pauses are short, so a 2400ms setting is enough to clear them. The pause this product exists for is a word-finding block in aphasia, which the literature puts at roughly 5 to 10 seconds, and no corpus of dysarthric read speech can show it. The 6000ms value is a ceiling for that case, not a delay we pay on every turn: when both ears agree a sentence has finished, ForceEndpoint closes it early, which is why the median wait on fluent speech measured 1.7 seconds rather than six.
  • What this measures, and what it does not. One ear, no ForceEndpoint and no second stream, so each number is about the setting rather than about our orchestration on top of it. Twelve sentences is a small sample and is reported as one. Run 2026-09-29.
node --env-file=.env eval/turnbench.mjs --n 12
node eval/turnbench-summarise.mjs

Your words, double-checked

The patient ear is told the caller's own words, so it hears them more easily. That is also the danger: a word it is listening for can be heard because it was expected. We tested it with four slurred ways of saying a medicine.

Vocabulary probe
SaidEar told the caller's wordsEar not told
am low dippyamlodipineAMLO Dippy
amla deepinamlodipineAmla Deepin
amblo, dipineamlodipineamblyodipine
um, am, am loum, amlodipineum, am. Am. Am low

The boosted ear produced “amlodipine” 4 times out of 4; the unbiased ear, 0. So rule 11: a word from the caller's list that only the listening-for-it ear heard, or that only partly came out, is offered as a choice (“Amlodipine, or metformin?”) and never relayed on trust. Measured 18 September 2026 with eval/probe-vocabulary.mjs.

A failure the live checks missed, and the audit caught

We publish this one on purpose. In an earlier run of the harder call, the caller said “I need to refill my ... am low dippy prescription, please.” In context, both live ears wrote “amlodipine”, so they agreed, no rule fired, and it was relayed.

The failure
SourceWhat it had
Patient ear (live)I need to refill my amlodipine prescription, please.
Fast ear (live)I need to refill my... Amlodipine prescription, please.
Relayed to the pharmacistRobert says: I need to refill my amlodipine prescription, please.
Careful transcript after the call (universal-3-5-pro)I need to refill my... I'm low. Dippy prescription, please.
Self-audit verdictSome relayed words were not in the careful transcript: amlodipine.

When both ears are wrong the same way, no live check can know. The self-audit can, because the pre-recorded model listens more carefully and was not told what to expect. This is why every call is graded afterwards, and why the product is positioned as a first line with a human behind it.

Check it yourself: 12 of 12 checks pass

Every figure above is derived from a run file in the repository rather than typed into the page. So the question is not whether we made a typo, it is whether those files still say what we claim. npm run verify scores every relay in both benchmarks again from scratch, against ground truth we did not write, rebuilds the totals from those fresh scores, and exits non-zero if anything differs. This table is that command's output, not a summary of it.

Verification checks
CheckResultWhat it found
NEGBENCH rows score the same todaypass48 relays re-scored against the Harvard Sentences, all identical
NEGBENCH published totals match its rowspass3 conditions rebuilt from 48 rows
AUDIOBENCH rows score the same todaypass120 sentences re-scored against the TORGO prompts, all identical
AUDIOBENCH published totals match its rowspass120 sentences, 8 speakers, every arm rebuilt from the rows
AUDIOBENCH still reports its own failurespass11 wrong relays, 11 questions that were not needed, self-audit caught 3 of 11
FLUENTBENCH rows score the same todaypass60 fluent sentences re-scored against the LibriSpeech references
FLUENTBENCH published totals match its rowspass60 sentences from 35 readers, every total rebuilt
FLUENTBENCH still reports what the caution costspass11 of 60 became a question rather than a relay
recorded-call.json: every word the agent spoke was approvedpass2 replies, 24 approved words, 0 unapproved
recorded-call-choice.json: every word the agent spoke was approvedpass2 replies, 24 approved words, 0 unapproved
recorded-call-dissent.json: every word the agent spoke was approvedpass2 replies, 25 approved words, 0 unapproved
the engine tests passpass65 of 65 passing, 0 failing

Run 2026-09-29, with 65 of 65 engine tests passing. One of the checks asserts that the unflattering numbers are still there, so a future run cannot quietly drop them.

Reproduce it

git clone <this repository>
cd one-moment && npm install
npm run verify                   # re-derives every number above from the committed runs
cp .env.example .env            # add your AssemblyAI key
npm test                         # the engine: floor, evidence, Dissent, semantic patience
node --env-file=.env apps/orchestrator/scripts/record-sample-call.ts   # a full call, recorded
node --env-file=.env eval/negbench.mjs --n 16 --seed 11                # the benchmark

Where the waiting time came from, and what the literature says about it

We have not tested this with a person who has aphasia. That is the honest limit of the whole project, and no amount of benchmarking substitutes for it. What we can do is say where each design decision came from, and then check it against research done by people who did work with them.

Design decisions against published research
The decisionWhy we chose itWhat the literature says
Wait 6 to 9 secondsMeasured, not reasoned: on one recording, default settings split a 6.0-second word-finding pause into two turns, so the patient ear was set to hold past it.Optimal response-time cutoffs for people with aphasia cluster at approximately 5 to 10 seconds, measured across 10 people with aphasia, in an assessment context where they are typically allowed up to 30 seconds. Evans et al. 2020 The window we picked by measurement sits inside the window the literature identifies. We did not know that when we picked it.
Target the phone, not the roomA trained partner is the best-evidenced help in conversation, and the person on the other end of a phone call cannot be trained.The phone is a documented barrier in its own right: it strips the gesture and facial cues people with aphasia rely on, and phone difficulty is associated with greater social isolation. Greig et al. 2008 Assessment by telephone has been reported as anxiety-provoking in a way that itself degrades spoken output.
Ask, rather than guessA relay that invents one word is worse than no relay, because the caller cannot hear what was said in their name.In the one study we found of voice assistants and aphasia, 8 participants, recognition failures caused frustration but did not destroy acceptance where the device met a real need, and a participant valued getting answers “without being perceived as dumb”. The authors argue a voice assistant is less stigmatising than other assistive technology because it is used in spite of the impairment rather than because of it. Nunez Macias et al. 2023

What this does not establish. None of these studies tested this system, or any automated conversation partner. They establish that the problem is real, that the waiting window we chose is the right order of magnitude, and that people with aphasia will use a voice assistant that earns it. They do not establish that this one helps anybody. Only testing with people who have aphasia would do that, and that is the next step rather than a claim.

The research it rests on

The central inference, stated with its limit: the best-evidenced help for aphasia in conversation is a trained partner, and the partner on a phone call cannot be trained. No study has tested an automated partner. This is an untested application of a well-evidenced principle.

Aphasia and partner training

Speech recognition and AI

Our own accessibility audit is on the accessibility page.