How scoring works
No black box. Every number, explained.
Each take is measured on nine things a listener can hear. Here is what each one means, how we measure it, and where the measurement can be wrong.
A real attempt, or no score
Before anything is scored, the engine checks that it heard a person speaking: short voiced syllables switching on and off several times a second, with moving pitch. Noise, music, humming, counting or a word list get no score.
Measured from your voice, not guessed
Pace, pauses, hesitations, pitch and your start time come from the audio itself, 100 measurements a second. We tested every measure against recordings with known answers before showing it to you.
Compared with you, not with folklore
Pace bands, hesitation trends and goals are set from your own takes. There is no zero-um target and no "authority" pitch rule, because the research does not support them.
One fix at a time
People improve fastest working on one thing. Every take ends with one fix and one number to beat, and your next take checks whether it worked.
Your voice stays on your device
Analysis runs in your browser. Audio is not uploaded or stored on our servers.
The nine measurements
Pace
Needs your voiceHow many words per minute you speak across your answer, pauses included.
Good looks like: Anywhere in your band scores full marks: 120 to 195 words per minute to start, then a band centred on your own natural pace once we know it.
Why it matters and how we measure it
- Why it matters
- Listeners follow a wide range of speeds easily. In lab studies faster, fluent speech was rated as more credible, not less. The real risk is slow, halting delivery.
- How we measure it
- Counted from the audio itself: we find every syllable in your voice and convert to words, or use your transcript when it is complete. Tested at about 11% error in quiet rooms.
- Where it can be wrong
- Very echoey rooms blur syllables and can read a little slow. Your own trend matters more than any single number.
Hesitations
Needs your voiceUm, uh, drawn-out sounds and filler words like "you know", per 100 words.
Good looks like: 3 or fewer per 100 words scores full marks. 12 or more scores zero.
Why it matters and how we measure it
- Why it matters
- Some are normal: everyday speech has about 6 per 100 words, and research finds little effect on how competent people seem. Fewer, shorter hesitations are easier to follow, so we track yours against your own baseline, never against zero.
- How we measure it
- Um and uh are heard in the audio, because most transcription engines quietly delete them. Filler words come from your transcript. Tap any hesitation in your replay to hear it.
- Where it can be wrong
- The audio detector is tuned to be careful: it misses some hesitations rather than inventing them, and echo makes it more cautious still.
Confident language
Needs your wordsHow often you soften claims with phrases like I think, kind of, maybe, sorry.
Good looks like: 1 or fewer per 100 words scores full marks.
Why it matters and how we measure it
- Why it matters
- When listeners pay attention, hedged claims persuade less. Hedging real uncertainty is honest; hedging your own recommendation is not.
- How we measure it
- We search your transcript for softening phrases, counted per 100 words.
- Where it can be wrong
- Needs a transcript. Some hedges are right: use judgment.
Flow
Needs your voiceWhether your speech flows or stalls in long gaps of two seconds or more.
Good looks like: Half a long stall per minute or fewer scores full marks.
Why it matters and how we measure it
- Why it matters
- Short pauses give you time to think and are fine. Long stalls read as lost place. Pause counts themselves are not scored: there is no good evidence for a magic number.
- How we measure it
- From the audio: every silent gap inside your answer, measured to the hundredth of a second. Gaps of 2 seconds or more count as stalls.
- Where it can be wrong
- Loud background noise can hide very short pauses, which is why only long gaps are scored.
Vocal energy
Needs your voiceHow much your voice moves in pitch and loudness from syllable to syllable.
Good looks like: About 7 semitones of pitch range, with clear loudness contrast, scores near full marks. Under 3 reads as flat.
Why it matters and how we measure it
- Why it matters
- In a study of startup pitches, how passionate founders seemed predicted funding. Pitch range and loudness contrast are our audio proxy for part of that energy.
- How we measure it
- We track the pitch of your voice 100 times a second, correct octave errors, and measure the range you actually use (10th to 90th percentile), plus the loudness of each syllable. Tested within 0.2 semitones of a lab reference.
- Where it can be wrong
- Your natural pitch is never graded: only how much you move from your own centre.
Loudness contrast
Needs your voiceHow much louder your stressed syllables are than the rest: the contrast that makes key words stand out.
Good looks like: About 5 dB of contrast between syllable peaks scores full marks.
Why it matters and how we measure it
- Why it matters
- Contrast creates emphasis, but the direct evidence is thin, so this is a supporting cue, not a headline score.
- How we measure it
- The loudness at the centre of every syllable we detect, and how much it varies across the take.
- Where it can be wrong
- Moving closer to or farther from the microphone changes this number.
Landing statements
Needs your voiceWhether your statements end with a settled, falling tone.
Good looks like: Most phrase endings falling by a semitone or more scores full marks.
Why it matters and how we measure it
- Why it matters
- Falling at the end of a statement may read as more confident. Questions and lists are meant to rise, so this is a light cue, never a verdict.
- How we measure it
- For every phrase you speak, we compare the pitch of its last quarter-second with the rest of the phrase.
- Where it can be wrong
- We cannot tell a question from a statement by sound alone, so this carries little weight.
Pace gears
Needs your wordsA retired measure of how much you changed speed.
Good looks like: Not scored.
Why it matters and how we measure it
- Why it matters
- We found no evidence behind it, so it is no longer scored.
- How we measure it
- Not measured.
- Where it can be wrong
- Retired.
Timing
Needs your voiceWhether you finished close to the target length for the drill.
Good looks like: Within 15% of the target length scores full marks.
Why it matters and how we measure it
- Why it matters
- Real meetings have a clock. Landing in the time you are given is a practical skill.
- How we measure it
- Total take length compared with the drill's target.
- Where it can be wrong
- None. This is a stopwatch.
Speaking structures
In Situations we also check whether a well-known structure is audible in your words. This is separate from your score, because structure and delivery are different skills.
PREP
Point, Reason, Example, Point. The simplest structure for giving an opinion or answer.
3-2-1
Three points, two pieces of evidence, one takeaway. Great for introductions and updates.
BRIEF
Background, Reason, Information, End with the ask. For updates, requests and pitches.
STAR
Situation, Task, Action, Result. The answer shape interviewers score behavioural questions against.
SBI
Situation, Behavior, Impact. The clearest way to give feedback without blame.
Label and ask
Name the emotion (“It sounds like…”), then ask an open “what” or “how” question.
Privacy settings · This is practice guidance, not certification or clinical assessment.