Skip to content

How scoring works

No black box. Every number, explained.

Each take is measured on nine things a listener can hear. Here is what each one means, how we measure it, and where the measurement can be wrong.

A real attempt, or no score

Before anything is scored, the engine checks that it heard a person speaking: short voiced syllables switching on and off several times a second, with moving pitch. Noise, music, humming, counting or a word list get no score.

Measured from your voice, not guessed

Pace, pauses, hesitations, pitch and your start time come from the audio itself, 100 measurements a second. We tested every measure against recordings with known answers before showing it to you.

Compared with you, not with folklore

Pace bands, hesitation trends and goals are set from your own takes. There is no zero-um target and no "authority" pitch rule, because the research does not support them.

One fix at a time

People improve fastest working on one thing. Every take ends with one fix and one number to beat, and your next take checks whether it worked.

Your voice stays on your device

Analysis runs in your browser. Audio is not uploaded or stored on our servers.

The nine measurements

Pace

Needs your voice

How many words per minute you speak across your answer, pauses included.

Good looks like: Anywhere in your band scores full marks: 120 to 195 words per minute to start, then a band centred on your own natural pace once we know it.

Why it matters and how we measure it
Why it matters
Listeners follow a wide range of speeds easily. In lab studies faster, fluent speech was rated as more credible, not less. The real risk is slow, halting delivery.
How we measure it
Counted from the audio itself: we find every syllable in your voice and convert to words, or use your transcript when it is complete. Tested at about 11% error in quiet rooms.
Where it can be wrong
Very echoey rooms blur syllables and can read a little slow. Your own trend matters more than any single number.

Hesitations

Needs your voice

Um, uh, drawn-out sounds and filler words like "you know", per 100 words.

Good looks like: 3 or fewer per 100 words scores full marks. 12 or more scores zero.

Why it matters and how we measure it
Why it matters
Some are normal: everyday speech has about 6 per 100 words, and research finds little effect on how competent people seem. Fewer, shorter hesitations are easier to follow, so we track yours against your own baseline, never against zero.
How we measure it
Um and uh are heard in the audio, because most transcription engines quietly delete them. Filler words come from your transcript. Tap any hesitation in your replay to hear it.
Where it can be wrong
The audio detector is tuned to be careful: it misses some hesitations rather than inventing them, and echo makes it more cautious still.

Confident language

Needs your words

How often you soften claims with phrases like I think, kind of, maybe, sorry.

Good looks like: 1 or fewer per 100 words scores full marks.

Why it matters and how we measure it
Why it matters
When listeners pay attention, hedged claims persuade less. Hedging real uncertainty is honest; hedging your own recommendation is not.
How we measure it
We search your transcript for softening phrases, counted per 100 words.
Where it can be wrong
Needs a transcript. Some hedges are right: use judgment.

Flow

Needs your voice

Whether your speech flows or stalls in long gaps of two seconds or more.

Good looks like: Half a long stall per minute or fewer scores full marks.

Why it matters and how we measure it
Why it matters
Short pauses give you time to think and are fine. Long stalls read as lost place. Pause counts themselves are not scored: there is no good evidence for a magic number.
How we measure it
From the audio: every silent gap inside your answer, measured to the hundredth of a second. Gaps of 2 seconds or more count as stalls.
Where it can be wrong
Loud background noise can hide very short pauses, which is why only long gaps are scored.

Vocal energy

Needs your voice

How much your voice moves in pitch and loudness from syllable to syllable.

Good looks like: About 7 semitones of pitch range, with clear loudness contrast, scores near full marks. Under 3 reads as flat.

Why it matters and how we measure it
Why it matters
In a study of startup pitches, how passionate founders seemed predicted funding. Pitch range and loudness contrast are our audio proxy for part of that energy.
How we measure it
We track the pitch of your voice 100 times a second, correct octave errors, and measure the range you actually use (10th to 90th percentile), plus the loudness of each syllable. Tested within 0.2 semitones of a lab reference.
Where it can be wrong
Your natural pitch is never graded: only how much you move from your own centre.

Loudness contrast

Needs your voice

How much louder your stressed syllables are than the rest: the contrast that makes key words stand out.

Good looks like: About 5 dB of contrast between syllable peaks scores full marks.

Why it matters and how we measure it
Why it matters
Contrast creates emphasis, but the direct evidence is thin, so this is a supporting cue, not a headline score.
How we measure it
The loudness at the centre of every syllable we detect, and how much it varies across the take.
Where it can be wrong
Moving closer to or farther from the microphone changes this number.

Landing statements

Needs your voice

Whether your statements end with a settled, falling tone.

Good looks like: Most phrase endings falling by a semitone or more scores full marks.

Why it matters and how we measure it
Why it matters
Falling at the end of a statement may read as more confident. Questions and lists are meant to rise, so this is a light cue, never a verdict.
How we measure it
For every phrase you speak, we compare the pitch of its last quarter-second with the rest of the phrase.
Where it can be wrong
We cannot tell a question from a statement by sound alone, so this carries little weight.

Pace gears

Needs your words

A retired measure of how much you changed speed.

Good looks like: Not scored.

Why it matters and how we measure it
Why it matters
We found no evidence behind it, so it is no longer scored.
How we measure it
Not measured.
Where it can be wrong
Retired.

Timing

Needs your voice

Whether you finished close to the target length for the drill.

Good looks like: Within 15% of the target length scores full marks.

Why it matters and how we measure it
Why it matters
Real meetings have a clock. Landing in the time you are given is a practical skill.
How we measure it
Total take length compared with the drill's target.
Where it can be wrong
None. This is a stopwatch.

Speaking structures

In Situations we also check whether a well-known structure is audible in your words. This is separate from your score, because structure and delivery are different skills.

  • PREP

    Point, Reason, Example, Point. The simplest structure for giving an opinion or answer.

  • 3-2-1

    Three points, two pieces of evidence, one takeaway. Great for introductions and updates.

  • BRIEF

    Background, Reason, Information, End with the ask. For updates, requests and pitches.

  • STAR

    Situation, Task, Action, Result. The answer shape interviewers score behavioural questions against.

  • SBI

    Situation, Behavior, Impact. The clearest way to give feedback without blame.

  • Label and ask

    Name the emotion (“It sounds like…”), then ask an open “what” or “how” question.

Privacy settings · This is practice guidance, not certification or clinical assessment.