Voice recognition and auto-prompting: how speech-following teleprompters work
A speech-following prompter matches what it hears to where it is in the script and scrolls to match. This is how that works, where it fails, and why the operator still matters.
| surface | feed | state |
|---|---|---|
| Stage confidence monitor | presenter view | live |
| Operator laptop | mirror | live |
| Stream Deck | webhook | armed |
| Comfort monitor — FOH | mirror | live |
How voice following works
A speech-following prompter runs speech recognition on the presenter's microphone, aligns the recognised words against the script it already has, and scrolls so the current line sits at the reading position. It is not transcribing for its own sake — it has the text already. It is solving a much narrower problem: given roughly these words, where in this known script are we?
That narrowness is what makes it work at all. General dictation has to consider every word in a language. A prompter only has to consider the next few hundred words of one document, which means it can tolerate a poor recognition result and still land in the right place. A word misheard entirely still leaves enough matching neighbours to locate the line.
The practical consequence is that a prompter which follows well on a scripted read can still fail badly the moment the speaker leaves the script — because there is no longer a position in the document that matches what it is hearing.
Where it breaks
The failures are consistent across every implementation, because they are properties of the technique rather than of any one product.
- Ad-libbing
- The speaker goes off script to thank someone, answer a heckle or react to the room. There is nothing to match, so the safe behaviour is to hold position and wait rather than to guess.
- Repeated phrases
- A script that says "thank you very much" three times gives three equally good matches. Position, not just wording, has to break the tie.
- Room noise and applause
- Applause is the worst case, because it arrives exactly at the moments the script is most likely to move on.
- Accent and pace
- Recognition quality varies by voice. A prompter that follows one presenter perfectly may lag noticeably on the next.
- Two people talking
- A panel or a co-host handover means two voices, and possibly two microphones, against one script.
What the operator needs
Voice following handles the ordinary case. The operator exists for the rest, and the interface they need is not the presenter's view — it is a view of where the system thinks the speaker is, how confident it is, and how far that has drifted from where they actually are.
Manual override has to be immediate and has to be obvious, because it will be used under pressure in a dark room. An operator who has to leave an automatic mode before they can scroll has already lost the moment they were trying to fix.
When not to use it
Speech following suits scripted delivery: an awards script, a keynote written to be read, a broadcast link, a formal welcome. It suits unscripted delivery badly, and a panel discussion not at all.
For bullet-point speakers the honest answer is usually not to follow the voice but to give them a static set of prompts and let the operator advance them. A prompter chasing a speaker who is not reading anything is a distraction on a confidence monitor, and the speaker will notice it moving.
How EventBlok fits
Everything above applies to any speech-following prompter. This is what EventBlok's does, so you can judge whether it is relevant.
The EventBlok prompter follows the speaker's voice and gives the operator a mirror showing position, pace and drift, with manual control available at any time. The part that is specific to it is where the script comes from: it reads the same running order as the showflow, so a session that moves in the programme moves for the prompter too. The script is a view of the event record rather than a file loaded into a separate application before doors.
Recognition runs on-device using Moonshine, with Deepgram available as a cloud fallback. That split is worth understanding rather than taking on trust, because it decides how the prompter behaves on a bad day: on-device recognition does not depend on a working internet connection and does not send the presenter's audio anywhere, which matters in a basement, in a marquee, and in any room where the content is confidential. When the cloud fallback is in use, audio is being sent to a third-party service — so if that is not acceptable for a particular event, it is a question to settle before doors rather than during them.
For how the prompter sits inside the running order rather than beside it, see teleprompter and running order.

Common questions
How does a voice-following teleprompter work?
It runs speech recognition on the presenter's microphone and matches the recognised words against the script it already holds, then scrolls so the current line sits at the reading position. Because it only has to search one known document rather than a whole language, it can tolerate poor recognition and still find the right place.
What happens when the speaker goes off script?
There is no longer any position in the script that matches what is being heard, so the correct behaviour is to hold position and wait rather than guess. This is exactly why an operator with immediate manual control still matters — the system will lose an ad-libbing speaker, and a human has to put it back.
Is voice following suitable for a panel discussion?
No. Speech following assumes one voice reading one script. A panel is several voices and no script, so there is nothing to follow. Panels are better served by static prompts or by nothing at all.
Does the EventBlok prompter send audio to the cloud?
Not by default. Recognition runs on the device using Moonshine, so it works without an internet connection and the presenter's audio stays on the machine. Deepgram is available as a cloud fallback, and when that is in use audio is sent to a third-party service — which is a decision worth making before the show if the content is confidential.
Does room noise or applause affect it?
Yes. Applause is the hardest case, because it tends to arrive at exactly the points where the script moves on. Microphone choice and placement matter as much as the software — a prompter fed from a good speech mic follows far better than one fed from an ambient room mic.
Plan the event. Publish the programme. Run the show.
One event spine, many views. Build the programme once, then give every person the right live view of it.