Voice recognition and auto-prompting: how speech-following teleprompters work
A speech-following prompter matches what it hears to where it is in the script and scrolls to match. This is how that works, how it stays with a speaker who ad-libs, and what the operator sees.
| surface | feed | state |
|---|---|---|
| Stage confidence monitor | presenter view | live |
| Operator laptop | mirror | live |
| Stream Deck | webhook | armed |
| Comfort monitor — FOH | mirror | live |
How voice following works
A speech-following prompter runs speech recognition on the presenter's microphone, aligns the recognised words against the script it already has, and scrolls so the current line sits at the reading position. It is not transcribing for its own sake — it has the text already. It is solving a much narrower problem: given roughly these words, where in this known script are we?
That narrowness is what makes it work at all. General dictation has to consider every word in a language. A prompter only has to consider the next few hundred words of one document, which means it can tolerate a poor recognition result and still land in the right place. A word misheard entirely still leaves enough matching neighbours to locate the line.
The moment that separates prompters is what happens when the speaker leaves the script. There is no longer a position in the document that matches what is being heard, and the system has to recognise that and hold — not guess, and not scroll away from the speaker.
Where it breaks
The hard cases are the same for every implementation, because they are properties of live speech rather than of any one product. What separates products is how they behave in them.
Ad-libbing
The speaker goes off script to thank someone, answer a heckle or react to the room. EventBlok recognises the departure, holds its place calmly, and rejoins the moment the script resumes — it does not guess, and it does not scroll away from the speaker.
Repeated phrases
A script that says "thank you very much" three times gives three equally good matches. Position, not just wording, has to break the tie.
Room noise and applause
Applause is the worst case, because it arrives exactly at the moments the script is most likely to move on.
Accent and pace
Recognition quality varies by voice. A prompter that follows one presenter perfectly may lag noticeably on the next.
Two people talking
A panel or a co-host handover means two voices, and possibly two microphones, against one script.
What the operator needs
Voice following does the work. The operator's interface is not the presenter's view — it is a mirror of where the system is in the script, how confident it is, and the speaker's pace against it, with manual override the whole time: skip back or forward to catch up, or take over on smooth mouse control.
The question is never whether it will lose the speaker. It is how quickly a human can put it back.
Manual override has to be immediate and has to be obvious, because it will be used under pressure in a dark room. An operator who has to leave an automatic mode before they can scroll has already lost the moment they were trying to fix.
When not to use it
Speech following suits scripted delivery: an awards script, a keynote written to be read, a broadcast link, a formal welcome. It suits unscripted delivery badly, and a panel discussion not at all.
For bullet-point speakers the honest answer is usually not to follow the voice but to give them a static set of prompts and let the operator advance them. A prompter chasing a speaker who is not reading anything is a distraction on a confidence monitor, and the speaker will notice it moving.
How EventBlok fits
Everything above applies to any speech-following prompter. This is what EventBlok's does, so you can judge whether it is relevant.
The EventBlok prompter follows the speaker's voice and gives the operator a mirror showing position, pace and drift, with manual control available at any time. The part that is specific to it is where the script comes from: it reads the same running order as the showflow, so a session that moves in the programme moves for the prompter too. The script is a view of the event record rather than a file loaded into a separate application before doors.
Recognition runs on-device using Moonshine, with Deepgram available as a cloud fallback. That split is worth understanding rather than taking on trust, because it decides how the prompter behaves on a bad day: on-device recognition does not depend on a working internet connection and does not send the presenter's audio anywhere, which matters in a basement, in a marquee, and in any room where the content is confidential. When the cloud fallback is in use, audio is being sent to a third-party service — so if that is not acceptable for a particular event, it is a question to settle before doors rather than during them.
For how the prompter sits inside the running order rather than beside it, see teleprompter and running order.

A presenter working to a prompter that is following their voice
Nicolae Ciobota
Founder, EventBlok. Live, hybrid and broadcast producer, showcaller and technical director, in production since 1996.
Common questions
It runs speech recognition on the presenter's microphone and matches the recognised words against the script it already holds, then scrolls so the current line sits at the reading position. Because it only has to search one known document rather than a whole language, it can tolerate poor recognition and still find the right place.
EventBlok recognises there is no longer a matching position, holds its place, and picks the speaker straight back up when they return to the script. The operator keeps manual override throughout — skip back or forward to catch up, or take over on smooth mouse control — because on a live show a human must always be able to take the wheel.
No. Speech following assumes one voice reading one script. A panel is several voices and no script, so there is nothing to follow. Panels are better served by static prompts or by nothing at all.
Not by default. Recognition runs on the device using Moonshine, so it works without an internet connection and the presenter's audio stays on the machine. Deepgram is available as a cloud fallback, and when that is in use audio is sent to a third-party service — which is a decision worth making before the show if the content is confidential.
Yes. Applause is the hardest case, because it tends to arrive at exactly the points where the script moves on. Microphone choice and placement matter as much as the software — a prompter fed from a good speech mic follows far better than one fed from an ambient room mic.
MORE FROM THE JOURNAL
- GUIDE
What is a showflow? A practical guide to the live-event running order
A showflow is the cue-level running order a production crew calls a live event from. This is what belongs in o…
- GUIDE
How to write a run of show for a live event
A run of show is the minute-by-minute document the production team delivers from. Here is the structure that w…
- GUIDE
Event agenda vs showflow vs run sheet: what each document is actually for
These three documents describe the same event to three different audiences, and confusing them is why events e…