Voice following

A voice activated teleprompter that follows the words, not the clock.

Most prompters scroll at a fixed speed and the speaker chases the text. Wispim listens to what you actually said and moves the script to match — pause, hurry or go back a line, and it comes with you.

Coming to the App Store 6 free takes. iPhone, iOS 16 or later.

What “follows your voice” actually means

Voice mode recognises your speech and scrolls to the words you actually said. Pause to breathe or let a car go past, and the script waits with you. Speed up and it keeps up. Say an earlier line again and it scrolls back to that line.

That last one gives away that the tracking is positional, not a clever speed control: the app matches what it heard against where you are in the text, so backwards is as ordinary as forwards and there is no drift to correct. There is nothing to count down to, either — Voice and Sound wait for your first words.

Why a voice activated teleprompter beats a fixed speed

A fixed-speed prompter asks you to do two jobs at once: deliver the line, and manage your pace against a machine that has no idea what you just said. The tell on camera is rarely that someone is reading — it is that they are reading slightly too fast or slightly too slow, eyes flicking sideways to find the place. Following the words removes the second job.

Sound mode, and exactly when to pick it

Sound mode advances on the presence of speech rather than on the words: it moves while it hears you and stops when you stop, without matching what it heard against the script. Two situations call for it. The first is a language your iPhone cannot recognise — Sound works in any language, because it is not listening for one. The second is a condition where recognition would struggle: a noisy room, a distant microphone, a script full of names and numbers. The trade is the obvious one — Sound does not know you repeated a line, so it cannot scroll back on its own.

Constant, and the mode you did not choose

Constant is the conventional teleprompter: a fixed speed in words per minute, kept because sometimes it is the right tool. It is the only mode that does not need the microphone, and the only one with a start delay — off, 5, 10 or 20 seconds, time to get into position, shown inside the floating window as well as full screen.

It is also the safety net. Voice following keeps up while you record in Camera, Instagram and TikTok — the three that are tested — and if some other app does take the microphone mid-take, Wispim falls back to Constant by itself rather than freezing mid-sentence. The floating window page has that hand-off in both directions; all three modes sit side by side on the home page.

Where the recognition runs

Recognition runs on your iPhone wherever your language supports it: Wispim asks iOS for the on-device recogniser whenever one exists. Where your language has no on-device model, iOS sends the audio to Apple to convert it to text — Apple’s pipeline, not ours, and the audio does not reach us either way.

You are told twice, in your own language: the language picker marks those languages before a take, and a banner says so during one that ends up on the server. You will not read “offline teleprompter” here, because for some languages that would be untrue. Your scripts, the audio and the transcripts are never sent to us; the privacy policy has the rest.

Three language counts, and they are different numbers

The three language counts
WhatHow many
Voice following Every language your iPhone can recognise — 60+ locales on current iOS, and the list moves with iOS.
The app’s interface 16 languages, every screen. All sixteen have speech recognition.
Spoken commands 10 languages, phrases editable. The other six interface languages have none, and the app says so.

When it mishears you

It will, occasionally — a homophone, a name, a sentence lost to a passing truck. Recovery is what makes a following prompter usable, and none of it asks you to stop recording and start the take again:

  • Tap a word. Reading carries on from the word you tapped. Every mode, and it never pauses.
  • Browse mode. Drag the script to read ahead or check a line you passed. Following pauses, an arrow points back to where you left, and you can start reading anywhere.
  • Spoken commands. Pause, continue and jump between sentences by saying so — ten languages, phrases editable.
  • Volume buttons. One sentence back or forward. Marked Experimental in the app, and they do not work inside camera apps, which take the buttons.

Tapping and dragging are full-screen tools: the floating window does not accept touch at all, an iOS constraint and the reason the spoken commands exist. Choose your mode before a take that will live in that window — over Camera, or over a Zoom or Meet call.

Questions

How does a teleprompter know where you are in the script?

Wispim asks iOS to recognise what you are saying and matches it against the script, then scrolls to the words you actually said. Because it tracks your position in the text rather than counting seconds, pausing holds the script still, hurrying moves it faster, and saying an earlier line again scrolls back to that line.

What happens if I improvise or skip a line?

The script stays where you left it and picks you up again when you come back to written words. If you want to move it yourself, tap any word and reading carries on from there without pausing, or drag the script in browse mode — an arrow points back to the line you left, and starting to read anywhere lets it find you again.

Does voice following work in my language?

It works in every language your iPhone can recognise — 60+ locales on current iOS, which covers all sixteen of the app’s interface languages. Spoken commands are a shorter list, ten languages. If your language is not recognised at all, Sound mode still follows you, because it listens for speech rather than for words.

Does voice following need an internet connection?

It depends on your language, and the app tells you which before you start. Where your language has an on-device model, recognition runs on your iPhone. Where it does not, iOS sends the audio to Apple to convert it to text, so that language needs a connection — the language picker marks those languages in advance, and a banner says so during the take. Sound mode follows the presence of speech rather than the words, and Constant does not use the microphone at all.