Transcription now runs on your own machine, for free. English speech is transcribed on-device by default — no API key, no per-minute cost, nothing sent to a provider. Also: every voice picker now lets you hear the voice before you choose it, scenes stop leaving silence at the end, and your AI agent can finally edit the spoken cut itself.
Transcription runs on-device, free
- English transcription now happens on your Mac or PC by default, using the speech engine that ships with the app. No API key is needed and no audio leaves your machine for it. A cloud key stays the fallback for what on-device cannot do: other languages, an install that has not downloaded the speech model yet, or a failed run.
- Previously the on-device engine only ran when you had *no* API key configured at all — so anyone with a key kept paying per minute while a fully working local engine sat unused on disk.
- Settings now shows which engine will actually run, so you are never guessing.
- Transcription errors tell the truth now. A billing problem ("you have no credits remaining") used to appear as "Connection error." after a 30-second wait. The real reason now appears in under a second, and it names the on-device alternative.
Hear a voice before you pick it
- Play/stop buttons in every voice picker — the voiceover settings, the scene panels, the cloning screen. Choosing a narrator no longer means generating a whole video to find out.
- Voice cloning got an audition step: browse for your own audio file and hear the clone speak before you commit it to a project.
- Server-side and agent-driven builds now honor the voiceover choice you made in Settings — including the free local one — instead of quietly falling back to a paid provider.
Scenes fit their voiceover
- A scene now ends when its narration ends. Generated scenes used to keep a fixed length, so a short line left seconds of silence before the next scene. Lengths are measured from the audio itself, both ways — shrinking away trailing silence and growing when the narration runs long.
Your AI agent can edit the spoken cut
For people who drive NeoMotion from Claude, Codex, Cursor or ChatGPT:
- Speech editing by script. Your agent can read the timeline as an editable transcript, remove filler words and long pauses, cut or reorder what is said — and every caption, overlay, sound effect and speaker-frame stays locked to the right word automatically.
- Captions are editable: restyle them, highlight individual words, fix a misheard word (which corrects it everywhere at once), and export subtitles as SRT or VTT.
- New things it can make: AI video clips, background music, and sound effects — each lands in your asset library as a normal, reusable file.
- Safer project handling: deleting a project is now undoable, and one progress check covers every long-running job instead of a different one per task.
- Agents are also steered to the higher-quality scene-authoring path by default, and are stopped from silently spending your managed AI credits when they meant to use their own.
Fixes
- The red label that flashed at the start of every scene is gone. A stacked-caption feature meant for talking-head edits was being applied to every generated video, stamping the first few words of the narration over the opening seconds in a red box nobody asked for.
- Scene sliders that did nothing now work. In some AI-authored scenes the position and size controls in the Properties panel wrote values the scene ignored, so nothing moved. The app now refuses to accept a scene whose controls are not actually connected.