Sign inStart creating

Narration Mode

Book recaps and narrated shorts: the voiceover is generated first and the visuals are cut to fit it.

Open it →

Studio → a project → Narration.

For book recaps, narrated shorts and documentary pieces, where scenes are illustrated rather than acted.

### The pipeline runs backwards on purpose

Normal order: write, shoot, cut picture, lay audio against it.

Narration order: generate the voiceover first, measure how long each segment actually takes to say, then generate a visual to fill exactly that long.

This is not a stylistic choice. In a narrated piece the narrator's track is the spine — you cannot trim it without cutting words. Generate picture first and every segment needs re-timing or regenerating to fit the voice. Generate audio first and the durations are known before a single picture credit is spent.

### Estimates are never presented as measurements

Until a voiceover exists, durations are word-count estimates and are labelled as such. A word-count guess is out by 15–20% on ordinary prose, and 20% of a 40-second segment is eight seconds of black.

Illustration stays disabled until at least one segment has a real measured duration. Generating picture against an estimate produces shots that are confidently the wrong length, and you find out after the credits are gone.

Once some are measured, the page shows the drift between estimate and reality — which is the clearest possible argument for the ordering.

### Splitting

Segments are split by rhythm, not word count: paragraph breaks first, then sentence ends. A visual that changes mid-sentence reads as a mistake.

### Cost

No one is on screen speaking, so there is no lip-sync requirement — the single most expensive capability in the router. Narration segments route to the open-model tier by default. This is the high-volume, low-cost category and it is priced like one.

Did this answer your question?

Related

All help articlesAsk the community