Studio → Voices.
Four ways in, and they end the same place: hear a test line before you cast it.
### 1. The library
Twelve presets with a known sound and no consent requirements. Ready to cast immediately.
### 2. Describe one
"Gravelly, mid-fifties, slight Irish lilt, speaks slowly."
What was recognised — gender, age, accent, pitch, pace, texture, energy — is shown back to you. Anything the parser did not classify is still passed to the engine verbatim, so the detail you cared enough to type is never silently dropped.
A described voice cannot be cast until it has been auditioned. That gate is deliberate. "Warm" means something different to you and to the model, and the gap is invisible until you hear it — casting a whole film in a voice nobody has listened to means finding out after two hundred lines of dialogue.
### 3. Upload a recording
WAV, MP3, M4A, OGG or WebM. About 30 seconds of clean, varied speech gives a good clone; under 10 seconds is guesswork, and you are told the length before it uploads.
### 4. Record one here
Press record and read the provided passage — it has varied vowels, a clipped line and a warm one on purpose. A single flat sentence tells the model far less about how you actually speak.
Noise suppression and automatic gain are switched off for the capture. They flatten exactly the texture a clone needs.
### Consent is enforced, not requested
Uploading or recording someone's voice creates a clone of a real person. A cloned voice cannot be cast until consent is recorded — the assignment is refused in code, not warned about.
Recording your own voice is self-attested: the consent record still exists for the audit trail, but requiring you to formally consent to yourself is friction with no protective value.
Consent can be revoked, and revocation takes effect at use rather than at assignment — the voice stops being sent to any provider from the next generation onward, including for characters already cast in it.
### Auditioning
Pick a test line, an emotion and a language, and press Hear it. Not right? Another take re-renders with a different seed — the same voice, a different read. Previous takes are kept so you can compare.
Without TTS_API_KEY auditions render on the local placeholder engine. The timing and the line are real; the timbre is not. It is useful for checking length and pacing and tells you nothing about how the voice sounds — and it says so rather than letting you cast on it.
### Voices reach the video model, not just the TTS pass
Models with native audio — Seedance, Veo, Kling — accept a voice reference and generate the speech as part of the shot. The mouth and the voice come out of one generation, so they cannot disagree: the mismatch is not corrected, it never occurs.
Models without native audio fall back to TTS-then-lip-sync, so the reference is offered rather than required. The Router picks the path.