Which voice cloning API works directly in a TTS request?
Condition a line from a voice description, a description with reference audio, or reference audio with its transcript. You do not need to create a saved voice first.
Which voice conditioning method should I use?
Description only
Describe the voice and delivery in natural language when you do not have a recording.
Description plus reference
Combine a voice description with reference audio when the recording and written direction should work together.
Reference plus transcript
Use a clean recording and its exact spoken transcript for reference-led voice conditioning.
How do I condition or clone a voice?
Pick one conditioning path
Use voice description, voice description with reference audio, or reference audio with reference transcript.
Do not mix description and transcript
Voice description and reference transcript are mutually exclusive in one request.
Generate with permission
Use only recordings and identities you are authorized to use, and preserve consent records for saved voices.
What else should I know?
Can voice description work without reference audio?
Yes. Voice description can define identity and performance on its own.
Can I combine voice description and reference audio?
Yes. A description can guide the request while the reference recording contributes acoustic information.
Can I send voice description and reference transcript together?
No. Use voice description or reference transcript, not both. Reference transcript belongs to the reference-audio-led path.
