How should I choose between tts-1-hd and tts-1?
tts-1-hd emphasizes quality, while tts-1 emphasizes real-time use. For course dubbing, article read-alouds, and narration that needs review before publication, try HD first; for real-time feedback, focus on evaluating tts-1. It is recommended to compare using the same actual script rather than interpreting the model difference as a fixed multiple of speed or audio quality.
How do I choose a voice for tts-1-hd?
Use voice to select alloy, echo, fable, onyx, nova, or shimmer; the default is alloy. Preview each one using a short script containing terms, numbers, and ordinary narration, then decide which voice suits the content. Once selected, keep the same configuration within the same series to help maintain a consistent audio production style.
Can tts-1-hd use my own recordings to customize a voice?
The creation method here is to enter text and select a preset voice, and it should not be used as a workflow for customizing a voice from a reference recording. voice contains a voice name, not a file address or person description. When voice cloning is needed, choose a product that explicitly supports reference audio, and handle authorization and production requirements separately.
Can the generated audio be used for playback and editing?
You can choose mp3, opus, aac, flac, wav, or pcm depending on the use case; the default output format is mp3. After calling, save it as binary audio or pass it to a player rather than parsing it as a text response. For editing, you must also arrange visual synchronization, background music, and final export in production software.
How do I submit a script and adjust the reading speed?
Submit input text to POST /v1/audio/speech, use tts-1-hd for model, and set voice, response_format, and speed as needed; the default value of speed is 1.0. In actual production, first preview the speaking speed with a short script, then process the main text; the text should also retain reasonable punctuation and paragraph boundaries.