People who have worked in content operations, training materials, or product documentation have probably seen a process like this: after the copy is finalized, it is sent to a voiceover colleague, then there is waiting for scheduling, waiting for revisions, and waiting for export; after receiving the audio, it is uploaded to cloud storage or object storage, and finally the link is filled back into the page.
None of the steps are difficult, but once there are too many stages, even an audio clip of just a few dozen seconds may take one or two days. Especially for content such as internal training, feature updates, and event notifications, what is truly needed is often not complex production, but a clear, stable audio file that can be played directly.

When copy, generation, uploading, and distribution are scattered across different tools, the voiceover workflow can easily be prolonged.
¶ The pain point is not only “generation”
Many TTS demos look very simple: enter text and get audio. But when put into real business scenarios, several other problems arise:
- After the API returns audio data, you still need to save the file yourself, upload it to storage, and generate a link;
- Long-text generation takes longer, and it is difficult to determine the task status after an HTTP connection is interrupted;
- Operations teams need to know which audio clips have been completed, while developers also need to record the cost of each generation;
- MP3, WAV, and PCM are intended for different use cases, so the same configuration cannot be used for all of them.
Therefore, the more practical goal is not “being able to generate a voice clip,” but turning text into an audio result that is playable, traceable, and able to continue flowing through the process.
¶ Get a playable link with one API
The AceDataCloud Fish TTS API endpoint is:
POST https://api.acedata.cloud/fish/tts
Use the platform Token in the request headers:
authorization: Bearer {token}
content-type: application/json
A minimal request only needs to provide the text and output format:
{
"text": "The product update has been completed. Welcome to try the new features.",
"format": "mp3"
}
After successful generation, the response provides audio_url. It points to the platform CDN and can be placed directly into a web player, or passed to message notifications, video composition, or asset management systems. The cost in the response can be used to record the actual consumption for this generation.
This step eliminates the common process of “receiving binary data—uploading to object storage—configuring an access URL.” However, for formal business use, it is still recommended to retain a copy of the audio according to your own archiving strategy, rather than treating an external link as the only backup.
¶ Choose the right output format first
The three common formats each have their own focus:
- MP3: Small in size, suitable for web playback, message sharing, and most content distribution;
- WAV: Suitable for editing, mixing, and loudness-processing workflows;
- PCM: Suitable for further processing in clients or audio-processing pipelines. The API returns 16-bit PCM in a WAV container.
You can also set the sample rate through sample_rate, and choose between 64, 128, and 192 kbps through mp3_bitrate. Rather than pursuing the highest specifications from the beginning, it is better to first determine a standard based on the delivery channel, reducing repeated transcoding later.
¶ Do not keep waiting on the connection for long text
For short text of just a few sentences, synchronous calls are usually sufficient. Articles, courses, and longer narrations may take more than ten seconds or even longer. If the frontend keeps waiting, network fluctuations can easily lead to duplicate submissions and unclear statuses.
At this point, you can add callback_url to the request:
{
"text": "Here is a longer course narration…",
"format": "mp3",
"callback_url": "https://example.com/webhooks/fish-tts"
}
The API first returns task_id and started_at, and then POSTs the final result to the callback address after the task is completed. You can also actively query the task using the Fish Tasks API. callback_url is an extension added by AceDataCloud relative to the official API, making it more suitable for batch and long-text business scenarios.

Tasks are generated asynchronously after submission, and results automatically return to the content system, so operations teams do not need to keep watching the page.
¶ A simple but sufficient workflow
In actual implementation, the process can be kept to six steps:
- Content staff submit and confirm the copy;
- The backend calls Fish TTS;
- Short content obtains results synchronously, while long content saves the
task_id; - Update the task record after receiving
audio_urlandcost; - The audio enters previewing, the page player, or downstream video workflows;
- Handle failed tasks based on error codes to avoid unconditional repeated submissions.
It is recommended that the task table record at least: business content ID, text version, output format, task ID, audio URL, cost, status, and completion time. This way, after the copy is modified, it is possible to clearly know which version of the audio it corresponds to.
¶ Generation completed does not mean it can be published directly
AI voiceover can shorten the delivery time for standardized content, but content intended for public release still needs to be reviewed by listening. Focus on checking numbers, English abbreviations, proper nouns, sentence breaks, and volume. Content such as brand advertisements and character performances that have higher emotional requirements may also require human recording or further post-production.
A more appropriate division of work is: the API is responsible for stable generation and delivering links, while the team is responsible for the text, listening experience, and publication decisions. This can improve efficiency without treating automation itself as a guarantee of quality.

