Is digitalhuman a general-purpose text-to-video model?
It is mainly used for lip-synced narration based on face images or videos. The text field is used for narration content and should not be understood as being able to generate arbitrary videos solely from scene descriptions. To create complex environments, shots, or stories, choose a more suitable creation method.
Must character assets be videos?
They do not have to be limited to videos; digital human creation also supports starting from face images. The API provides image_url and video_url respectively, which can be prepared according to the asset type. For first-time use, it is recommended to first create a short test clip, confirm the character's performance, and then use the same assets for subsequent tasks.
Can I use my own recording or script?
The API provides audio_url, as well as text and voice_id fields, allowing creation around recordings or scripts. The two approaches should not be mixed without validation; first choose the narration method, test the corresponding asset combination, and then use successful configurations for ongoing production.
Can I clone a voice directly in a generation request?
The supporting MCP tool supports cloning voices through short reference audio, but the video generation API does not have a dedicated field for cloning samples. Voice preparation should be completed separately before arranging narration generation; do not confuse character-driving audio with voice cloning reference materials.
How do I obtain the video after generation?
Submit a task through POST /digital-human/videos; the response provides the video URL and task-related fields. When using an asynchronous workflow, you can use callbacks or MCP task queries to obtain progress, determine from the returned status whether a usable result is available, and then download the video for post-production.