Turn static key visuals into sound-enabled multi-shot short films
wan2.6-i2v is the standard image-to-video model in the Alibaba Wan 2.6 series. It uses an image as the starting frame, then uses text to describe actions, camera movement, and plot. It supports audio-video synchronization and multi-shot storytelling, making it suitable for turning product key visuals, character illustrations, or scene concept art into short films. Compared with creation methods that only add slight motion to an image, it is better suited to visual expression with a clear beginning, development, and conclusion.
Input parameters and result formats vary by service. Use the public API for this model and follow its guide for generation, task retrieval and editing operations.
Specifications and API features
Creation method
Generate video from an image starting frame + text prompt
POST /wan/videos;model=wan2.6-i2v;action=image2video
Task delivery
Supports asynchronous tasks, returning a task ID and video result link
Duration, resolution, and frame rate are native model specifications; the platform submits requests and retrieves results through image URLs, prompts, and task APIs.
Core capabilities
Develop actions from key visuals
The image provides the opening composition, while the prompt explains what happens next. You can design shots around product displays, character actions, or environmental changes without having to rebuild the entire opening from text. Clearly specify the subject, direction of movement, and camera changes to make static assets the visual starting point of a short film.
Organize stories with shot sequences
Multi-shot storytelling is suitable for writing a short film as a concise shooting plan—for example, showing the environment first, then presenting the subject's actions, and finally closing on a product shot. The creative focus is not merely stacking style terms, but explaining what happens in each segment and their sequence, so limited duration can carry a clear information structure.
Express through sound and image together
wan2.6-i2v has native audio-video synchronization capabilities, allowing visual actions and sound atmosphere to be considered together during creation. Prompts can describe desired scene sounds and emotions to help convey the creative goal of a sound-enabled short film. The final video should still be reviewed by watching and listening, especially to ensure that the audio matches the actions and overall rhythm.
Use Cases
Animating Product Key Visuals
Use a product key visual or campaign image as input, describe the camera moving closer, subject presentation, and background changes, and generate a short video for product introductions or social sharing. When creating it, first determine whether to emphasize the appearance or the usage scenario, then design actions around one key point to avoid requesting too many detail changes at once.
Character Illustration Story Previsualization
Use a character illustration or concept art as the starting frame, together with a brief plot description, to generate a previsualization video of action and camera progression. It is suitable for discussing character entrances, emotional turns, and scene atmosphere. The deliverable can serve as storyboard communication material, after which the creative team can choose directions worth continuing to produce.
Campaign Visual Short Video Drafts
Starting from an existing poster or scene concept image, try different actions, camera movements, and sound atmospheres to create comparable video drafts. It is suitable for discussing opening appeal and narrative order before formal production; text titles, brand marks, and calls to action can be typeset consistently in post-production for easier control of the final presentation.
How to Choose This Model
Choose i2v, t2v, or r2v Based on Your Assets
If you already have a clear starting image and want to develop it into a dynamic scene, choose wan2.6-i2v. If you only have a text concept, choose wan2.6-t2v; if you want to extract a person's appearance from an existing video and use it in a new scene, consider wan2.6-r2v. The three are designed for creation based on image, text, and video references respectively, and should not be mixed simply by replacing asset fields.
Choose Standard or Flash Based on Your Iteration Goal
wan2.6-i2v is the standard image-to-video model, while wan2.6-i2v-flash emphasizes fast generation and is suitable for first exploring multiple camera directions. Compared with wan2.5-i2v-preview, 2.6 explicitly supports multi-shot storytelling and also offers a wider native duration range. If the task requires plot progression within a short video rather than displaying a single action, 2.6 is a better fit.
Get Started
First Prepare the Input for This Task
Prepare a publicly accessible first-frame image, submit it using image_url, and describe the subject's actions and camera movement in prompt.
Select the Actual Model and Output
Specify model=wan2.6-i2v and action=image2video for /wan/videos, starting with 5 seconds and 720P; Wan 2.6 does not use Wan 3's 30-second duration, automatic duration, or all-purpose media parameter.
Retrieve the Video and Check Audio and Visuals
Use async=true to save task_id, query /wan/tasks or receive a callback; after completion, retrieve video_url, check the subject, actions, and sound, then proceed to editing.
Trial Recommendation: Animate the Main Product Image
Input and Goal
Use the product image as the first frame, keep the cup centered in the frame, slowly push the camera in, slightly vary the window-side lighting, keep the background quiet, and do not switch scenes.
Acceptance and Next Steps
Confirm that the image is used as the starting point, the cup remains consistent, and the camera movement is completed; choose 720P or 1080P, and do not use Wan 3's smart duration or unified media workflow.
Usage Boundaries
This is an image-to-video model. Joint control of first and last frames, video continuation, or multimedia file references should not be treated as its default capabilities. Images are used to establish the starting frame; when video reference is needed for character transfer, choose the corresponding r2v model instead of using a video URL as an image input.
The native output range is 720P, 1080P, and 2–15 seconds. Do not plan for longer durations or lower resolutions based on other Wan models. For long stories, it is recommended to split them into multiple short clips and organize them through editing; multi-shot creation should also allow enough time to avoid overly crowded actions and transitions.
An image starting point and text instructions do not mean frame-by-frame locking. Product details, character appearance, shot continuity, and sound performance all need to be checked in the final video; for text and logos that must be presented accurately, it is recommended to add them in post-production. Prioritize keeping prompts concise and clear, then gradually add complex actions.
Frequently Asked Questions
What input does wan2.6-i2v require?
Use the image URL image_url to provide the starting frame, and use prompt to describe the subject's actions, scene changes, and camera movement. When calling, select model=wan2.6-i2v and explicitly set action=image2video to avoid using the default text-to-video operation. Prompts should be written as brief shooting instructions.
How long and what resolution can the generated videos be?
It natively supports videos from 2–15 seconds, with resolutions of 720P or 1080P, output as 30 fps MP4 files. When choosing, first consider whether the action can be fully expressed within the target duration; for longer stories, split them into multiple short clips and edit them together instead of cramming all plot points into a single generation.
How can a short film have multiple shots?
This model natively supports multi-shot storytelling. In the prompt, describe the opening, main action, and ending frame in sequence, clearly specifying the subject, action, and transition relationship for each shot. Each segment should revolve around the same subject and plot, reducing conflicting composition requirements; the number of shots and transition effects are not locked in item by item, so check whether the transitions are natural after generation.
Does it support videos with sound?
wan2.6-i2v natively supports audio-video synchronization and can be used to create short films with sound. You can describe the sound atmosphere and its relationship to on-screen actions in the prompt, but this is not equivalent to precise dubbing or voice cloning. Before publishing, you should still preview the finished video and check the relationship between the audio content, rhythm, and visuals.
How do I obtain the generated result after submission?
You can set async=true to submit asynchronously and obtain a task_id, then query the final task status through /wan/tasks. Successful results provide a video link, and the response also defines video dimensions and thumbnail fields. The business workflow should first confirm that the task succeeded before proceeding to preview, save, and publish.