See the model's 30-second continuity, multi-source reasoning, reference fidelity, visual texture, and native sound in real Wan 3.0 examples.
Thirty seconds creates room for a beginning, transition, and payoff without stitching together short clips. Use it for continuous camera moves, one-take scenes, and more developed narrative beats.
Wan 3.0 can reason over text, images, audio, video, documents, and public web pages. This dashboard currently exposes text, start/end image, and image/video reference workflows.
Reference-to-video is designed to hold faces, hairstyles, physique, clothing, accessories, product structure, logos, materials, scene blocking, and visual style steady across the shot.
Restrained micro-expressions coordinate with body language while skin, fabric, metal, light, and motion retain detail instead of flattening into a synthetic look.
Wan 3.0 composes generated audio with the image for more convincing space, motion, and impact. Play this example with sound to hear the audiovisual result.
Reference material can guide product geometry, visual identity, camera language, and pacing for advertising, automotive, apparel, software, and other production workflows.
Four decisions take you from source material to a production-ready first pass.
Use Text-to-Video for a scene built from language, Image-to-Video for a start frame and optional end frame, or Reference-to-Video when identity and source consistency matter.
Upload clear, high-resolution images or short reference videos that show the details the model must preserve.
Describe what happens over time—not just what the first frame looks like. Name the subject, action, camera path, continuity constraints, style, and sound.
Choose 480p, 720p, or 1080p, a platform-ready aspect ratio, and a 5–30 second duration. Use Prime when fidelity matters more than cost.
Wan 3.0 has enough duration to interpret sequence and camera language. Tell it what changes, what stays fixed, and what the audience should hear.
Subject + Ordered action + Camera path + Continuity + Look + Sound
Identify the character, product, place, or reference source
Sequence actions and transitions across the full duration
Specify framing, lens feel, movement, and final composition
Name every identity, geometry, colour, or style detail that must not drift
Direct ambience, dialogue, effects, silence, and music
One continuous 30-second tracking shot follows a young filmmaker through a rain-soaked night market. She passes steaming food stalls, turns into a narrow neon alley, then reaches a rooftop as dawn breaks over the city. Natural handheld camera movement, realistic reflections and fabric motion, footsteps, market ambience, distant traffic, and a restrained cinematic score.
Use Image 1 as the exact product reference. Place the watch on a dark stone pedestal while the camera makes a slow 180-degree orbit. Preserve the dial markings, case shape, logo, brushed-metal finish, and strap in every angle. A narrow beam of light travels across the surface, with subtle mechanical clicks and low ambient music.
Begin with the quiet coastal village shown in the start image at blue hour. The camera glides down the main street as shop lights switch on and people enter the frame. End on the supplied final image of the lighthouse after sunrise, with one continuous transition and consistent architecture.
Keep the character from Image 1 identical throughout the scene. She walks through the environment in Video 1, pauses beside the window, then turns toward camera with a restrained smile. Preserve her face, hairstyle, red coat, silver earrings, and body proportions. Match the reference camera movement and natural room ambience.
A premium thirty-second launch film for an electric concept car: macro details of the headlights and textured interior, a clean match cut to the car accelerating along a wet mountain road, then a wide sunrise reveal. Controlled reflections, realistic tire spray, precise logo geometry, deep engine-like sound design, no on-screen text.
Wan 3.0 is Alibaba's latest video generation family. It supports text-to-video, start/end image-to-video, and multi-reference video creation, with native audio, up to 1080p output, and clips as long as 30 seconds. Prime variants trade a higher credit cost for premium fidelity.
The model supports 2–30 seconds in one pass. This dashboard provides 5, 10, 15, 20, 25, and 30-second presets. Longer duration gives a prompt more narrative room, but it works best when you describe enough ordered action to fill the full shot.
Text-to-Video needs a prompt. Image-to-Video accepts a required start image and an optional end image. Reference-to-Video accepts up to 10 images and 5 videos, with reference videos totaling no more than 15 seconds.
The underlying Wan 3.0 reference model supports doc, xls, ppt, pdf, txt, key, pages, numbers, and md files, plus public web links, with one file or link per generation, up to 100MB and 50 pages. This dashboard currently exposes text, image, and video inputs rather than document or URL upload.
Wan 3.0 is designed to preserve facial features, hairstyle, physique, clothing, accessories, product structure, logos, materials, scene blocking, and stylistic tonality. Clear references from useful angles and explicit continuity instructions produce the strongest result.
Both versions expose the same core text, image, and reference workflows. Prime is the higher-cost option for shots where visual fidelity and reference accuracy matter more than generation price; standard Wan 3.0 is the economical choice for iteration.
Yes. Wan 3.0 generates audio with the picture, including ambience, effects, dialogue cues, and music direction from the prompt. Alibaba notes that audio texture still has room to improve, so review important sound in post-production.
Alibaba identifies audio texture and text accuracy as areas that still need improvement. Avoid depending on generated typography for exact spelling, and plan a finishing pass when precise text or high-fidelity sound is required.
Credits scale linearly with duration and resolution. Standard Wan 3.0 follows $0.05, $0.10, and $0.20 per second at 480p, 720p, and 1080p. Prime follows $0.068, $0.14, and $0.28 per second. For 5 seconds, that is 6/12/24 credits for standard and 9/17/34 credits for Prime.
The Wan 3.0 family supports instruction-based changes to visuals, plot, and dialogue. The six models on this dashboard currently cover text-to-video, image-to-video, and reference-to-video rather than a dedicated editing workflow.