The new all-in-one flagship model for expressive audio & videoView plans →
Vidu Q4 is the newest AI video model from ShengShu Technology, released as a preview on October 7, 2026. Unlike models that add sound afterwards, Vidu Q4 generates picture and audio together, so dialogue, sound effects and lip movement come out of the same generation.
This site gives you Vidu Q4 in a simple web workspace. Use reference-to-video to blend up to 12 images and 3 voice clips into one consistent scene, or image-to-video to bring a single photo to life. Every Vidu Q4 clip runs 3 to 16 seconds and can be rendered from 540p up to 4K.
Reference-to-video and image-to-video, with native audio
Combine up to 12 reference images and 3 voice clips — characters, style, scenes and sound stay in sync.

Picture and sound are generated together. Reference up to 3 voices for natural emotion and matched lip movement.
Multi-shot stories with smooth camera cuts in a single generation, rendered in 2K or 4K.
Bring one image to life in 3–16 seconds, keeping its aspect ratio — with fluid action and vivid effects.

Four steps from idea to finished clip. You don't need editing software or a GPU — Vidu Q4 runs in your browser.
Pick Reference to Video when the same characters, products or voices must appear across shots. Pick Image to Video to animate one photo.
Add up to 12 reference images and 3 voice clips, or a single first frame. Save a character as a subject to reuse it in every new project.
Write what happens, how the camera moves and what we should hear. Insert each reference as a tag so the model knows who is who.
Set 3–16 seconds, an aspect ratio and a resolution from 540p to 4K. If a run fails, its credits are refunded automatically.
The key numbers for Vidu Q4 Preview as offered on this site.
Benchmark data as of October 2026.
Vidu Q4 shines wherever a character, product or voice has to stay the same from shot to shot.
Keep the same cast, costumes and voices across episodes. Vidu Q4 cuts between shots inside one clip, so a scene plays out like a real edit.
Turn a product photo and a brand voice into a polished spot in 4K, then make vertical and square versions for every channel.
Show a product from several angles with logos and colors kept intact. Upload reference shots and let the model handle the camera.
Create 9:16 videos with dialogue and sound already in place, ready for TikTok, Reels and Shorts without extra editing.
Test performances, storyboards and pacing before a shoot. It gives directors a fast way to see and hear a scene.
Bring portraits, artwork or travel shots to life with image-to-video, keeping the original framing while adding natural motion.
Vidu Q4 follows long, specific prompts well. These habits give the most reliable results.
Say which reference is the main character and list what must not change: face, outfit, logo, colors.
Split a longer clip into timed beats, such as 0–3s close-up, 3–7s tracking shot. Vidu Q4 follows the sequence and cuts between shots.
Use film language: close-up, low-angle tracking, slow push-in, hard cut. Clear camera notes give smoother motion.
Spell out the dialogue, background noise and whether you want music. Add a voice clip to keep a character's voice consistent.
Exclusions such as "no extra people, no captions" keep the frame clean and on-brief.
Test ideas at 540p or 720p to save credits, then render the final version in 1080p, 2K or 4K.
Vidu Q4 Preview supports two modes, both with native audio: image-to-video, which animates a single picture, and reference-to-video, which combines several images and voice clips into one consistent video. Text-only prompts are not supported in Q4 yet.
Image-to-video starts from one image plus a description of the motion and camera. Reference-to-video takes 1–12 reference images and up to 3 audio clips, so characters, objects, scenes and voices stay consistent throughout the video.
Up to 16 seconds per generation. With image-to-video you can choose any length from 3 to 16 seconds.
540p, 720p, 1080p, 2K and 4K. The 2K and 4K outputs use 10-bit color depth.
Reference-to-video offers 16:9, 9:16, 1:1, 4:3 and 3:4. Image-to-video keeps the aspect ratio of the photo you upload.
Yes. Audio is generated together with the picture. You can add up to 3 reference audio clips to keep a character’s voice and emotional delivery consistent, with lip movement matched to the speech.
Most videos finish within a few minutes. Longer clips and higher resolutions take more time, and you can leave the page — finished videos stay in your recent results.
Vidu Q4 raises the output ceiling to 4K with 10-bit color, where Vidu Q3 Pro tops out at 1080p. It also scores higher in blind tests: 1,179 versus 1,056 for Vidu Q3 Pro in the Artificial Analysis image-to-video arena (October 2026).
ShengShu Technology released Vidu Q4 Preview on October 7, 2026. It is an early public version that comes ahead of the full Q4 model.
AI short dramas, ads and marketing, e-commerce product videos, social content, and film pre-visualization such as testing performances, storyboards and pacing.
Creating an account is free. Generating videos uses credits, and longer or higher-resolution videos use more. See the Pricing page for current plans and any free credits on offer.
Signing up is free and needs no credit card. New accounts may get starter credits to try Vidu Q4; after that, credits come from a monthly plan or a credit pack.
No. If a generation fails, the credits it used are returned to your account automatically.
No. This is an independent site and is not affiliated with Vidu or ShengShu Technology.