Made with this app
How it works
Wan 2.2 S2V turns a single still photo and a spoken audio clip into a short MP4 video where the subject moves and speaks in natural, speech-driven sync. Upload your photo, attach up to 7.5 seconds of audio, and the model generates lip movement and facial motion that follows the rhythm and cadence of the speech. It outputs at 480p, 580p, or 720p, making it a practical tool for digital presenters, branded spokespersons, and anyone who needs a realistic talking avatar without recording live video.
How to use this app
- 1
Describe what you want to create, or upload your source file.
- 2
Pick your options, then press Generate.
- 3
Watch your result appear in the gallery within moments.
- 4
Download, share, or generate again with a new idea.
What you can make
Digital Presenters for Slide Decks
Drop a professional headshot and a recorded voiceover into Wan 2.2 S2V to produce a talking presenter clip you can embed in a presentation or landing page. The speech-driven motion keeps the avatar looking attentive and natural rather than stiff or looping.
Localized Spokesperson Clips
Record a translated voiceover from the same source script, attach it to a single brand spokesperson photo, and generate a localized avatar video for each language. The model follows the new audio cadence without requiring a new photo shoot.
Animated Memorial or Tribute Photos
Give a cherished portrait a voice by pairing it with a recorded message. Wan 2.2 S2V generates subtle, lifelike facial motion that makes the photo feel present and engaged rather than artificially animated.
Social Media Talking Head Content
Produce short, polished talking-head clips at 720p for social feeds without camera equipment. Pair a clean portrait with a scripted audio clip to create consistent on-brand content at scale.
E-Learning Character Narrators
Pair an illustrated or photographic character with a narration track to build an animated instructor for a course module. The audio-length output means the video runs exactly as long as the lesson segment, with no padding required.
Prompt ideas to try
- A confident professional in a grey blazer speaks directly to camera in a softly lit modern office, maintaining steady eye contact throughout.
- A friendly customer service representative with a warm smile delivers a short welcome message, head slightly nodding as she speaks.
- An elderly man seated by a window speaks slowly and thoughtfully, natural wrinkle movement and soft ambient light visible throughout.
- A young woman with natural hair speaks enthusiastically, subtle head movement and expressive brow motion matching her upbeat audio delivery.
- A corporate spokesperson in formal attire addresses the viewer with measured, authoritative speech and minimal background distraction.
- A historical portrait-style subject speaks in a composed manner, lighting consistent with a classical painted background.
Why creators use this app
- Photo + Audio Input
- Speech-Driven Motion
- 480p / 580p / 720p
- Audio-Length Output
Tips for better results
Keep Audio Under the Limit
The audio limit is 7.5 seconds. Trim your clip precisely before uploading. Audio that cuts off mid-sentence can produce abrupt motion at the end of the video, so plan your script to land a complete thought within the limit.
Use a Clear, Front-Facing Photo
Wan 2.2 S2V generates speech-driven motion from facial landmarks in the source image. A front-facing portrait with visible lips and neutral expression gives the model the clearest reference and typically produces the most accurate lip sync.
Match Resolution to Your Use Case
Choose 720p for final deliverables where quality matters and 480p for rapid iteration or small-format embeds. Settling the resolution before you finalize your audio avoids regenerating the same clip at a different quality tier.
Write a Descriptive Prompt
The model accepts a text prompt alongside the photo and audio. Use it to describe the setting, lighting style, and mood you want. A prompt like 'warm studio lighting, steady eye contact, professional demeanor' helps guide motion style and visual coherence.
When to choose this app
Choose Wan 2.2 S2V when you need speech-driven facial motion that responds to the actual cadence and rhythm of your audio rather than a generic mouth loop. Compared to SadTalker, which is positioned as a budget avatar option, Wan 2.2 S2V offers higher resolution tiers up to 720p and accepts a guiding text prompt. Compared to Kling Avatar v2, which focuses on versatile lip sync for any character type, Wan 2.2 S2V is purpose-built around the photo-plus-audio pipeline with audio-length output, making it the cleaner choice when your source material is already a portrait and a recorded voiceover.
Frequently asked questions
What file formats does the audio input need to be in?
The app accepts an audio file paired with your photo. Use a clean, clear recording with minimal background noise for best results. The model derives all facial motion from the speech signal, so audio quality directly affects the naturalness of the output.
Can I use an illustrated or AI-generated portrait instead of a real photograph?
Yes. Wan 2.2 S2V works from any still image that contains a clearly visible face with recognizable facial landmarks. Illustrated characters, digital paintings, and AI-generated portraits can all be animated, though results are most reliable when the face is front-facing and unobstructed.
Does the output video include the original audio track?
The output is an MP4 video. The audio you upload drives the motion, and the rendered file includes that audio so playback is already in sync without requiring a separate merge step in your editing software.
What happens if my audio is shorter than 7.5 seconds?
The output duration matches your audio length exactly, which is what the audio-length output feature means. If your clip is three seconds, you get a three-second video. There is no padding or looping added after the speech ends.
How does the text prompt affect the final video?
The prompt guides visual style, mood, and environment rather than overriding the photo content. Describing lighting, setting, and the subject's demeanor helps the model produce motion and atmosphere consistent with your intent, particularly when your source photo has an ambiguous or plain background.
Is 720p the maximum resolution available?
720p is the highest resolution tier currently available in Wan 2.2 S2V. If your project requires 4K lip-sync output, Sync-3 Lipsync, which handles video dubbing at 4K, may be more appropriate for that specific requirement.
Which AI model powers this app?
This app runs on Wan 2.2 S2V, available through Arteza with no separate account or setup.
Can I use the results commercially?
Yes. Content you generate is yours to use, subject to our content licenses.
How long does a generation take?
Most generations finish in under a minute, and you can watch progress live in the gallery.