Made with this app
How it works
OmniHuman v1.5 transforms a single portrait photo and an audio file into a fully animated talking avatar video, complete with synchronized lip movement, natural facial expressions, and speech-matched gestures. The output is a clean MP4 at 720p or 1080p, making it practical for virtual presenters, branded spokespersons, e-learning narrators, and anyone who needs a polished talking-head video without camera equipment or a studio.
How to use this app
- 1
Describe what you want to create, or upload your source file.
- 2
Pick your options, then press Generate.
- 3
Watch your result appear in the gallery within moments.
- 4
Download, share, or generate again with a new idea.
What you can make
Virtual Course Instructors
Upload a professional headshot and a recorded lecture script to produce an on-screen instructor who speaks, gestures, and reacts naturally. OmniHuman v1.5 keeps facial expressions aligned with vocal tone, so the avatar feels engaged rather than robotic across the full 60-second audio window at 720p.
Multilingual Product Demos
Record a voiceover in any language, pair it with one brand-representative photo, and generate a localized spokesperson video in minutes. Because OmniHuman v1.5 derives gestures from speech rhythm, the body language adapts to each language's natural cadence without any manual keyframing.
Corporate Internal Announcements
HR and communications teams can turn a single executive headshot into a professional video message for company-wide distribution. The 1080p output option ensures the video holds quality on large screens, while perfect lip sync maintains credibility when the executive cannot appear on camera directly.
Social Media Persona Videos
Content creators who prefer to stay behind the scenes can build a consistent on-screen persona from one portrait. Feed in scripted audio for each post and receive an MP4 avatar that maintains the same face, expressions, and gestural style across every video in a series.
Prototype Marketing Spots
Before committing to a live shoot, marketing teams can validate scripts and visual direction by generating a realistic presenter video from a stock or concept photo. The 30-second 1080p limit is well matched to typical ad and social video formats, making review cycles faster and cheaper.
Prompt ideas to try
- Animate this portrait with a calm, authoritative tone: the speaker explains a new product feature in 25 seconds, nodding gently and maintaining steady eye contact throughout.
- Use this headshot and audio clip to create a 30-second 1080p welcome message for a company onboarding video, with warm expressions and natural hand gestures.
- Generate a 60-second 720p talking avatar of this person delivering a motivational speech, with gestures that rise in energy as the audio reaches its climax.
- Animate this portrait to match an enthusiastic product testimonial audio file, showing raised eyebrows and broad smiles at the positive moments in the script.
- Create a multilingual spokesperson video at 1080p: use this professional headshot with a 28-second Spanish-language audio narration, preserving natural gesture rhythm.
- Produce a 720p avatar video of this illustrated character portrait delivering a 45-second explainer, with measured gestures and focused expressions suited to a technical audience.
Why creators use this app
- Single Photo Input
- Perfect Lip Sync
- Natural Expressions
- Gesture Generation
- Turbo Mode
- 720p/1080p Output
Tips for better results
Choose a Clean Portrait
OmniHuman v1.5 builds all animation from a single photo, so image quality matters. Use a well-lit, forward-facing portrait with a neutral background. Avoid heavy filters, extreme angles, or partial face crops, as these constrain how accurately the model can generate lip movement and expressions.
Match Audio Length to Resolution
The app supports up to 60 seconds at 720p and 30 seconds at 1080p. Plan your script to fit within the correct limit before recording. Trimming audio after the fact often creates abrupt endings, so write and time your narration first, then choose the resolution that fits your audio duration.
Use Expressive Audio for Better Gestures
Gesture generation in OmniHuman v1.5 is driven by speech rhythm and emotion in the audio. A flat, monotone recording produces minimal movement. Record with natural pacing, varied emphasis, and clear emotional intent to get a more dynamic, lifelike avatar with appropriate head tilts and body gestures.
Use Turbo Mode for Fast Iteration
Turbo Mode accelerates generation, making it ideal for testing different audio takes or portrait choices before committing to a final render. Draft your avatar at 720p with Turbo Mode first, confirm the result meets your needs, then generate the polished 1080p version for distribution.
When to choose this app
OmniHuman v1.5 is the right choice when realism and expressive synchronization are the top priorities. Compared to SadTalker, it produces noticeably more natural facial movement and full gesture animation rather than head motion alone. Compared to Sync-3 Lipsync, it generates a complete animated avatar from a single photo rather than requiring existing video footage as input. If your goal is a credible virtual presenter from minimal source material, OmniHuman v1.5 is the most complete solution in the lineup.
Frequently asked questions
Can I use an illustrated or stylized portrait instead of a photograph?
OmniHuman v1.5 can animate non-photographic portraits such as digital illustrations or rendered characters, but results are most reliable with realistic human portraits. Highly stylized images with exaggerated proportions or non-human features may produce inconsistent lip sync and expression accuracy.
Does the model preserve the identity and appearance of the person in the photo?
Yes. OmniHuman v1.5 keeps the face, skin tone, hair, and overall appearance consistent with the input portrait throughout the generated video. It animates the existing face rather than replacing it, so the avatar looks like the person in the original photo.
What audio formats and quality levels work best?
While the app accepts standard audio files, clear speech recorded at a consistent volume level without background noise or heavy compression produces the most accurate lip sync. Avoid heavily processed audio with reverb or overlapping voices, as these can confuse the rhythm detection that drives gesture generation.
Is there a way to control the intensity of gestures or expressions?
Gesture and expression intensity in OmniHuman v1.5 are derived automatically from the emotion and rhythm of the audio rather than manual controls. Adjusting the delivery, pacing, and expressiveness of the recorded audio is the primary way to influence how animated the avatar appears in the final video.
Can the avatar speak in any language?
OmniHuman v1.5 generates lip sync and gestures from the audio waveform and its emotional content rather than from specific language rules, so it works across languages. Lip accuracy may vary slightly for languages with phoneme patterns less common in the training data, but overall results are broadly multilingual.
What should I do if the lip sync looks slightly off in places?
Re-recording the audio with cleaner pronunciation and a slightly slower pace often resolves minor sync issues. You can also try Turbo Mode to quickly test an alternate audio take before spending credits on a full 1080p render. Background noise in the original audio is the most common cause of sync drift.
Which AI model powers this app?
This app runs on OmniHuman v1.5, available through Arteza with no separate account or setup.
Can I use the results commercially?
Yes. Content you generate is yours to use, subject to our content licenses.
How long does a generation take?
Most generations finish in under a minute, and you can watch progress live in the gallery.