Get 5-day unlimited access to Seedream v4.5 + moreup to 25% off

Discount expires in --

Made with this app

How it works

1Describe what you want to create, or upload your source file.
2Pick your options, then press Generate.
3Watch your result appear in the gallery within moments.
4Download, share, or generate again with a new idea.

The Fabric 1.0 app on Arteza takes a single portrait photo and an audio track and produces a natural talking-head video with accurate lip sync, delivered as an MP4 at 480p or 720p. The output runs exactly as long as your audio, up to 60 seconds per generation. It is built for creators, educators, marketers, and small teams who need a believable on-screen presenter without a camera, a studio, or a professional actor.

How to use this app

  1. 1

    Describe what you want to create, or upload your source file.

  2. 2

    Pick your options, then press Generate.

  3. 3

    Watch your result appear in the gallery within moments.

  4. 4

    Download, share, or generate again with a new idea.

What you can make

Course and Training Narrators

Upload a professional headshot and a recorded lesson audio to produce a presenter-led segment for an online course. Fabric 1.0 keeps the lip sync tight throughout the clip, so learners see a natural speaker rather than a static image with voice-over laid on top.

Social Media Spokesperson Clips

Record a 30-second brand message, pair it with a brand ambassador photo, and export a 720p MP4 ready for Instagram Reels or TikTok. The audio-length output means no manual trimming, and the natural lip sync holds up at close viewing distances common on mobile feeds.

Multilingual Product Announcements

Record the same script in multiple languages, then run each audio file against one portrait to generate localized presenter videos. Because Fabric 1.0 drives animation from the audio signal, swapping languages requires no reshoots, only a new audio file per generation.

Internal Communications and Updates

Create a weekly leadership message using a single approved executive photo and a fresh audio recording each week. The result is a consistent, professional presenter clip that feels more personal than a text email and requires no video production resources.

Demo Reel Stand-Ins

Agencies and freelancers can produce client-facing presenter demos quickly by pairing a stock portrait with a scripted pitch audio. At 720p the output is sharp enough for proposal decks and website embeds, letting clients visualize a finished video before committing to production.

Prompt ideas to try

  • A professional woman in a navy blazer delivers a 45-second product launch announcement, facing the camera directly with natural eye contact and calm delivery.
  • A friendly male presenter in his 30s reads a 30-second welcome message for a corporate onboarding video, speaking clearly with a relaxed but confident tone.
  • A studio-lit portrait of a news anchor reads a 60-second briefing, maintaining composed facial expressions and precise lip movement throughout.
  • A young female educator in a bright classroom setting narrates a 50-second tutorial introduction, looking engaged and approachable while speaking at a measured pace.
  • A headshot of a startup founder against a plain background delivers a 40-second investor pitch summary with steady gaze and deliberate pacing.
  • A close-up portrait of a customer service representative speaks a 35-second product FAQ response, conveying warmth and clarity in each sentence.

Why creators use this app

  • Photo + Audio Input
  • Natural Lip Sync
  • 480p / 720p
  • Audio-Length Output

Tips for better results

Choose a Front-Facing Portrait

Fabric 1.0 derives lip sync and facial movement from your photo, so a straight-on portrait with the face unobstructed and well-lit produces the most accurate results. Avoid heavy shadows across the mouth area, as these can reduce lip sync fidelity.

Keep Audio Clear and Clean

The model reads the audio signal to drive mouth movement, so a recording with minimal background noise, consistent volume, and clear enunciation yields tighter sync. Record in a quiet room and normalize your audio levels before uploading.

Match Audio Length to Your Goal

Output length equals audio length up to the 60-second limit. Plan your script to fit within that window rather than trimming afterward. For longer content, break your script into multiple 60-second segments and run separate generations.

Pick 720p for Embeds and Presentations

Choose 720p when the video will appear on a website, in a slide deck, or on a large screen. Reserve 480p for messaging apps or social stories where file size matters more than pixel density, keeping exports fast and lightweight.

When to choose this app

Fabric 1.0 is the right choice when your starting point is a still photo and a pre-recorded voice track and you want the output to match the audio duration exactly. Compared to Kling Avatar v2, which focuses on expressive full-body motion, Fabric 1.0 prioritizes tight, natural lip sync from minimal inputs. Compared to Sync-3 Lipsync, which re-syncs existing video footage, Fabric 1.0 generates the entire talking-head video from scratch using only a portrait and audio, making it faster to set up from zero assets.

Frequently asked questions

What file formats work best for the portrait photo input?

A high-resolution JPEG or PNG with a clear, front-facing face works best. The photo should have the subject centered, the face well-lit, and no large obstructions like sunglasses or masks across the mouth, since Fabric 1.0 uses the facial geometry in the image to anchor lip movements.

Can I use a synthetic or AI-generated portrait as the input photo?

Yes. Fabric 1.0 treats any portrait image the same way, whether it is a real photograph or an AI-generated face. The lip sync quality depends on face clarity and front-facing orientation rather than whether the subject is a real person.

What happens if my audio is longer than 60 seconds?

The app enforces a 60-second audio limit per generation. If your recording is longer, you will need to split it into segments of 60 seconds or less and run each segment as a separate generation, then combine the resulting MP4 files in a video editor.

Does the avatar reproduce background elements from the original photo?

Yes. The output video preserves the background and framing from your uploaded portrait. If you want a specific backdrop, edit or composite the portrait before uploading. The model animates the face within the original image rather than replacing the scene.

How does 480p compare to 720p for social media use?

720p offers noticeably sharper detail on lips and facial features, which matters for close-up talking-head clips on platforms like YouTube or LinkedIn. 480p is acceptable for smaller display contexts such as chat thumbnails, stories, or messaging apps where the player window is compact.

Will background music in my audio track affect lip sync accuracy?

A voice track mixed with loud background music can confuse the model's audio analysis and reduce lip sync quality. For best results, upload a voice-only track. If you want background music, add it in post-production after downloading the generated MP4.

Which AI model powers this app?

This app runs on Fabric 1.0, available through Arteza with no separate account or setup.

Can I use the results commercially?

Yes. Content you generate is yours to use, subject to our content licenses.

How long does a generation take?

Most generations finish in under a minute, and you can watch progress live in the gallery.

Explore more apps