Image + audio to expressive digital human video
Turn a single image, a voice or music track, and an optional prompt into a lifelike speaking or singing video with synchronized lips, emotion-aware movement, and camera control.
Preview an OmniHuman result and fill its inputs.
Make her sing confidently into a microphone with natural lip sync
Use OmniHuman 1.5 for digital avatars, character performances, AI music videos, marketing clips, and story-driven short videos from a compact set of inputs.
Generate an MP4 digital human video from one reference image and one audio file. Clear, high-resolution faces and clean audio usually produce the most stable result.
Speech, singing, rhythm, and emotional tone can guide lip sync, facial expression, pauses, gestures, and body movement instead of only moving the mouth.
Add prompt instructions for action order, camera movement, focus changes, environmental reactions, or visual details that should appear during the performance.
Use people, pets, animated characters, or multi-subject scenes. Optional subject detection can help select the target speaker when you need a specific character to speak.
Create a digital human video with clean inputs, optional subject targeting, and a concise performance prompt.
Use a JPG, PNG, or WebP image under 10MB. For best quality, choose a clear subject with visible facial detail; multi-subject images can use target speaker selection.
Upload MP3, WAV, AAC, OGG, or MP4 audio under 10MB. Keep it under 60 seconds; 15 seconds or less is recommended. Run subject detection when you need a specific speaker.
Choose 720p or 1080p, add a short prompt for gestures, emotions, camera movement, or scene interaction, then generate and download the final MP4 video.
Answers to common questions about OmniHuman 1.5 inputs, prompts, target speakers, video quality, and credits.
Use our AI image prompt gallery to design scenes and characters, then bring them to life with OmniHuman 1.5.
Browse AI image prompts →