Model availability and access depend on your workspace and plan.
About the model
Mera Avatar, explained.
Mera Avatar is a video model by Skytells. It supports image to video, text to video, audio to video, fast, avatar, lipsync.
Supported capabilities
Image to video
Text to video
Audio to video
Fast
Avatar
Lipsync
What you get back. Your generated video appears directly in your Relam chat.
Working with Mera Avatar
Shape the result.
Seed
Optional input
Set for reproducible generation.
Audio
Optional input
Optional uploaded audio to drive avatar speech. If provided, this is used instead of voice_script and voice settings.
Image
Required input
First frame. Supports jpg, jpeg, png, webp.
No Op
Optional input
Health check mode - returns status without inference.
Hyper Motion
Optional input
Skytells' adaptive inference engine for lip-sync under extreme motion, multi-character interactions, cinematic camera movement, and sports scenarios.
Video Prompt
Optional input
Optional visual prompt describing how the person should appear or behave while speaking.
Questions & answers
More about Mera Avatar.
Mera Avatar is a video model by Skytells. It supports image to video, text to video, audio to video, fast, avatar, lipsync.
Mera Avatar is a video model by Skytells. It supports image to video, text to video, audio to video, fast, avatar, lipsync.
Skytells is the vendor behind Mera Avatar, a video model available in the Relam catalogue.
Skytells is the vendor behind Mera Avatar, a video model available in the Relam catalogue.
Set for reproducible generation.
Set for reproducible generation.
Optional uploaded audio to drive avatar speech. If provided, this is used instead of voice_script and voice settings.
Optional uploaded audio to drive avatar speech. If provided, this is used instead of voice_script and voice settings.
First frame. Supports jpg, jpeg, png, webp. This input is required.
First frame. Supports jpg, jpeg, png, webp. This input is required.
Health check mode - returns status without inference.
Health check mode - returns status without inference.
Skytells' adaptive inference engine for lip-sync under extreme motion, multi-character interactions, cinematic camera movement, and sports scenarios.
Skytells' adaptive inference engine for lip-sync under extreme motion, multi-character interactions, cinematic camera movement, and sports scenarios.
Optional visual prompt describing how the person should appear or behave while speaking.
Optional visual prompt describing how the person should appear or behave while speaking.
Optional style instructions for how to speak voice_script, such as tone, pacing, accent, or emotion. These instructions are not spoken.
Optional style instructions for how to speak voice_script, such as tone, pacing, accent, or emotion. These instructions are not spoken.
Exact words the avatar should say. Required when no audio file is uploaded.
Exact words the avatar should say. Required when no audio file is uploaded.
Disabled if empty.Mention what you do NOT want in the video, e.g. "subtitles, text, blurry, low quality, frames, watermark, titles, scene change". We recommend using multiple keywords at once.
Disabled if empty.Mention what you do NOT want in the video, e.g. "subtitles, text, blurry, low quality, frames, watermark, titles, scene change". We recommend using multiple keywords at once.
Disable safety filter for prompts and input image. See Skytells's Responsible AI docs.
Disable safety filter for prompts and input image. See Skytells's Responsible AI docs.
Strength of the Negative Prompt. Optimal value can differ for different video lengths (Experimental Feature)
Strength of the Negative Prompt. Optimal value can differ for different video lengths (Experimental Feature)
When true, skip automatic enhancement of the visual video prompt and use video_prompt directly.
When true, skip automatic enhancement of the visual video prompt and use video_prompt directly.
Voice to use when generating speech from voice_script. Supported values are Zephyr (Female), Puck (Male), Charon (Male), Kore (Female), Fenrir (Male), Leda (Female), Orus (Male), Aoede (Female), Callirrhoe (Female), Autonoe (Female), Enceladus (Male), Iapetus (Male), Umbriel (Male), Algenib (Male), Despina (Female), Erinome (Female), Laomedeia (Female), Achernar (Female), Algieba (Male), Schedar (Male), Gacrux (Female), Pulcherrima (Female), Achird (Male), Zubenelgenubi (Male), Vindemiatrix (Female), Sadachbia (Male), Sadaltager (Male), Sulafat (Female), Alnilam (Male), Rasalgethi (Male). The model default is Zephyr (Female). Your workspace may offer a narrower selection depending on your plan.
Voice to use when generating speech from voice_script. Supported values are Zephyr (Female), Puck (Male), Charon (Male), Kore (Female), Fenrir (Male), Leda (Female), Orus (Male), Aoede (Female), Callirrhoe (Female), Autonoe (Female), Enceladus (Male), Iapetus (Male), Umbriel (Male), Algenib (Male), Despina (Female), Erinome (Female), Laomedeia (Female), Achernar (Female), Algieba (Male), Schedar (Male), Gacrux (Female), Pulcherrima (Female), Achird (Male), Zubenelgenubi (Male), Vindemiatrix (Female), Sadachbia (Male), Sadaltager (Male), Sulafat (Female), Alnilam (Male), Rasalgethi (Male). The model default is Zephyr (Female). Your workspace may offer a narrower selection depending on your plan.
Resolution of the video. Supported values are 720p, 1080p. The model default is 720p. Your workspace may offer a narrower selection depending on your plan.
Resolution of the video. Supported values are 720p, 1080p. The model default is 720p. Your workspace may offer a narrower selection depending on your plan.
Language/accent target for generated speech. Supported values are English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean, Hindi. The model default is English (US). Your workspace may offer a narrower selection depending on your plan.
Language/accent target for generated speech. Supported values are English (US), English (UK), Spanish, French, German, Italian, Portuguese (Brazil), Japanese, Korean, Hindi. The model default is English (US). Your workspace may offer a narrower selection depending on your plan.
Open your Relam workspace and look for Mera Avatar in the video model selector. Its required inputs are image. Available models, controls and access depend on your workspace and plan.
Open your Relam workspace and look for Mera Avatar in the video model selector. Its required inputs are image. Available models, controls and access depend on your workspace and plan.
Voice Prompt
Optional input
Optional style instructions for how to speak voice_script, such as tone, pacing, accent, or emotion. These instructions are not spoken.
Voice Script
Optional input
Exact words the avatar should say. Required when no audio file is uploaded.
Negative Prompt
Optional input
Disabled if empty.Mention what you do NOT want in the video, e.g. "subtitles, text, blurry, low quality, frames, watermark, titles, scene change". We recommend using multiple keywords at once.
Disable Safety Filter
Optional input
Disable safety filter for prompts and input image. See Skytells's Responsible AI docs.
Strength Negative Prompt
Optional input
Strength of the Negative Prompt. Optimal value can differ for different video lengths (Experimental Feature)
Disable Prompt Upsampling
Optional input
When true, skip automatic enhancement of the visual video prompt and use video_prompt directly.
Output & configuration
The details you control.
These options come from Mera Avatar’s model specification. Your workspace shows the options included with your plan.
Voice
Voice to use when generating speech from voice_script.