Mera 2.0
The flagship Omni model. Build a scene where motion, sound, and emotion stay connected for up to two minutes.
Every frame. Every feeling.
Mera 2.0 is a flagship video generation model by Skytells, bringing believable motion, realism, lip sync, audio, and emotional performance into the same scene.
Video. Audio. A performance that belongs together.
The Mera family
Text to video. Image to video. Audio-driven lip sync, avatars, and reference-guided creation. Mera brings these modes into one Omni family. Mera 2.0 and Mera 2.0 Avatar can generate up to 120 seconds of video, maintaining a consistent subject and emotional continuity with a deep understanding of meaning.
The flagship Omni model. Build a scene where motion, sound, and emotion stay connected for up to two minutes.
The avatar family. Mera 2.0 Avatar carries a consistent subject and emotional performance through videos of up to 120 seconds.
The Recast family extends Mera’s creative possibilities, alongside the flagship and Avatar models.
We chose Classical Arabic to challenge Mera’s pronunciation, phrasing, and emotional understanding. Its precise articulation and sustained vocal passages ask a great deal of even an experienced speaker. A word of love can carry longing, tenderness, or loss. Watch how the face follows that meaning as you listen to the voice.
Physics that gives a scene its grounding. Movement, contact, and momentum are designed to feel connected to the world around them.
Light, texture, and human expression work together. The small details make the whole scene feel more convincing.
Audio and lip sync belong to the same moment. Speech timing, facial delivery, and sound carry the scene together.

Classical Arabic brings precise articulation, breath, and musical phrasing into one performance. Listen for the held notes. Watch how the mouth stays with the sound.
Tenderness, longing, and sorrow can live in the same phrase. The voice carries those shifts; the eyes and face make them visible.
A sustained vowel asks more of a scene than a single spoken syllable. Vocal timing, facial movement, and the continuity of the shot need to stay connected.
Voice understanding
Mera listens to the original audio, understands its emotional meaning, and adapts the face, lip sync, and movement to the voice. Models that receive only a transcript work from written words, without the speaker’s original tone or delivery.
Mera hears the voice. Other models read the words.
Hears the original voice. Understands its emotion.
Mera hears pitch, breath, pauses, and emotion in the original voice. It understands the meaning behind the delivery and adapts lip sync, facial expression, and movement to it.
Receive a transcript instead of hearing the voice.
The same phrase, written down.
A transcript contains the words, but leaves out the sound of the original performance. These models infer emotion from text and generate delivery without hearing the speaker’s tone.
| What shapes the performance | Mera | Other models |
|---|---|---|
| Original voice | Heard directly | Transcribed into words |
| Emotional meaning | Understood through voice and delivery | Inferred from written text |
| Expression & lip sync | Adapt to the speaker’s vocal emotion | Generated from text instructions |
Audio understanding
Mera listens to the voice itself. The rhythm of a sentence. A hesitation. The lift in a laugh. Vocal delivery gives it context for emotion, expression, and lip sync.
A system that reduces speech to plain text keeps the words. Pitch, intensity, pauses, and the texture of the original voice can be left behind. Generating a new voice from that text does not automatically recover the original delivery.
Listening to audio preserves the cues that give the same words a different feeling. Mera uses that vocal context to connect the sound of a performance with the face and movement on screen.
Why delivery matters: research on combining audio and transcripts for emotion recognition.