Skip to main content
Visual DubbingExplainer· 4 min read· in Technology

Amazon Prime Video Deploys AI to Alter Actors' Lips for Dubbed Shows

Prime Video has launched a generative AI feature that modifies on-screen lip movements to match translated audio, debuting the technology on the German series Maxton Hall.

By Elena Castillo

Streaming Executives 45%Localization Professionals 30%Original Cast Members 25%
Streaming Executives
Prioritize seamless viewing experiences to maximize the global reach of local content.
Localization Professionals
Advocate for human artistry in translation and voice acting against automated efficiency.
Original Cast Members
Value the authenticity of the unedited physical performance over phonetic synchronization.

Perspectives this story doesn't cover

  • Deaf and hard-of-hearing viewers who rely on accurate lip-reading
  • Labor unions representing traditional dubbing actors

Key terms

Visual Dubbing
The process of using artificial intelligence to alter the video of an actor's face so their lip movements match a translated audio track.
Neural Radiance Fields (NeRFs)
A machine learning technique that reconstructs complex 3D scenes from 2D images, used here to map the actor's face.
Generative Adversarial Networks (GANs)
An AI framework where two neural networks contest with each other to generate highly realistic synthetic images or video.
Synchronization Loss
A training metric in AI models that penalizes the system when the generated lip movements fail to match the exact timing of the audio.

Key points

  1. Amazon Prime Video has launched a generative AI feature that modifies actors' lip movements to match dubbed audio.
  2. The technology is debuting exclusively on the English dub of the hit German romance series Maxton Hall.
  3. The system uses Neural Radiance Fields and GANs to reconstruct the lower half of the actor's face.
  4. While it improves lip-syncing, critics warn it alters the original physical performance and emotional nuance.
  5. AI localization is projected to reduce dubbing costs by up to 90% across the streaming industry.

The binding constraint of video localization is anatomical: human mouths form different shapes to produce different languages. A German "nein" and an English "no" require entirely different facial muscle movements. Historically, dubbing studios have had to choose between translating a script accurately and matching the actor's lip movements. Amazon is now attempting to bypass that constraint entirely by altering the video to match the audio.

On Wednesday, September 9, 2026, Amazon Prime Video launched a generative AI feature that modifies the on-screen lip movements of actors to sync with human-dubbed audio tracks. The technology is debuting exclusively with the English dub of the hit German romance series Maxton Hall.[1][2][6]

The rollout represents a shift from audio-only AI translation to direct visual manipulation. While Prime Video announced plans to expand the capability to "additional titles" in the future, the current deployment is strictly limited to Maxton Hall, a show that ranks as the platform's most-watched international original series to date.[1]

To understand how the system works, it is necessary to look past the broad "AI" marketing umbrella. The underlying framework relies on Neural Radiance Fields (NeRFs) and Generative Adversarial Networks (GANs). Instead of generating a completely new face, the system isolates the lower portion of the actor's face and reconstructs it frame by frame.[5]

How generative AI isolates and reconstructs the lower face to match new audio.

When the dubbed English audio requires an "O" mouth shape, the AI generates the corresponding lip movement and blends it back into the original footage. Amazon's researchers previously detailed a visual speech-to-speech translation (AVS2S) framework that integrates synchronization loss into the training process, ensuring the generated lips match the exact millisecond timing of the new audio track.[3][5]

When the dubbed English audio requires an "O" mouth shape, the AI generates the corresponding lip movement and blends it back into the original footage.

This visual modification is the second phase of Amazon's broader localization overhaul. In March 2025, the company launched an AI-aided audio dubbing pilot for 12 licensed titles, including El Cid: La Leyenda. That earlier system used a hybrid approach where AI generated translated audio and human localization professionals reviewed it for cultural accuracy.[3]

"AI-aided dubbing is only available on titles that do not have dubbing support, and we are eager to explore a new way to make series and movies more accessible and enjoyable," Raf Soltanovich, Vice President of Technology at Prime Video, stated during the 2025 pilot launch.[3]

The new visual dubbing feature flips that model: it pairs human voice actors with AI-generated video. The English audio track for Maxton Hall is recorded by human dubbing artists, and the AI's only job is to warp the original German actor's face to match the English syllables.[1][6]

The technology introduces a strict trade-off between accessibility and performance authenticity. While it solves the cognitive dissonance of mismatched lips, it fundamentally alters the original actor's physical performance. The lower face is critical for conveying subtle emotions—a smirk, a quivering lip, a clenched jaw—which the algorithm may overwrite in its pursuit of phonetic synchronization.[4]

Damian Hardung, the lead actor in Maxton Hall, highlighted the existing challenges of localization in a December 2025 interview. Noting that voice actors often have only two days to a week to dub an entire season, he explained the difficulty of matching the original performance. "To get into that emotional state where it actually feels true, what you're saying, it's almost impossible in a dubbing studio," Hardung said. Replacing the actor's actual facial muscles with a GAN-generated approximation adds another layer of separation between the set and the audience.[4]

AI localization technologies drastically reduce the time and cost required to dub international series.

Despite these artistic concerns, the economic incentives for streaming platforms are massive. Industry data indicates that AI localization technologies are already reducing dubbing costs by 70% to 90% and shortening turnaround times from weeks to days. By making international content feel native to English-speaking audiences, platforms can amortize the cost of local productions across a global subscriber base of hundreds of millions.[5]

The success of the Maxton Hall deployment will test whether audiences prioritize seamless lip-syncing over unedited performances. If viewers accept the modified footage without entering the uncanny valley, visual dubbing will likely become the default standard for international streaming. For now, the technology remains a highly controlled experiment on a single show.[1]

Sources

Source coverage

6 outlets

3 viewpoints surfaced

Streaming Executives 45%Localization Professionals 30%Original Cast Members 25%
  1. [1]The VergeStreaming Executives

    Amazon Prime Video’s new AI tech matches lips to dubbed audio

    Read on The Verge
  2. [2]Pressbee

    Amazon's Prime Video is promising to make English-language dubs look more natural

    Read on Pressbee
  3. [3]SlatorStreaming Executives

    Amazon Prime Video Launches AI Dubbing Pilot

    Read on Slator
  4. [4]CinemaBlendOriginal Cast Members

    Why Maxton Hall's Damian Hardung Doesn't Dub Himself In English

    Read on CinemaBlend
  5. [5]PitchAvatarLocalization Professionals

    Visual Dubbing: Matching Lips to Audio

    Read on PitchAvatar
  6. [6]AmazonStreaming Executives

    Visual dubbing using AI and VFX technologies enhances the viewing experience

    Read on Amazon

Comments

Stay informed

Every angle. Every day.

Get Technology stories with full source coverage and perspective breakdowns delivered to your inbox.