Wednesday, September 16, 2026
Advertisement
HomeAI Video GenerationWhat Video Specs Actually Matter for Synthesia AI Dubbing

What Video Specs Actually Matter for Synthesia AI Dubbing

0
(0)

Synthesia rolled out a significant update to its AI video dubbing tool on July 15, 2026. Called Dubbing 2.0, the update touches nearly every part of the localization pipeline: lip-sync tracking, voice generation, translation timing, and the editing workflow. For anyone using the tool to translate existing video into other languages, the obvious next question is practical: what do you actually need to do to your source footage to get a good result?

The answer is more limited than many guides suggest. Synthesia does publish a clear technical checklist for what a source video needs to be file format, size, resolution, frame rate, audio sample rate. What it does not publish is a set of filming rules specific to dubbing: no stated lighting requirements, no camera-angle limits, no guidance on how close a face needs to be to the lens. That gap matters, because it’s easy to find filming advice online that sounds authoritative but was written for a different product, or even a different company, entirely.

Here’s what Synthesia has actually confirmed, what changed with Dubbing 2.0, and where the real uncertainty lies.

Advertisement

What Synthesia AI Dubbing Actually Requires From Your Source Video

According to Synthesia’s own documentation, a video uploaded for dubbing needs to meet the following:

  • File format: MP4, WebM, or MOV
  • File size: up to 5GB
  • Duration: up to 2.5 hours
  • Resolution: up to 4K (3840×2160)
  • Frame rate: 23.98 to 60 frames per second
  • Audio sample rate: 8 to 96 kHz

Alternatively, a video can be dubbed directly from a YouTube link rather than an uploaded file, with no stated duration limit through that path.

RequirementSpecification
Accepted formatsMP4, WebM, MOV
Max file size5GB
Max duration2.5 hours
Max resolution4K (3840×2160)
Frame rate range23.98–60 fps
Audio sample rate range8–96 kHz
Alternative sourceYouTube link

That’s the confirmed ceiling on what the system accepts. What’s missing is any documented floor. Synthesia states a maximum resolution of 4K, but nowhere specifies a minimum resolution below which lip-sync or voice cloning quality is expected to suffer. The same is true for lighting, framing, and camera angle — none of these appear in Synthesia’s dubbing documentation at all. Anyone preparing footage for dubbing is working from a known technical envelope, not a known quality threshold.

What Changed With Dubbing 2.0

Synthesia describes Dubbing 2.0 as a rebuilt stack affecting the “entire workflow,” from transcription through final output. According to the company, the update includes:

  • An upgraded lip-sync model designed to track fine mouth movements and hold sync through fast cuts, scene transitions, and scenes with multiple speakers — areas where Synthesia says the previous version tended to drift
  • A revised audio engine intended to produce more natural pacing, more faithful accents, and a wider emotional range in the cloned voice
  • A translation engine that Synthesia says better respects the original video’s timing, reducing how often a translated line runs long and forces a segment to stretch
  • Support for glossary terms (such as product names) inside dubbing, so terminology stays consistent across languages
  • An editing workflow that lets users correct a single transcript segment and regenerate just that portion, rather than reprocessing the whole video, without using additional credits

These are Synthesia’s own characterizations of the update, published on its company blog. No independent, third-party benchmark of Dubbing 2.0’s actual lip-sync accuracy or voice quality — measured against the prior version or against competing tools — was identified. That doesn’t mean the claims are false; it means they haven’t yet been externally verified, and readers should treat them as the company’s assessment of its own product rather than an established, measured result.

Lip Sync vs. Voice-Only Dubbing: Cost and Plan Differences

One concrete, documented factor worth planning around is cost. Synthesia charges differently depending on whether lip sync is enabled:

OptionCredits per minute
Dubbing with lip sync240
Dubbing without lip sync (voice only)120

Lip sync effectively doubles the credit cost of a dubbing job. It’s also worth noting that lip sync isn’t available at all on Synthesia’s Basic-tier plan — it requires a higher plan tier. Teams budgeting a localization project, especially at scale, will want to factor both the cost multiplier and the plan restriction into their planning before committing footage to the pipeline.

What Synthesia Does Not (Yet) Document About Filming Conditions

This is the part of the picture that’s easy to get wrong, because the internet is full of confident-sounding advice about how to film for AI dubbing that doesn’t actually come from Synthesia.

As of now, Synthesia has not published:

  • A minimum resolution requirement for dubbing
  • Lighting or contrast guidelines specific to the dubbing feature
  • A maximum head-tilt or facial angle for reliable lip tracking
  • Any stated “padding” or minimum face size within the frame
  • A recommended camera-to-subject distance

There is also no page on Synthesia’s site specifically titled something like “Source Quality Requirements” for dubbing. The closest thing that exists is the technical spec table above.

It’s worth being explicit about a distinction that’s easy to blur. Synthesia does publish filming guidance elsewhere on its site — but it’s written for a different feature entirely: creating a Personal Avatar or a Studio Express-1 Avatar, where a person’s face and voice are used to build a reusable AI avatar. That guidance includes advice like ensuring even lighting, keeping both eyes visible, and avoiding busy backgrounds. It’s genuinely useful advice, but it was written for the process of building an avatar from scratch, not for translating the audio of an existing video. Applying it to dubbing is an assumption, not a documented requirement.

It’s also common to find similarly specific-sounding filming tips published by other AI video companies — DeepBrain AI’s “AI Studios,” Synthesys AI, Dubly.AI, Sync Labs, and ElevenLabs among them. Given how closely some of these names resemble Synthesia’s, it’s easy to see how their guidance could get misattributed. None of it describes how Synthesia’s dubbing tool behaves.

Bulk Dubbing and API Options

For teams with more than one video to localize, Synthesia’s in-app Dubbing page supports uploading and processing up to 10 videos in a single batch, with shared settings applied across the group. Larger volumes are handled through Synthesia’s Dubbing API or its Excel Add-in, which allow programmatic or spreadsheet-driven dubbing at scale. Neither route introduces additional footage-quality requirements beyond the specs already listed above.

One policy detail is worth flagging separately from the technical requirements, because it’s often confused with them. Synthesia states that dubbing a video does not require consent from the person appearing in it, because the process preserves the speaker’s original appearance and intent — unlike building an avatar or cloning a voice for reuse, which does involve consent requirements. That said, the person uploading the video is still responsible for holding the appropriate rights and permissions to the underlying content itself. This is a legal and content-rights point, not a footage-quality one, but it belongs on any pre-upload checklist.

Conclusion

The picture that emerges from Synthesia’s own documentation is narrower than a “how to film for AI dubbing” headline might promise. There is a real, usable technical checklist: stick to MP4, WebM, or MOV; stay under 5GB or 2.5 hours; keep resolution at or below 4K; keep frame rate between roughly 24 and 60fps; keep audio sample rate between 8 and 96kHz. There’s also a real cost consideration — lip sync costs twice as much as voice-only dubbing and isn’t available on the entry-level plan.

What isn’t yet available is Synthesia-specific guidance on the filming conditions that most affect perceived quality — lighting, angle, framing, distance from camera. Until Synthesia publishes that, the most honest approach is to follow the documented technical limits, apply ordinary video-production judgment (clear audio, a visible and unobstructed face, minimal camera shake), and treat any more specific numeric claims about lighting or angle tolerance — wherever they’re found — with skepticism unless they’re sourced directly to Synthesia’s own materials. Whether Synthesia adds this kind of guidance as Dubbing 2.0 matures is one of the things worth watching.

FAQ

What file formats does Synthesia AI Dubbing accept?
MP4, WebM, and MOV files are supported, along with dubbing directly from a YouTube link.

What’s the maximum video length or file size for Synthesia AI Dubbing?
Uploaded files can be up to 5GB or up to 2.5 hours long. YouTube-link dubbing has no stated duration limit.

Does Synthesia AI Dubbing support 4K video?
Yes, Synthesia’s documentation lists support for resolutions up to 4K (3840×2160). No minimum resolution is specified.

Do I need special lighting or camera angles for accurate lip-sync in Synthesia?
Synthesia hasn’t published lighting, framing, or camera-angle requirements specific to its dubbing feature. Guidance of that kind that does exist on Synthesia’s site applies to building Personal Avatars or Studio Express-1 Avatars, a different feature.

How much more does lip-sync dubbing cost compared to voice-only dubbing?
Lip-sync dubbing costs 240 credits per minute, versus 120 credits per minute without lip sync — double the cost. Lip sync also isn’t available on Synthesia’s Basic plan.

What changed in Synthesia’s Dubbing 2.0 update?
Released July 15, 2026, Synthesia says the update improves lip-sync stability through fast cuts and multi-speaker scenes, produces more natural-sounding voice output, better preserves original video timing during translation, adds glossary support, and allows single-segment editing without extra credit cost. These are company claims and haven’t been independently benchmarked.

Do I need consent from the person in the video to dub it with Synthesia?
No — Synthesia states that dubbing doesn’t require the subject’s consent, since it preserves their original appearance and intent. The uploader is still responsible for holding proper rights to the source content.

Was this article helpful?

Rate this article from 1 to 5 stars.

Average rating: 0 / 5. Reader ratings: 0

No ratings yet. Be the first to rate this article.

Thank you for your feedback

Help us improve this article.

What information was missing or could be improved?

RELATED ARTICLES
Advertisement

Most Popular