What goes in and what comes out

Who it is for

Programming, distribution and versioning teams

Rights holders, studios, networks and streaming services taking their own library into new markets, rather than language vendors pricing a localization project for someone else's catalogue.

What you start with

A locked master and a speaker list

The released cut of each episode, a list of who speaks in it, and written consent for any performer voice you intend to reproduce. VisionStory does not clear rights for you.

What you get

Dubbed, lip-synced editions per market

One reviewed edition per territory in any of 88 languages, each speaker carried on their own voice, ready for your review chain and your own delivery to platforms and broadcasters.

What changes when versioning runs on your own library

Most teams reach this page after one specific failure: a season went out in a new market and something about the audio was wrong in a way the spreadsheet never captured. On the rights-holder side the library is yours, so the recurring question is how to hold a whole slate together across months of releases rather than how to price a single project for a client.

Give every speaker in the scene their own voice

Three people are in the room. The host pushes back, the guest talks over her, the third voice comes in from a monitor half a second later. Run that through a system built for narration and all three come back as one performer doing an impression of a conversation. The words survive. The scene does not, and a viewer in that market hears a version of the show nobody would have signed off on in the original language.

Multi-speaker video dubbing treats the cast as a cast. Each speaker is assigned a distinct voice in each language and keeps it, with lip sync tied to the picture so a reaction still lands on the face that made it. You decide how close each match should sit to the original performer: a warmer read for the host, something plainer for the analyst who only appears in the cold open. VisionStory produces the editions; casting judgement stays with the people who know the show.

Keep episode fourteen sounding like episode one

Nobody complains about episode one. The complaint arrives at episode four, when the German edition comes back and the lead sounds like a different person, because that batch was ordered eight weeks later and whoever handled it made a reasonable choice with no record of the earlier one. Viewers who watch a season in a weekend hear the seam immediately. The team that has to answer for it usually finds out from a support thread rather than from anyone in the pipeline.

Fixing the voice for each recurring speaker turns that into a decision you make once. A cloned voice built from an authorized recording, or a selection from the 1,000+ voices in the library, becomes the entry for that character in that language and is reused when the next batch of episodes goes through. Keep the mapping of speaker to voice to language somewhere your team can read it, because a season that spans a year of releases will outlive whoever set it up.

Move a back catalogue that stalled on cost

There are several hundred hours sitting on a shelf: finished shows, cleared music, rights that still run for years. They are not in the new market for one reason, which is that the quote to voice them came back higher than anyone could justify against a title that already earned out. So the library stays where it is, and every renewal conversation about that region starts from the same short list of titles.

When dubbing runs in-house from the locked masters you already hold, the arithmetic changes and older titles become worth versioning again. Start with the episodes that have the clearest audio and the fewest overlapping speakers, since those come back needing the least review, then work outward. Two things still gate it and neither is technical: your underlying rights have to cover dubbed derivatives in that territory, and any performer voice you reproduce needs consent that reaches that far.

Track the version tree before it tracks you

One episode rarely stays one episode. A scene is trimmed for one broadcaster, a brand is obscured for another, a rating card is added in a third territory, and each of those cuts wants its own audio and its own subtitle file. Six months in, a request for the current French version has three plausible answers and the person who knows which is correct is on leave. Nothing here is exotic; it is version count multiplying faster than anyone planned for.

The discipline that holds is dubbing only from a locked cut, and rebuilding rather than patching when the cut changes. Each territory edition is generated from a named source master, so a trim that arrives late means reissuing the affected editions from the new master instead of splicing the old audio onto it. VisionStory hands back files with the source they came from; the naming, the rights window and the record of which edition is live stay in the media management system your distribution team already runs.

Settle the voice rights before you clone anyone

The request sounds harmless in a production meeting: keep the narrator's voice in Spanish, the audience already knows it. Then someone asks what the original contract actually says about a synthetic reproduction, and the answer is usually that it was written before anybody had to think about it. Meanwhile the performer has a view of their own, an agent has another, and a title that has already been promoted for a date is waiting on both.

Treat the consent as part of the deliverable. Voice cloning at VisionStory requires a recording the speaker authorized you to use, and for a program library that authorization should name the territories, the term, the titles it covers and what happens when the licence lapses. Where consent is missing or unclear, cast a library voice instead and keep the show moving. The technical step takes minutes; the paperwork is the part that decides whether an edition can stay in the market.

How programming and distribution teams use VisionStory

Start from a locked master and a list of who speaks in it, and note at each step what the local distribution lead still has to sign off before an edition reaches a broadcaster, a platform or an audience.

Before any edition is released, the local distribution lead confirms the translated dialogue, the cultural adaptation and the rating and cut requirements for that territory, and confirms that the consent behind every cloned voice covers this region and this term.

  1. 01

    Version the locked master, not a cut still in review

    AI Video Translator takes a source master your team has already locked and produces dubbed, lip-synced editions across 88 languages, with each speaker in the scene carried through as a separate voice. Work from the released cut only. An edition built from a version still in review has to be reissued the moment a trim lands, and by then the file may already be sitting with a broadcaster. Order the markets that have a contracted date first, and keep the source master name attached to every edition it produced.

    Tool AI Video Translator
    AI Video Translator
    Create dubbed, lip-synced language versions from one approved source video.
  2. 02

    Fix the voice each recurring speaker uses

    Voice Cloning builds a reusable voice from a recording the performer authorized, which is what lets the same host sound like the same host in episode one and episode twenty. Before you clone anyone, confirm the consent covers the territories, the term and this use, and store it with the title rather than in someone's inbox. Where you do not have it, pick from the 1,000+ library voices and record the choice against that speaker. Either way, the mapping of speaker to voice to language is what a season is built on.

    Tool Voice Cloning
    Voice Cloning
    Create one reusable voice from an authorized source recording.
  3. 03

    Clear the source audio that fights the dub

    Location sound is rarely clean. Room tone, an air handler, traffic under a window, a lapel mic rubbing on a jacket: all of it sits under the dialogue you are versioning, and it is worst on exactly the older material you want to move first. Remove Noise lowers that bed while keeping the voice intelligible, so the reference read going into a dub is the performance rather than the room. Do this pass on the episodes flagged in review, not on everything, and listen to a passage with two speakers overlapping before you accept the result.

    Feature Remove Noise
    Remove Noise
    Reduce background noise while preserving clear, intelligible speech.
  4. 04

    Repair archive picture without moving a frame

    Catalogue titles carry the damage of whatever they were finished on: soft transfers, compression artefacts from an old delivery, noise in every dark scene. AI Video Enhancer clears blur, noise and compression damage without changing the timeline or the audio, which is the property that matters here, because a dub already synced to the picture stays synced. Check faces in close-up and any on-screen text a territory has to read, and accept that a title beyond repair is one you re-master properly or leave out of the slate.

    Tool AI Video Enhancer
    AI Video Enhancer
    Restore blur, noise and compression damage in existing footage without changing its timing or audio.
  5. 05

    Deliver at the resolution the platform asks for

    A platform deal usually specifies what it will accept, and a standard-definition master from a decade ago does not meet it. AI Video Upscaler produces a larger, cleaner delivery master up to 4K from the version you have restored, so an older title can go out against the same spec as your current slate rather than sitting in a lower tier. Run it last, after the picture repair and after the audio editions are approved, then spot-check motion and skin tones at full size before the file leaves for the platform.

    Tool AI Video Upscaler
    AI Video Upscaler
    Upscale an existing low-resolution video for a sharper delivery master, up to 4K.

Frequently asked questions

  • Each speaker is treated as a separate part rather than folded into one narration track. You assign a distinct voice per speaker per language and keep it across the run, and lip sync ties the dubbed dialogue to the picture so a reaction still lands on the face that made it. Overlapping dialogue is the hardest case in any dubbing workflow, so review those passages first: if a crosstalk moment reads clearly in the dubbed edition, the calmer scenes almost always will too.

Users Love VisionStory

Discover why content creators and marketers trust VisionStory for their AI video needs. From powerful features to an effortless user experience, our community can’t stop raving about the results they achieve with VisionStory.

See all reviews on G2