Voice cloning in post-production: where it helps, where it cheats

Three months ago, dubbing a 20-minute interview with two speakers through premium-quality AI voice cloning without a cloud solution was nearly impossible. And even with a cloud solution, the result could be very uncertain. Today the technology is moving extremely fast and the question is no longer “does it work?”, but “where does it help, where does it cheat”.

The Tikehau Capital case: two speakers, an international audience

Tikehau Capital commissioned a 20-minute interview between two executives. Target audience: their investors, a significant share of them outside France. The obvious solution on paper: subtitles. Except that watching a 20-minute interview without hearing the executives’ real voices asks an English speaker for a serious cognitive effort.

The brief was never to pretend the executives speak perfect English. On the contrary: the AI-dubbed version is clearly disclosed, in the credits and in the description. The goal is simple. Allow the foreign audience to consume the film like a podcast: on headphones, in the car, while walking. Even easier than with subtitles.

Why we no longer use ElevenLabs as a first choice

As recently as three months ago, ElevenLabs was about the only solution able to produce a clean French-language voice clone at professional quality. The audio quality out of their API was above the market. But it came with a security wall.

To clone a voice on ElevenLabs in professional mode, you had to:

  • Bring the person physically into a studio
  • Record them live in a controlled environment
  • Capture between 30 minutes and 2 hours of clean voice without interruption
  • Pass the platform’s biometric validation

Impossible to send a recording remotely. Impossible to clone from an existing file. Logical on the security side, tedious on the client workflow side.

Today, that security can be bypassed

Open source models that can be fine-tuned locally now make it possible to clone a voice from an existing recording: a previous interview, a public talk, footage from a shoot. The ElevenLabs wall can be worked around.

We understand that this can be frightening. It is legitimate. But in practice, it makes the work far easier. An executive based abroad no longer has to block half a day to come and record 30 minutes of voice in our studio. That logistical barrier used to kill entire projects. This technique is becoming industrialised, and with it the ability to reach a wider audience, more easily.

How much source recording for which result

The constraint that remains: the quality of the clone depends directly on the source material. Concretely:

Source recording duration Clone quality Recommended use
10 minutes Acceptable but not premium. Some passages fall short, certain intonations sound less natural, and significant touch-up work is needed to rescue the result Internal communication, short demos, low-stakes content
30 minutes and more Premium, a clone usable in cinema-grade production Long interviews, corporate films, brand content series

To get a truly solid clone, aim for 30 minutes of good recording per speaker. At 10 minutes, the result remains usable but you can feel it is not the voice in all its finesse. For premium work, it is not enough.

The French accent: a bug turned signature

A less expected advantage of these tools: you can generate a great many languages while keeping the French accent. For authenticity, that is precious. A French executive addressing an Anglo-Saxon investor in perfectly neutral English sounds fake. With a light French accent, he sounds authentic. It really is him speaking.

Modern models know how to modulate that parameter. On Tikehau, that dosage is what made the client say “it really sounds like them”.

Properly tuned, local solutions rival the cloud

The reflex is to assume the cloud always offers better. On voice, that has become false. Cloud solutions remain good, sometimes even better, but they are more limiting: restrictive security, dependence on a third party, a voice that leaves your infrastructure.

Properly tuned, local models produce equivalent, even superior quality. The trade-off: you have to know how to tune. Temperature, seed, context length, silence handling. Badly tuned, a local model produces mush. Well tuned, it renders a voice at 95% of natural on 90% of sentences.

On the client side, one thing must be understood: the more premium the target quality, the more time it takes. That means redoing the AI recordings again and again until the quality is very good. With well-tuned local solutions, you get there. It is goldsmith’s work, sentence by sentence.

And honestly, I cannot imagine what it will look like in a week. The current pace is weekly.

How long it really takes

For Tikehau, two speakers, complete English dubbed version: one week of work. Not so much to generate the audio. A few hours per voice are enough once the clone is tuned. But closing the loop on premium quality takes longer.

Over 20 minutes of speech, that represents between 200 and 300 sentences to validate one by one. Sometimes five attempts on the same sentence before the right intonation.

That time can be reduced if a large recording of each voice is available from the start, at least 30 minutes per speaker, to build a solid clone. The phase that consumes the project is the hunt for the last 2% of imperfections. Not the generation itself.

Dubbing actors vs voice cloning: not the same craft

The ethical debate exists. It deserves to be laid out honestly.

The advantages of cloning are real: it is fast, it reaches an audience sooner, it costs less. But it is not exactly the same work as a dubbing actor’s.

A professional dubbing actor does not really imitate the voice of the person they dub. They offer an interpretation. It is an actor’s craft, not a copy. AI voice cloning does the opposite: you get the person’s actual voice, their inflections, their breathing. Two different techniques answering two different needs:

  • Human dubbing: fiction, animation, content where embodiment matters more than identification (films, series, video games)
  • AI voice cloning: corporate, interviews, speeches where it is the person themselves who must be heard, in a language they do not master

What we do, what we refuse

Our internal rule on voice cloning:

Yes: allowing a brand, an executive, an expert to reach a non-French-speaking audience without losing their vocal identity. With an explicit “AI dubbing” disclosure in the credits.

No: making someone say something they never said. Cloning a voice without written consent. Hiding the AI disclosure in unreadable legal fine print.

The line is not in the technique. It is in the transparency towards the final audience.

And in three months?

Honestly, we do not know. The pace of voice models has become weekly. What took a week for Tikehau a month ago will soon take three days. The current constraints, 30 minutes of source recording, fine tuning, manual sentence-by-sentence validation, are melting away.

What does not change is the question of meaning. Why clone this voice? For whom? Under what disclosure? That grid is what decides whether the tool helps or cheats. Not the technology itself.

Frequently asked questions

How much voice recording is needed for a professional-quality clone?

With 10 minutes of clean recording you get a usable clone, but clearly not a premium one. A few passages sound less natural and you have to compensate in post. For solid professional quality, aim for at least 30 minutes per voice. The longer and cleaner the source, the less post-production time is needed.

Should you always disclose that a voice was cloned by AI?

Yes. At SIGNS, it is a non-negotiable condition. The disclosure appears in the film’s credits and in the distribution description. The goal is never to make people believe the person speaks the target language. It is to make content accessible to an audience that could not consume it otherwise, without deception.

Local solutions or cloud solutions: which to choose?

Properly tuned, local solutions now rival the cloud, sometimes surpass it, and guarantee the voice never leaves our infrastructure. Cloud solutions remain good, sometimes even better, but more limiting on security and workflow. The local trade-off: you must know how to tune temperature, context and silence handling.

Does voice cloning replace human dubbing actors?

No, they are two different crafts. A dubbing actor does not really imitate the voice of the person they dub: they offer an interpretation. With voice cloning you get the exact voice. Cloning serves corporate work, interviews, the speech of an executive who does not speak the target language. Human dubbing remains irreplaceable for fiction, animation and any content where embodiment matters more than identification.

How long does a 20-minute video dubbed by voice cloning take?

For the Tikehau Capital project, two speakers, complete English version: about one week of work. Raw generation takes a few hours per voice; the rest goes into sentence-by-sentence quality control, typically 200 to 300 validations over 20 minutes of speech. With 30 minutes of source recording per speaker available from the start, that time can be reduced.

M
Maxime Vaux
Director & Head of Production, Signs
SHARE
S I G N S.

Got an audiovisual project in mind?
Let's talk it through.

CONTACT US
PREVIOUS Ce que chaque ligne de votre… NEXT Comment intégrer 30 % d'IA dans…