VoiceKit Open VoiceKit
Voice cloning dataset preparation

Build a voice cloning dataset from audio you already have.

Turn long interviews, podcasts and videos into clean, single-speaker WAV clips with aligned transcripts—without manually cutting a timeline.

Why preparation matters

Raw recordings contain the problems voice models learn.

Background music, other speakers, repeated lines and inaccurate transcripts can reduce similarity and consistency.

A complete, inspectable export

Clean WAV clips, aligned transcripts and a ready-to-use ZIP.