A free AI voice cloning tool that learns your voice from a few seconds of it and says anything you type, in 10 languages, in your browser, with nothing uploaded.
5.0(2 ratings)Add a few seconds of your voice, type anything, and hear it said in your voice. It all happens in this browser, and your voice never leaves it.
Add your voice to start: record the script, upload a clip, or paste a link.
“Advaita Vedanta philosophy considers Atman as self-existent awareness, limitless, non-dual and same as Brahman.”
“An inter-religious calendar provides significant religious observations for several years.”
“It was destroyed but has now been rebuilt and decorated by Nepali artisans.”
“The duo performs exclusively on analog synthesizers, especially Moog synthesizers.”
“Most species are highly agile, and regularly leap several metres between trees.”
“The antibody also reacts positively against junctional nevus cells and fetal melanocytes.”
“For asymmetric hydrocyanation, popular chiral ligands are chelating aryl diphosphite complexes.”
“It was such a gradual movement that he found it only by noticing the dots.”
“The strip's humor occasionally satirizes modern American culture, and deliberate anachronisms are rampant.”
“Live, in-studio performances by artists are also regularly scheduled.”
“Tablecloths can be made of almost any material, including delicate fabrics like embroidered silk.”
“Motivation theories can be classified broadly into two different perspectives: content and process theories.”
“He consistently attracted large crowds on his travels, some among the largest ever assembled.”
“Less than one third of the world's Armenian population lives in Armenia.”
“Mary told me that she got to meet up with you while she was back in San Francisco.”
“Reality television shows played an important, influential role on the charts during the decade.”
“Nakamura would follow up by opening video arcades featuring Atari games.”
Your voice and your text never leave this browser. The models download from our server once, and nothing you record is uploaded.
This free voice cloning tool learns your voice from a short recording, then says anything you type in that voice. It runs on your own computer's graphics card, in your browser, so your voice never leaves it.
Record yourself reading a short script, upload a clip, or paste a link to a Reel, a TikTok or any video of you talking. It speaks 10 languages, and there's no account, no watermark and no limit.
I built it because the voice cloners I pay for all upload your voice to their servers. Here's how it works, how close it gets to the real thing, and how to use it the right way.
What is an AI voice cloner?
An AI voice cloner makes a digital copy of a voice from a recording of it. You type any text, and the copy reads it out in that voice, with its tone, pace and accent.
Most voice cloners run on a company's servers. This one runs an AI model on your own device, so nothing is uploaded.
How voice cloning works
Every voice cloner I've tested does the same three things, each in its own way.
- It turns sound into speech codes – each second of audio becomes a short run of numbers, like letters for sound, and numbers become sound again at the end.
- A language model for sound writes those codes – the way a chatbot predicts the next word, it predicts the next bit of speech, from a huge amount of recorded speech it learned from.
- It gets your voice from you – from a voice print, from an example of you, or from training on you.
That last step is where voice cloners differ most. This table shows the three ways.
- How it gets your voice
- A short summary of your voice's color, measured from your sample
- What you get
- Close in color, rarely you
- How it gets your voice
- It hears seconds to minutes of you and your words, then keeps talking as if the recording went on
- What you get
- Your tone, accent and pace, and your microphone too
- How it gets your voice
- The model's own settings change to learn you from 30 minutes to 3 hours of recordings
- What you get
- How you say every sound, your rhythm and your habits
This tool and Voicebox use an example of you. ElevenLabs uses an example for its instant clones, and training for its Professional Voice.
That's why this tool asks for your words: the model lines up your sound with them. It writes them down for you on your device, and you can fix any it heard wrong.
How to clone your voice
These are the steps, from a recording to a file you can use.
The first time, the voice model downloads once and your browser keeps it. After that it starts in seconds.
Clone your voice from a link
Already talking in a Reel, a TikTok or a podcast clip? Paste its link instead of recording again.
It reads public posts on Instagram, TikTok, X, Facebook, LinkedIn and Pinterest, and any link to a video or sound file, like an MP4 or an MP3. These are the steps.
It picks the 30 seconds with the most talking. Music under your voice gets copied along with it, so move the window to a part without any.
The check is there because a link can hold anyone's voice. It compares your sentence with the clip on your device, and a voice that doesn't match can't be used.
YouTube links don't work, because YouTube doesn't allow downloads. If it's your video, download it from YouTube Studio and upload the file.
Instant clones vs trained voices: 30 seconds or 3 hours
This tool, like Voicebox, makes an instant clone. The model reads your sample as part of its instructions every time it speaks, and it uses up to 30 seconds of it.
A trained voice is made once, from hours of you, and then speaks without any sample. This table puts the two side by side.
- Instant clone, like this tool
- 10 to 30 seconds here; 1 to 2 minutes at ElevenLabs
- Trained voice, like ElevenLabs Pro
- 30 minutes at the least, 2 to 3 hours for the best
- Instant clone, like this tool
- Seconds
- Trained voice, like ElevenLabs Pro
- Hours on a big graphics card
- Instant clone, like this tool
- Your browser
- Trained voice, like ElevenLabs Pro
- A company's servers, or a powerful computer
- Instant clone, like this tool
- Your voice in one recording
- Trained voice, like ElevenLabs Pro
- How you say every sound
Does a longer sample make a better clone?
Not past 30 seconds. An instant clone learns nothing from extra minutes: every sentence just gets slower to make.
Here's the same laptop recording of me, short and long.
| Sample | Sounds like me |
|---|---|
| 13 seconds | 0.81 on average over 7 takes (0.78 to 0.83) |
| 29 seconds | 0.82 |
A clean sample matters more than a long one: noise, music and other voices get copied too.
Long recordings matter for tools that train a model on you. ElevenLabs asks for 1 to 2 minutes for an instant clone, and 30 minutes or more for its Professional Voice.
ElevenLabs says the same about its own instant clones: 1 to 2 minutes of clear audio is enough, and more than 3 minutes can even make the clone less stable.
Can I give it 30 minutes or more?
Yes, and it works: upload a long recording or paste a link to a long video, and the tool finds the 30 seconds with the most talking in it. You can move them on the waveform to your best-sounding part.
It can't train on all 30 minutes in your browser. Training a voice means hours on a big graphics card, far more than a browser has.
What I learned trying to train my own voice
To see how far training takes it, I tried training the same kind of model on 2 hours of my own solo sessions, on my MacBook Pro with an M1 Max.
It didn't work there. The training ran out of memory next to my other work, and the voice it made spoke nonsense, while the untrained model with my voice print said every word right.
Training is built for big NVIDIA graphics cards, which is why ElevenLabs trains its Pro Voices on its own servers. With a clean 30-second sample, this tool already came level with my Pro Voice, so the sample is where your time pays off.
How close it gets
My benchmark is my ElevenLabs Pro Voice. ElevenLabs trained it on long recordings of me, 30 minutes or more, as it does for every Pro Voice, and it's the voice I use for my own videos.
So I gave this tool 30 seconds of one of my sessions, recorded on a good microphone. Then this tool and my Pro Voice each said a sentence from later in the same session, which this tool never heard.
Each got 3 takes, and I kept each one's best, which is what I'd tell anyone to do. This table has the scores.
- How close to my voice
- 0.81 to 0.86
- Words right
- All
- How close to my voice
- 0.81
- Words right
- All
- How close to my voice
- 0.79
- Words right
- All
- How close to my voice
- -0.05
- Words right
- Not a clone
Each score comes from a speaker recognition model and runs from -1 to 1: the higher, the closer to my real voice saying that sentence. My real voice from other parts of the same session is the ceiling, and a stranger sits near zero.
On this test, 30 seconds came level with hours of training. Every take from both said every word right, and this tool's 3 takes were steadier: 0.78 to 0.81, against 0.69 to 0.79 for ElevenLabs.
A score only goes so far, so play all three in the tool above, my real voice first, and judge by ear.
The sample makes the biggest difference. Cloned from 13 seconds of me reading on my laptop, the same tool scored 0.37 on this test and sounded flat.
Why some voice clones sound better than others
I've cloned my voice with this tool, Voicebox and ElevenLabs, and 17 other people's voices with this tool. These are the things that made the difference.
- The sample decides most of it – from 30 seconds of me presenting on a good microphone, this tool came level with my ElevenLabs Pro Voice; from 13 seconds of me reading on a laptop, it fell far behind.
- Training is steadier – a voice trained on hours of you learns how you say every sound, while an example only shows 30 seconds of it, so takes vary more.
- A bigger model hears more – Studio Max, the bigger of this tool's models, sounded more like me than Studio from the same sample.
- The sample is copied, room and all – my restaurant recording gave me a restaurant clone, and my laptop recording sounds like a laptop.
- The words have to match the sound – a clip cut mid-word made the model guess at that word, and it said the guess before my text.
- The model needs to know the language – left to guess, my clone drifted toward another accent, so this tool always tells it.
- Every take is a little different – the model picks each bit of speech with some chance in it, so make 2 or 3 takes and keep the best.
- Some languages carry your voice better – in my test, my clone came out closest to me in English and furthest in Chinese.
The first one matters most, and it's the one you control: record your 30 seconds the way the next section shows.
Real voices, cloned: 17 accents
My own voice is one test, so I cloned 17 more people: women and men, from England to Hong Kong. Each one gave their voice to Mozilla's Common Voice, a public collection of voices donated for speech technology.
Play any of them in the tool above: the real clip first, then the clone saying the same sentence. The clone never heard that clip; it learned from about 20 seconds of their other recordings.
- Real against real
- 0.81
- Clone against real
- 0.75
- Words right
- 1 heard differently
- Real against real
- 0.83
- Clone against real
- 0.77
- Words right
- All
- Real against real
- 0.86
- Clone against real
- 0.76
- Words right
- All
- Real against real
- 0.77
- Clone against real
- 0.70
- Words right
- All
- Real against real
- 0.79
- Clone against real
- 0.74
- Words right
- All
- Real against real
- 0.82
- Clone against real
- 0.75
- Words right
- All
- Real against real
- 0.83
- Clone against real
- 0.78
- Words right
- All
- Real against real
- 0.84
- Clone against real
- 0.75
- Words right
- All
- Real against real
- 0.86
- Clone against real
- 0.76
- Words right
- All
- Real against real
- 0.83
- Clone against real
- 0.74
- Words right
- 1 heard differently
- Real against real
- 0.86
- Clone against real
- 0.81
- Words right
- All
- Real against real
- 0.78
- Clone against real
- 0.77
- Words right
- All
- Real against real
- 0.81
- Clone against real
- 0.81
- Words right
- All
- Real against real
- 0.81
- Clone against real
- 0.75
- Words right
- All
- Real against real
- 0.80
- Clone against real
- 0.76
- Words right
- All
- Real against real
- 0.82
- Clone against real
- 0.71
- Words right
- 1 heard differently
- Real against real
- 0.66
- Clone against real
- 0.57
- Words right
- All
On average, the clones scored 0.75 against the real clip, where the same person's real voice scores 0.81. Two of them, the Malaysian and the South African voice, scored as high as the real person.
Every clip is 1 take, not the best of several. In 13 of the 17, the transcript matched every word; in the other 4, the speech recognizer heard 1 word differently.
Three voice cloning models
The model menu beside the text box has 3 AI voice cloning models. Studio Max is chosen from the start, because it sounds most like you.
- Best for
- Sounding most like you
- Download, once
- 2.1 GB
- Languages
- 10
- Best for
- A computer with less memory, or a faster first try
- Download, once
- 965 MB
- Languages
- 10
- Best for
- English with a laugh, a sigh or a whisper where you want it
- Download, once
- 649 MB
- Languages
- English
Studio Max and Studio both learn from your sample and its words, and both speak 10 languages: English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian. Turbo works a different way and is English only, but it's the one that laughs and sighs on cue.
If Studio Max won't load on your computer, pick Studio. It needs about half the memory, and in my test it sounded almost as close.
How fast it is
After the one-time download, Studio speaks about as fast as you'd read it out loud: 8 seconds of speech took 8 seconds on my MacBook Pro with an M1 Max. Studio Max takes a little longer, about 10 seconds for 6 to 8 seconds of speech.
The download is the slow part, once. On my connection, the first sentence took 91 seconds with Studio and 117 with Studio Max, most of it the download.
Here's the full list of what it does.
| Spec | This voice cloning tool |
|---|---|
| Price | Free: no account, no watermark, no limit |
| Runs on | Your own graphics card, in the browser (WebGPU) |
| Browsers | Chrome or Edge on a computer; recent phones are next |
| Models | Studio Max, Studio and Turbo |
| Download | 2.1 GB, 965 MB or 649 MB, once, kept by your browser |
| Your sample | 3 to 30 seconds; 10 to 30 is best |
| Sample from | A recording made here, a file, or a link |
| Links | Instagram, TikTok, X, Facebook, LinkedIn and Pinterest posts, or a video or sound file |
| Its words | Written down on your device (a 77 MB model, once), and you can fix them |
| Text | Up to 5,000 characters at a time |
| Languages | 10 with Studio Max and Studio, English with Turbo |
| Speed | Studio about real time, Studio Max a little slower |
| Sound out | 24 kHz mono, at podcast loudness (-16 LUFS) |
| Downloads | WAV or MP3 |
| Privacy | Your voice and your text never leave your browser |
How it compares with other voice cloning tools
Most voice cloning tools run on the company's servers and charge a monthly plan. I read each tool's own pages on September 28, 2026, and this table sums up what they say about cloning.
- Free to clone
- Yes
- Paid plans with cloning
- None needed
- Sample it asks for
- 10 to 30 seconds
- Free to clone
- No
- Paid plans with cloning
- From $6 a month
- Sample it asks for
- 1 to 2 minutes; 30+ for a Pro Voice
- Free to clone
- No
- Paid plans with cloning
- From $29 a month
- Sample it asks for
- 30 seconds; 1 to 3 minutes for a pro clone
- Free to clone
- A short trial
- Paid plans with cloning
- From $16 a month
- Sample it asks for
- A 90-second script
- Free to clone
- No
- Paid plans with cloning
- From $29 a month
- Sample it asks for
- Your last 2 recordings there
- Free to clone
- In its iOS app
- Paid plans with cloning
- Not listed
- Sample it asks for
- Up to 10 seconds
- Free to clone
- Yes
- Paid plans with cloning
- From $11 a month
- Sample it asks for
- Not stated
- Free to clone
- Yes
- Paid plans with cloning
- Free, open source
- Sample it asks for
- A short sample and its words
ElevenLabs is still the one I use for my own videos. Its Pro Voice is trained on 30 minutes or more of your voice, and it only lets you clone your own.
Its free plan has premade voices, which work for some jobs, but it doesn't clone. In most cases I want my own voice, and that starts on a paid plan.
HeyGen, Descript and Riverside add voice cloning to a video or podcast editor. Riverside clones only the account owner, from the recordings you've already made there.
All of them clone on the company's servers except 2: Voicebox, a free desktop app for your own computer, and this tool, in your browser. Neither sends your voice anywhere.
How to get a clone that sounds like you
The model copies everything it hears, so your sample decides most of the result. These tips come from what I measured and from ElevenLabs' own recording guidance.
Your microphone and room
Start with the sound itself, since the clone keeps it.
- Record on the microphone you want to sound like – a clone of a laptop recording sounds like a laptop, and one from a good USB or XLR microphone sounds like that microphone.
- Sit about two fists from the microphone – the distance ElevenLabs gives for its Pro recordings: closer booms, farther picks up the room.
- Use a pop filter – it softens the hard p and b sounds, which the clone would copy too.
- Pick a quiet room with soft things in it – curtains, a sofa or a closet full of clothes cut the echo, and a noisy clip can go through the background noise remover first.
- Keep your level steady – ElevenLabs aims for -23 to -18 dB RMS with peaks under -3 dB, which in plain terms means loud and clear, and never clipping.
How you talk
The clone copies your delivery as closely as your voice.
- Talk the way you want it to sound – it copies your pace and your energy, calm or excited.
- Keep one style through the sample – a whisper next to a shout gives the clone two voices to choose from.
- Speak in full sentences – and end on a pause, which this tool also does for you when it cuts the sample.
Length, words and takes
These last few are about how you use the tool.
- Give it 20 to 30 seconds of just you – with no music and no one else talking.
- Check the words it heard – the clone lines up your sound with those words, so a wrong one makes worse speech.
- Pick Studio Max – the bigger model sounded more like me in every test.
- Tell it the language – or leave it on Match my text, which follows the language you type.
- Make 2 or 3 takes of the same text – every take is a little different, so keep the best.
Who is the voice cloning tool for?
It's for people who talk for a living and would rather not record every line twice.
- Video caption generator – YouTubers and course creators who fix a line of a voice-over in their own voice, then caption the video
- Audiogram generator – podcasters who turn a cloned intro or a quote into a clip for social media
- Speech to text – writers who draft by talking, then hear the script read back in their own voice
- How to make money podcasting – podcasters who record ad reads and intros for sponsors
- How to host a virtual summit – summit hosts who record an intro for every speaker
It isn't for copying someone else's voice: clone your own, or one you have the speaker's permission to use.
What people use a voice clone for
A clone of your own voice saves you from recording the same thing again. These are the uses I see most.
Fix a word
Say the corrected line in your voice instead of recording the take again.
Voice-overs
Type the script and keep your own voice on every video, even on a tired day.
Other languages
Speak Spanish, French or 8 more languages in your own voice.
Drafts you can hear
Listen to a script in your voice before you record it for real.
Famous mornings, in my AI voice
In the tool above, the mornings of Jeff Bezos, Dan Martell, Tony Robbins and Andrew Huberman are read by my clone. Each routine comes from my morning routines page, with its source there.
It's their routines in my voice, never their voices.
| Person | Their morning, in my cloned voice |
|---|---|
| A slow start he calls puttering: coffee, the papers, time with his family, then the gym | |
| Up at 4 without an alarm, water, 10 pages, quiet creative work, then a workout | |
| Water, a cold plunge, then 10 minutes of priming | |
| Outside light within an hour of waking, caffeine after 90 minutes |
Use it the right way
A voice is part of who someone is. These keep a cloned voice honest.
- Clone your own voice – or one you have the speaker's permission to use, in writing.
- Say it's AI – when you publish a cloned voice, tell your listeners.
- Keep words where they belong – use a clone for things its owner would say.
Parody and commentary that are clearly labeled are a different thing. When in doubt, ask the person first.
After you clone your voice
Bring it to podcast loudness with the audio enhancer, or turn a script into speech for a video and add captions with the video caption generator.
For a stock voice instead of yours, the text to speech tool reads in 15 voices.
Made a video with an AI tool, in a voice that isn't yours? The AI voice changer puts it in your voice, line by line, in the same timing.
What not to do with a cloned voice
Cloning a voice is legal, but some uses of a clone aren't. Stay clear of these.
- Cloning a celebrity or a politician – it passes a copy off as them, and some laws forbid it, like Tennessee's ELVIS Act.
- Faking a call or a message from someone – a cloned voice used to fool people is fraud.
- Selling with someone else's voice – an ad in a voice that sounds like a famous person reads as their endorsement.
ElevenLabs goes further and blocks the voices of celebrities outright. This tool asks you to confirm it's your voice, or that you have permission, before it clones.
A voice from a link also has to match yours: you read one sentence out loud, and it checks.
Want a pro voice?
The best pro voice
AI Voice Cloning FAQs
Questions about cloning your voice? Here's what to know.
AI voice cloning makes a digital copy of a voice from a recording of it.
Type any text, and the copy says it in that voice.
Click Add a voice, record the script on this page, upload a clip or paste a link, and tick the box.
Then type what it should say and click Speak it.
At least 3 seconds, and it uses up to 30.
10 to 30 seconds of just you talking, somewhere quiet, works best.
Not past 30 seconds: it doesn't train on your voice, it listens to your sample each time and uses up to 30 seconds of it.
In my test, 13 and 29 seconds of the same recording sounded as much like me.
Yes: upload it or paste a link, and the tool picks the 30 seconds with the most talking.
It can't train on all of it in your browser, since training a voice takes hours on a big graphics card.
Yes: click Add a voice, then Link, and paste the link to a public post.
It also reads X, Facebook, LinkedIn and Pinterest posts, and links to MP4 or MP3 files.
A link can hold anyone's voice, so the tool checks the voice in it is yours.
The check runs on your device and compares your sentence with the clip.
No, YouTube doesn't allow downloads.
If it's your video, download it from YouTube Studio and upload the file.
In my test, my clone scored 0.81 for sounding like me, where my real voice scores 0.81 to 0.86 and a stranger scores about zero.
On 17 more voices with 13 accents, the clones scored 0.75 on average, against 0.81 for the real person.
On my own voice, it came level with my ElevenLabs Pro Voice: 0.81 against 0.79, best of 3 takes each.
That was from 30 seconds of a clean recording, while ElevenLabs trained on hours of me; from a laptop recording, this tool scored far lower.
Studio Max, which is chosen from the start and sounds most like you.
Pick Studio if Studio Max won't load, and Turbo for English with a laugh or a sigh.
The Studio models learn your voice from your sample and its exact words.
It writes them down on your device, and you can fix any word it heard wrong.
10: English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French and Russian.
Leave the menu on Match my text and it follows the language you type.
Yes, with the Turbo model: pick an expression like laugh, sigh or whisper from the menu beside the text box.
It goes where your cursor is.
Only with their permission.
The tool asks you to confirm it's your voice, or that you have the speaker's permission, before it clones.
No, your recording and your text stay in your browser.
Only the models download, once, from our server.
Not yet: it needs a browser with WebGPU and enough memory, which means Chrome or Edge on a computer for now.
A version for phones is next.
Yes, it's free, with no account, no watermark and no limit on how much it says.
There's text to speech, a background noise remover, an audio enhancer and a speech to text tool.
The rest are in free tools.
I did.
I'm Navid Moazzez, and the voice cloning tool is one of my free tools on navid.me.
Read more about me.
Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.
More free tools
The most actionable AI newsletter for founders
Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.
No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.
P.S. Sign up now to get free access to my ultimate AI tools guide for creators.





















