# Voice Cloning APIs in 2026: A Complete Guide to Speaker Similarity, Consent Verification, and Pricing
The voice cloning landscape has evolved dramatically in 2026, with providers competing on audio quality, language support, consent safeguards, and cost efficiency. Whether you’re building a voice agent, localizing content, or creating brand-specific speech, choosing the right API requires weighing multiple factors at once. This guide breaks down the leading options and helps you match each service to your use case.
## Quick Comparison at a Glance
| Provider | Reference Audio Needed | Consent Verification | Commercial License Starting At | Languages Supported | List Price per 1M Characters |
|—|—|—|—|—|—|
| ElevenLabs | 1–2 min recommended / 30 min minimum | Rights attestation (IVC); Voice Captcha (PVC) | Starter, $6/mo | 70+ (v3) | $50 (Flash) to $100 (v3) |
| Cartesia | 10 s / up to 60 s on Sonic 3.6+ / 30 min | Permission required by terms | Pro, $5/mo | 44 | ~$37 to $50 |
| Inworld | 3–30 s / 10 min minimum (beta, English only) | Rights confirmation | Free On-Demand tier | 200+ languages and locales | $12.50 to $25 (TTS-2) |
| Gradium | 10 s / 30 min, 2 h recommended | Owner consent required by policy | XS, $13/mo | 5 | ~$36 to $58 |
| Fish Audio | ~10 s / 10 to 180 min | Live ownership check for pro clones | Plus, $15/mo | 83 | $15 per 1M UTF-8 bytes |
| Resemble AI | 10 s / 10 to 25+ min | Verifiable consent for Professional Clone | Business plan for cloning API | 23 | ~$30 (estimate) |
| Hume | 15 s / not documented | Upload from a consenting speaker | Creator, $14/mo | 11 | $50 to $150 by plan |
Note: Credit-plan rates are calculated by dividing the monthly plan price by the included characters. Resemble AI’s pricing page no longer lists TTS rates; tracker estimates place it at roughly $0.0005 per second as of mid-2026, converted here at approximately 1,000 characters per minute. All prices were verified in September 2026.
—
## Speaker Similarity: What the Data Reveals
One of the most meaningful ways to evaluate voice cloning quality is through blind testing. In September 2026, Hume published its Voice Replication Leaderboard, which assessed 11 voice cloning models using 25 reference voices and 7 different prompts. Three independent raters scored each generated clip on a scale of 1 to 5, judging how closely the output sounded like the original speaker.
Among the vendors covered here, Fish Audio’s s2-pro model achieved the highest score at 4.03. Cartesia’s Sonic 3.5 followed at 3.70, and ElevenLabs’ Multilingual v2 came in at 3.68. Cartesia’s Sonic 3.6-beta scored 3.63, and Inworld’s TTS-2 landed at 3.62. Notably, ElevenLabs’ newest v3 model ranked last among all 11 tested at 2.91.
However, a single similarity score doesn’t tell the full story. Sonic 3.6-beta actually scored highest on naturalness at 4.36, yet placed 8th for speaker identity. Inworld’s TTS-2 earned the top score for audio quality at 4.61. These nuances highlight the importance of looking at category-specific metrics rather than relying on one aggregate number. Full per-category breakdowns are available on the public leaderboard.
—
## Provider Breakdown
### ElevenLabs
ElevenLabs remains one of the most widely recognized names in voice cloning. It offers two cloning tiers: Instant Voice Cloning (IVC) and Professional Voice Cloning (PVC). For IVC, the documentation recommends providing 1 to 2 minutes of clean, high-quality audio. PVC requires significantly more — at least 30 minutes, and ideally 2 to 3 hours of training data, available on the Creator plan or higher.
ElevenLabs also implements one of the strictest consent verification processes in the industry. For PVC, the voice owner must read on-screen text aloud through a Voice Captcha system. IVC relies on a rights attestation process. API pricing is structured per character: $0.10 per 1,000 characters on Eleven v3 and Multilingual v2, while Flash, Turbo, and v3 Conversational modes cost $0.05 per 1,000 characters. The v3 model supports over 70 languages.
### Cartesia
Cartesia distinguishes itself with remarkably low audio requirements for instant cloning — just 10 seconds of input. Its Sonic 3.6 and newer models can leverage up to 60 seconds of reference audio to better preserve accents and vocal nuances. For professional-grade clones, Sonic requires 30 minutes or more of training data, and cloning begins at the $49 Startup plan.
Pricing is credit-based: Sonic charges 1 credit per character. Plans range from $5 for 100,000 credits up to $299 for 8 million credits. Sonic supports 44 languages, making it a strong option for projects that need broad language coverage with low latency.
### Inworld
Inworld offers one of the most accessible entry points into voice cloning. Instant clones work with as little as 3 seconds of audio, and providing samples of up to 30 seconds improves similarity significantly. The cloning process itself is free to use.
For real-time text-to-speech, TTS-2 starts at $25 per 1 million characters on demand, but the rate drops to $15 at a $300 monthly spend and further decreases to $12.50 at the $1,500 tier. A TTS-2 Flash tier begins at $15 and can fall to $7 at higher volume commitments. Professional cloning is currently in beta, requires 10 minutes of audio, and is limited to English. Despite the English-only beta for professional cloning, Inworld supports over 200 languages and locales overall, with 15 in its top quality tier.
### Gradium
Founded by researchers from Kyutai, Gradium offers clones starting from 10 seconds of audio. The free tier allows up to 5 clones for non-commercial use. Paid plans begin at $13 per month for 225,000 credits, billed at 1 credit per character. Professional Voice Clone requires a minimum of 30 minutes of audio and starts on the $340 monthly plan.
Gradium’s coverage is limited compared to competitors, supporting only five languages: English, French, German, Spanish, and Portuguese. This makes it most suitable for European-language projects, particularly those that want a free tier for experimentation.
### Fish Audio
Fish Audio charges a flat rate of $15 per 1 million UTF-8 bytes. English text, which averages roughly 1 byte per character, keeps costs predictable. However, languages like Chinese, Japanese, and Korean average approximately 3 bytes per character, effectively tripling the cost. The S2.1 Pro model covers 83 languages.
Instant clones work from approximately 10 seconds of audio, while professional clones require between 10 and 180 minutes of training data and mandate live ownership verification. Commercial use begins on the $15 monthly Plus plan, which covers verified voices that you own.
### Resemble AI
Resemble AI’s Rapid Clone feature can process 10 seconds of audio and deliver results in under a minute. The Professional Clone option requires 10 to 25 minutes of audio along with approximately 40 minutes of training time. It enforces explicit, verifiable consent from the voice talent — a significant differentiator in an industry where deepfake concerns are growing. Every output from Resemble includes PerTh watermarking for provenance tracking. The Chatterbox Multilingual model covers 23 languages.
An important note: Resemble AI has shifted its focus toward deepfake detection, and its Voice Cloning API now requires a Business plan or higher. The pricing page no longer lists TTS rates publicly, making cost estimation more difficult for prospective users.
### Hume
Hume’s Octave model can clone from as little as 15 seconds of audio, whether recorded live or uploaded from a consenting speaker. Octave 2 supports 11 languages. Commercial use starts at $14 per month on the Creator plan; Free and Starter tiers are limited to non-commercial use.
Hume’s pricing structure charges extra characters at $0.15 per 1,000 on the Creator tier, decreasing to $0.05 per 1,000 on the Business plan. It’s worth noting that Hume’s API access for voice cloning is listed only on the Enterprise tier; other plans create clones within the platform itself.
—
## Which API Fits Which Job
**Budget voice agents at scale:** Inworld and Fish Audio offer the most competitive pricing at high volumes, with per-unit costs that drop significantly on committed plans.
**Low latency with broad language coverage:** Cartesia leads here, combining fast response times with support for 44 languages and a straightforward credit-based pricing model.
**Brand voices with strict consent requirements:** ElevenLabs’ Professional Voice Cloning provides the most robust consent verification through its Voice Captcha system, making it ideal for enterprise brand deployments.
**European languages with a free test tier:** Gradium covers five major European languages and offers a free tier with up to 5 non-commercial clones.
**Cloning with watermarking and deepfake detection:** Resemble AI stands out by embedding PerTh watermarking on all outputs and prioritizing deepfake detection capabilities.
**Emotion-directed delivery:** Hume’s platform is designed to convey nuanced emotional expression, making it a strong choice for applications where tone and feeling matter as much as accuracy.
—
## Frequently Asked Questions
**Q: What is the best voice cloning API in 2026?**
A: There is no single winner across all dimensions. Based on Hume’s September 2026 leaderboard, Fish Audio achieved the highest speaker similarity among the vendors reviewed here. For consent safeguards, ElevenLabs offers the strongest verification process for professional clones. The best choice depends on your specific priorities — whether that’s cost, language coverage, latency, or ethical safeguards.
**Q: How much audio does voice cloning typically require?**
A: Instant cloning on most platforms needs between 3 and 30 seconds of reference audio. Professional cloning tiers generally require 10 to 30 minutes minimum, and many providers recommend 2 hours of training data for the best results. The required duration varies by platform and the quality threshold you’re targeting.
**Q: Which voice cloning APIs verify consent?**
A: Three providers document technical consent checks for professional clones: ElevenLabs (Voice Captcha), Fish Audio (live ownership verification), and Resemble AI (verifiable consent process). The remaining providers rely on attestation, terms of service agreements, or upload-side consent declarations rather than automated verification.
**Q: What is the most affordable voice cloning API per 1 million characters?**
A: Among self-serve list prices, Inworld’s TTS-2 Flash tier drops to $7 per 1 million characters at higher commitment levels, making it the cheapest option for high-volume usage. Fish Audio charges a flat $15 per 1 million UTF-8 bytes for its S2.1 Pro model, though keep in mind that CJK languages can triple the effective cost due to higher byte counts per character.
**Q: Can I use these APIs for non-commercial projects?**
A: Several providers offer free or non-commercial tiers. Inworld’s Free On-Demand tier and Gradium’s free tier (up to 5 clones) are explicitly for non-commercial use. Hume’s Free and Starter plans also restrict commercial usage. Always review the specific terms for your chosen provider before deploying.
—
## Conclusion
The 2026 voice cloning API market offers a diverse range of options, from budget-friendly services suitable for high-volume deployments to enterprise-grade platforms with rigorous consent verification and deepfake detection. Speaker similarity scores have become more transparent thanks to public leaderboards, enabling more informed comparisons than ever before.
Key trends worth watching include the narrowing gap between low-cost and high-fidelity models, the growing importance of consent verification and watermarking as industry standards, and the continued expansion of language support across platforms. As the technology matures, the decision framework is shifting from “which model sounds best” to “which model best fits your specific combination of cost, compliance, and scalability needs.”
Carefully evaluate your project requirements against the data above, test with free tiers where available, and remember that the best API is the one that aligns with both your technical goals and your ethical standards.
Thank you for reading



