🔒 Private by design — run it in your own VPC

Voice AI you can run inside your own walls

Studio-quality text-to-speech in 30 languages, with voice design and consent-based cloning — self-hosted in your own VPC, so your audio never leaves your network.

Built on the open-source, Apache-2.0 VoxCPM engine — no lock-in. Sign in to explore the interface; a paid plan is required to deploy.
Designed for regulated teams: 🏥 Healthcare🏦 Financial services🏛️ Government ⚖️ Legal📞 Contact centres
Capabilities

Everything the cloud players do — without sending your data away

One engine, three ways to make a voice, thirty languages.

🎨

Voice Design

Describe a voice in plain words — “warm, middle-aged, calm” — and get a brand-new voice. No recording needed.

🎛️

Controllable Cloning

Clone any voice from a short clip, with consent built in. Steer emotion, pace and style while keeping the timbre.

🌍

30 Languages

From English and Mandarin to Hindi, Arabic and Swahili — plus dialects. Just type; no language tag required.

🔊

48kHz Studio Audio

Crisp, broadcast-grade output with built-in super-resolution. Ready for video, IVR and audiobooks.

🔌

OpenAI-compatible API

Point your existing client at one base URL — a drop-in /v1/audio/speech endpoint to migrate off your current provider.

The difference

Your audio is sensitive. Keep it that way.

Most voice AI is cloud-only — every recording and transcript leaves your perimeter. Vocala lets you run the whole engine in your own VPC.

Cloud-only voice AI

  • Your audio + transcripts sent to a third party
  • Off-limits for many healthcare / finance / gov teams
  • Proprietary lock-in, opaque pricing
  • Data residency is whatever they decide

Vocala self-host edition

  • Engine runs inside your network — audio never leaves
  • Meets data-residency & compliance requirements
  • Open Apache-2.0 engine — no lock-in, commercial-ready
  • Production targets blocked by a built-in safety guard
  • Per-tenant usage metering · signed provenance + watermark
Talk to us about self-hosting →
Trust & governance

Every clip is disclosed, traceable, and tamper-evident

The controls regulated teams need — built in, not bolted on. This is what sets Vocala apart from a model with an API.

🔏

Signed provenance

Every generation returns an Ed25519-signed manifest — discloses it's AI, binds to the audio bytes, and is verifiable by anyone with the public key. No trust in us required.

💧

Inaudible watermark

An AudioSeal watermark embedded on synthesis and detectable after the fact — proven 0.0 on clean audio, 1.0 on watermarked.

Consent-first cloning

Cloning requires a recorded consent acknowledgement, stored in a per-voice consent ledger. Responsible by design.

🛡️

Content guardrails

Redact PII (email, phone, card) and block disallowed content by policy — before a word is ever spoken.

🗣️

Pronunciation control

Per-team lexicons teach the engine your brand names, drug names and tickers, so they're said right every time.

🏢

Self-host licensing

Run the whole stack in your VPC with an offline license; only usage counts — never audio — leave for billing.

How it works

From text to voice in four steps

Type or paste your text

Any of 30 languages. Long-form or a single line.

Pick, design or clone a voice

Use a preset, describe a new voice, or clone one with consent.

Generate on your own GPU

The engine runs in your VPC, your cloud, or a managed GPU — your audio never leaves your network.

Play, download, or call the API

Play in the Studio, export a 48kHz WAV, or hit the OpenAI-compatible API.

🎙️

Explore the interface

Create an account and sign in to explore the full Studio. Deploying the engine in your own environment requires a paid plan.

Sign in →
Pricing

One plan. Self-hosted.

Sign in to explore the interface for free. Deploying the engine is a paid, self-hosted license — you run it on your own hardware or cloud, so your audio never leaves your network.

Self-host

Enterprise

A$149/person/mo · minimum 2 people
Self-hosted voice AI for teams that can't send audio to the cloud.
  • Self-host in your VPC, cloud, or on-prem
  • 30 languages · voice design + consent-based cloning · 48kHz
  • Signed provenance + inaudible watermark on every clip
  • Content guardrails (PII redaction) + pronunciation control
  • Per-tenant usage metering · OpenAI-compatible API
  • Deploy bundle + offline license · onboarding support
Contact sales
FAQ

Questions, answered

What does “self-hosted” actually mean?

The voice engine runs on your own infrastructure — your cloud VPC or on-prem. Your text and audio never touch our servers. You manage it from the same Vocala control plane; only metadata (usage counts, plan) syncs.

How is this different from ElevenLabs or OpenAI TTS?

Quality is comparable, but those are cloud-only and proprietary. Vocala is built on the open Apache-2.0 VoxCPM engine, so you can run it yourself, avoid lock-in, and keep regulated data in-house — usually at lower cost at volume.

Is voice cloning safe and legal?

Cloning requires an explicit, recorded consent acknowledgement, stored in a per-voice consent ledger, and every generation is signed and watermarked. We do not support impersonation.

Which languages are supported?

30, including English, Chinese (and several dialects), Spanish, Hindi, Arabic, Japanese, French, German, Portuguese, Russian and more — no language tag needed.

How do I get started?

Create an account and sign in to explore the full interface for free. To deploy the engine in your own environment — your VPC, cloud, or on-prem — you'll need a paid plan (A$149 per person / month, minimum 2 people); contact sales and we'll issue your license.

Give your product a voice — without giving away your data.

Sign in to explore the interface. Become a paid customer to deploy in your own environment — your audio never leaves your network.