Real-time conversations
Listens first, dips its voice the moment you talk over it, and starts speaking before the whole reply is voiced.
Open-source voice infrastructure
Voxie runs the call: listening, turn-taking, interruptions and the voice, in 17 languages. It switches speech-to-text mid-call when the caller's language needs it. Your app only decides what to say.
Hears and speaks through
Listens first, dips its voice the moment you talk over it, and starts speaking before the whole reply is voiced.
A listener per call: a misheard turn is identified by language and moved to the listener that hears it.
Four request types over HTTP. Every turn arrives with the language it was heard in. You reply with text.
Failover at every layer. A provider going down degrades the call; it doesn't end it.
Turn conversations into action
From support lines to payment reminders, Voxie handles the hard parts of a live call, so your agent can focus on getting things done.
See how a call worksHow a call works
WebRTC (WHIP) from a browser, Twilio Media Streams for phone calls, or any SIP trunk through sip-server.
:8080Speech detection, turn-taking, barge-in, echo guard, and speech-to-text that adapts to the caller.
Voxie posts each turn to one HTTP endpoint of yours, with its language. You reply with text.
:8300Sentence by sentence: Kokoro on the GPU in ~130ms a sentence, Sarvam for Indian languages.
Give any app a voice: Voxie brings the call, the listening and the voice; your agent brings the knowledge and the tools.
See the quickstart
Answer questions in the caller's language, and hand off to a person when they ask.

A polite call about a failed payment. Softknock's REX runs its calls on Voxie.

Call back every lead in minutes, qualify them, and book the meeting.

Confirm, move or cancel appointments over the phone, in ten Indian languages and more.
Integrate
Voxie is open source. The images are published; the rest is your own app.
Add your provider keys and start the server and the voice router. You need keys for Deepgram, Groq and Sarvam (AssemblyAI is optional), and an NVIDIA GPU for the voice router.
git clone https://github.com/SomehowLiving/Voxie.git && cd Voxie
cp .env.example .env # DEEPGRAM_API_KEY, GROQ_API_KEY, SARVAM_API_KEY
docker compose -f compose.release.yaml up # server :8080 · voice router :8300 · test page :8400One config for every caller. Set your agent's URL and a shared key; Voxie sends it as a bearer token on every request.
[agent]
url = "https://your-app.example.com/voice/agent"
api_key = "…" # or STREAMCORE_AGENT_API_KEY
[stt]
provider = "adaptive" # a listener per callVoxie posts each event of the call to your endpoint and speaks what you send back.
type | When | Reply |
|---|---|---|
listen | Before speech-to-text starts | { language?, region? } |
greeting | The caller stayed silent | { text } |
chat | The caller finished a turn | { text, end_call? } |
oneshot | Background work | { text: "" } |
// A whole agent: Node 20+, no dependencies.
import { createServer } from "node:http";
createServer(async (req, res) => {
let body = "";
for await (const chunk of req) body += chunk;
const turn = JSON.parse(body); // { session_id, resource_id, type, text, language }
let reply = { text: "" };
if (turn.type === "listen") reply = {}; // or what you know: { language: "hi", region: "IN" }
if (turn.type === "greeting") reply = { text: "[lang:en] Hi, this is Voxie. How can I help?" };
if (turn.type === "chat") reply = { text: await think(turn.text, turn.language) };
res.writeHead(200, { "Content-Type": "application/json" });
res.end(JSON.stringify(reply));
}).listen(9101);
async function think(text, language = "en") { // your app goes here
return `[lang:${language}] You said: ${text}`;
}Published images
Both halves of Voxie are on GitHub's container registry, built for linux/amd64 with the Apache 2.0 licence and StreamCore's notice inside. Pin a version in production.
ghcr.io/somehowliving/voxie:845a2f1ghcr.io/somehowliving/voxie-voice-router:845a2f1docker pull ghcr.io/somehowliving/voxie:845a2f1
docker pull ghcr.io/somehowliving/voxie-voice-router:845a2f1
VOXIE_TAG=845a2f1 docker compose -f compose.release.yaml upStart each sentence with [lang:xx], for example [lang:ta]. Hindi and Marathi share a script, so text alone can't tell them apart. Answer only in languages the router's GET /health lists right now.
Set TWILIO_AUTH_TOKEN, then connect the call's audio to Voxie. No SIP trunk and no public IP: a public URL is enough.
<Response><Connect><Stream url="wss://<voxie-host>/twilio/media">
<Parameter name="resource_id" value="+919812345678"/>
</Stream></Connect></Response>Reply as text/event-stream and Voxie voices the first sentence while the rest is on its way. Send {"end_call": true} to hang up once the goodbye has played.
Languages, measured
Every result comes from real provider APIs, tested with scripted callers on clean and phone-quality audio.
No voice for these yet: your agent hears them, and answers in English.
Real voices. Measured results.
Measured with real APIs and scripted callers, never guessed.
Honest limits
Hosted and enterprise
Hosted Voxie, or help taking it to production on your own servers. Tell us what you're building and we'll get back to you.