We’re thrilled to announce private reasoning via Venice is available on OpenHome!
You can now add a Venice API key in ‘Settings’, pick Venice as your Text-to-Text provider, and point any Agent at a Venice model.
If Venice is new to you: it's a privacy-first AI platform with a wide model catalog, including open-weight models that run on its own infrastructure alongside proxied access to frontier models from Anthropic, OpenAI, Google, and others. Venice is built so that what you send it isn't stored, making it a natural fit for our privacy-first, local-first design.
Building towards a device you control
A speaker is not a browser tab. It sits in a room where people talk without thinking about it, and it hears everything that comes with that: arguments, symptoms, money, kids doing homework, half-formed ideas nobody would type into a search box. Almost every mainstream assistant built for that room was built by a company whose business depends on knowing what happens in it.
We think that's backwards. You bought the hardware. You hold the keys to the models. What happens in your home, and the record of it, should be yours to control.
That belief is why OpenHome works the way it does. Bring-your-own-key exists so your usage is billed by providers you chose and can leave. Abilities run sandboxed with no direct access to storage or infrastructure. Local models have been supported from early on for people who want nothing to leave the building at all.
The near-term goal is more concrete: reduce the number of companies that end up holding what gets said in your home. We work with strong providers at every layer and will keep offering them, because they're a big part of what makes the experience good. But a voice pipeline spreads your audio and your transcripts across several vendors by default, each with its own retention terms, and most people never read any of them.
We're building a privacy-first path where all three layers can run through providers that store nothing, and we can ship that this year.
We're doing it in two phases. Phase 1 is the reasoning layer, which we shipped today with Venice. Phase 2 is voice, which puts the audio layers on the same footing and completes the pipeline. Here's what that means in practice.
What each layer of a voice pipeline gives away
Every spoken turn on an OpenHome device runs through three separate systems, and they don't see the same thing about you:
Speech-to-Text captures raw audio off the microphone. That's the waveform, not a transcript. It carries your voiceprint. It also picks up whoever else was talking in the room, the TV in the background, and whatever was said in the couple of seconds before you meant to address the device at all. Of the three layers this one is the most identifying, and it's the one almost nobody thinks about.
Text-to-Text gets the transcript. Where STT has your voice, TTT has the content: what you asked, what you said to your kid, and whatever context the Agent has built up about your schedule and house. And because an Agent's behavior is shaped by its interaction history, this layer can reveal patterns over time, rather than one isolated request.
Text-to-Speech gets the model's reply. It has half of every exchange, and half is usually enough to infer the other half. It's the smallest of the three exposures, and it's the one we'll move last.
A closed assistant runs all three inside a system/pipeline controlled by one company and asks you to trust the total. We split them, which means we can move them one at a time and you can see where you stand at each step.
Local models and cloud models
We've supported local STT, TTT, and TTS since the beginning, and running the whole pipeline on device is still the strongest privacy position available. Nothing leaves the house. Local also works when your internet doesn't, costs the same whether it talks all day or sits quiet, and never gets rate limited or deprecated out from under you.
The constraint is compute. Edge silicon has a fraction of what a datacenter GPU does, and it shows up in specific places: proper nouns, accented English, noisy rooms, flat synthesized voices, and multi-step Abilities where a heavily quantized model loses the thread. Small models have improved significantly over the past two years, and it is likely this constraint will ease over the next year.
Today, cloud models outperform local models on accuracy, voice quality, and reasoning depth. Cloud also needs a connection, bills per use, and puts you on someone else's uptime. And all three layers end up with companies that keep what you send them.
That last cost is the one that Venice solves, removing retention as a trade-off of using the cloud. The rest of the tradeoffs stay real and stay yours to weigh.
Local stays supported, and you can mix layers, running local STT with cloud reasoning or whatever combination works best for you.
Phase 1: Reasoning (now live)
The layers aren't equally hard, which is why we tackled reasoning first. TTT is text in, text out, and the abstraction was already clean on our side. STT and TTS involve audio formats, streaming behavior, and latency budgets that we want to test properly rather than rushing.
On Venice's Private tier, which is the default, inference runs on Venice-controlled GPUs or on partner infrastructure under zero-retention terms. Venice stores no prompt and no response. Your transcripts stop accumulating at the reasoning provider.
Phase 2: Voice (in progress)
STT is next in the queue, followed by TTS. Together they're Phase 2, and they're responsible for the part of the pipeline that changes what happens to your actual audio rather than the text made from it.
Once all three layers are on Venice, one key covers the whole pipeline and no outside model provider retains any part of the conversation. Not the audio, not the transcript, not the reply.
You're still running on OpenHome, so that isn't full end-to-end privacy. It is, however, a meaningful step toward a voice assistant where you decide who handles each layer of the conversation.
Venice: Picking a privacy mode
OpenHome is bring-your-own-key, so your Venice plan determines what your speaker can reach. Venice runs four privacy modes and they're worth understanding before you pick one:
Anonymized routes your request to a frontier model like Claude, GPT, or Gemini through Venice's proxy, which strips your identity before the provider sees it. The provider still handles the content under its own retention policy. So it protects who you are, not what you said.
Private is the default. Inference runs on Venice-controlled GPUs or zero-retention partner hardware, and nothing gets stored.
TEE, or Trusted Execution Environment, needs Venice Pro. Inference happens inside a hardware-isolated enclave run by outside partners including NEAR AI Cloud and Phala Network. This is where the privacy guarantee stops being a promise and becomes a property of the hardware.
E2EE also needs Pro. Your prompt gets encrypted before it leaves, stays encrypted through Venice, and only gets decrypted inside a verified enclave, so nobody operating the infrastructure along the way can read it. It costs you: web search and memory are off in E2EE mode, and responses come back slower.
For most people Private is the right setting. But a speaker in your kitchen hears things a browser tab never will, and some people have a higher bar than average. If you're running an OpenHome device in a clinic, a law office, or a house where the threat model is real, TEE and E2EE let you verify the privacy claim instead of trusting it. Drop in a Pro key and those tiers show up for your Agents right away.
Unrestricted models
Venice's open-weight models ship with no content restrictions, and this deserves more than a footnote because it's one of the sharper lines between an open platform and a closed one.
Anyone who has built a Character or Companion has hit this. A character with a strong point of view refuses to hold it. A horror story goes soft at the scary part. A blunt coach turns diplomatic exactly when blunt was the whole idea. Stripping that layer out isn't about edge cases. It's about agents that stay in character.
We think the call belongs to you. OpenHome ships with safeguards on by default, and we'll keep building for people who want them: filtered model options, guardrails on kids and family Agents, sensible defaults on anything published to the Marketplace. Want a device that behaves conservatively out of the box? That's what you get without touching a setting.
But if you want a fully unrestricted AI model running on hardware you own, in a house you control, on a model with no retention behind it, that's what we're building. Nobody approves it. No policy team decides what your speaker is allowed to discuss with you.
Get started with Venice
Generate a Venice API key at venice.ai
Add it under Profile > Settings > API Keys
Pick Venice as your TTT Platform in Model Configuration Settings
Choose a model and start a conversation
Models behave differently in voice than they do in chat, where nobody notices half a second. Latency and turn-taking are the whole experience on a speaker. If you find a model that holds up well or falls apart, tell us. That feedback decides which models we put in front of people by default when the full stack ships.
One integration doesn't fix voice AI. What it does is hand another decision back to you: which model reasons over what gets said in your house, what that company does with it afterward, and what you're willing to trade for accuracy. Every version of OpenHome we ship is meant to widen that set of choices rather than narrow it, and privacy is one of the things we intend to keep buying you more of.
That's the direction we're building toward: a voice assistant that gets more capable without requiring you to give up control of what happens inside your home.
Come find us at discord.gg/openhome
Join our early access program at https://dev.openhome.com/
About Venice
Venice is a privacy-first AI platform founded in 2024 by Erik Voorhees, built on leading open-source models and designed around the idea that machine intelligence should be private and unrestricted by default. It serves over a million users across web, iOS, and Android, and its API covers text, image, video, audio, embeddings, and web search under a single key.
Check out venice.ai
Learn more at docs.venice.ai