“Who the hell is Alexa?” the voice came over my mobile phone, angry and accusatory. I made some excuse and hung up, knowing it was only the beginning. It was launch day: November 6, 2014.
Say hello to Alexa
Let me back up a bit. I was over a decade into my voice design career, and I’d been working at Amazon for about a year as the only voice/conversation designer across all three of its voice-enabled products: FireTV, Fire Phone, and Echo/Alexa. The first two of them used speech recognition software from a company called Nuance, which had been in the speech recognition industry for a long time. Echo/Alexa, however, used Amazon’s own speech recognition instead. This came as a great surprise to Nuance, which assumed it would be Amazon’s speech recognition vendor for the foreseeable future. Hence, the phone call from the Nuance program manager I’d been working with on the Fire Phone for a year.
Who, in fact, was Alexa? It wasn’t the first voice assistant; Siri had existed for a while, after all, on iPhones. How was Alexa any different? Well for one, it was a device whose primary modality was voice - not touch, like a phone. It sat in the background listening for its “wake word,” which was a new concept at the time.
Alexa had limited functionality at launch: play music, check the weather, set timers and alarms, and a few more features. However, its hands-free nature powered by a wake word, and its communal nature as a device that could be used by everyone in a room, were both unique, and leapfrogged the competition. Alexa and the Echo devices became hugely popular over the next decade.
Looking back, Alexa’s big leap had more to do with where it lived and how you got its attention than with how smart it was. I suspect the next leap will be similar.
Beyond words: a new primitive
Over the last couple of years, with LLM-based technologies like Open AI’s ChatGPT and Anthropic’s Claude, there has been a focus on advancing chat-based interfaces to be more human-like and act as personal assistants. Both have voice capabilities too, but they are secondary. Whatever happened to the voice-first assistant as the most cutting edge technology?
Chat and voice are different modalities, and they’re good for different things. A lot of what’s coming out now takes a chat experience and ports it to voice as though it were the same thing, which is a bit like treating a magazine article as a conversation.
Amazon’s Alexa+ and Google’s Gemini for Home exist, to be sure. However, they are more adaptations of the existing voice assistant to use LLM-powered technology to be more “conversational” vs. anything totally new and revolutionary, as Alexa was in 2014.
We propose something new; we believe the fundamental primitive should not be text. It should be voice itself. It has long been understood that the voice signal contains far more information than what is simply contained in transcribed text. Voice carries emotion, context, and meaning.
In fact in some cases, such as sarcasm, the entire meaning of the sentence can be the opposite of the written text. Consider the phrase “I’m so happy at my job!” vs. “I’m soooo happy <sarcastic> at my job” spoken two different ways.
Just as a human walks in and “reads the room” beyond the words spoken, the smart speaker - if truly intelligent - should glean more meaning than simple spoken text. It should provide the emotion and tone of the speakers, as well as capture any environmental audio that contains meaning: a glass breaking, dogs barking, babies crying. These all provide context about what’s happening in the environment. A truly intelligent speaker shouldn’t just understand the words someone says, it should understand the entire context and use that information to better anticipate user needs.
We call this “spatial intelligence”, and it is what we propose as the future of voice assistants.
The next generation: spatial intelligence
So, what’s next? We believe the next generation of devices will be able to understand the world around them via spatial intelligence.
Understanding the room matters because of what it lets a device do next: anticipate what people need, sometimes before they ask. For a decade, assistants have waited for a request. The next generation will be able to notice.
The washer finishes, and the speaker reminds you before the clothes sit overnight. You ask about your calendar while rushing out the door, and it hears you’re in a hurry and mentions only the meeting that matters.
Or take a note-taking app. Today it listens and writes things down. With spatial intelligence it can act more like a moderator. If the room goes quiet, or it senses confusion, it can speak up. Getting more meaning out of the room gives the device a better sense of when and how to act.
Alexa meant you could stop reaching for a device. Anticipation means you can skip a lot of the asking.
This is the future we are building at OpenHome, where AI is no longer trapped in chat but becomes an ambient intelligence layer woven into everyday life. OpenHome provides the hardware, platform, and community that let developers create voice-first experiences for the physical spaces people work and live in.
Through our open-source Voice SDK, local models, sensor layer, and DevKit, OpenHome turns a home device into an intelligent, conversational interface that can listen, sense, understand, and act in the real world. Anticipation happens between understanding and acting, when the device does something useful with what it notices.
Why anticipation matters
Anticipation removes the constraint that boxed in the skills era. On Alexa, a skill mostly spoke when spoken to. Anticipation needs developers who can decide when their work should speak up. Picture a speaker in an elderly man’s home that hears a thud and then nothing. It asks if he’s okay, and when he doesn’t answer, it texts his daughter. Nobody in that moment is going to say a wake word.
To be clear about where this stands: today, speech on OpenHome gets transcribed and a language model works from the text. The difference is developer control. Abilities can run in the background and keep notes the assistant reads, and some can be started by the assistant with no trigger phrase. Those are the first pieces of anticipation. The DevKit’s custom multi-microphone array, built to pick up sound across a room, is where the spatial work starts.
Control matters more once an assistant anticipates. In a walled garden, one company decides which needs get noticed and whose product answers them. On OpenHome, developers run whatever AI agent they want, whether that’s a commercial model on their own API key or an open model on their own hardware, and they can switch when something better comes along.
Get started with OpenHome
We’re accepting applications for the next round of DevKit hardware at dev.openhome.com. Spots are limited.
To learn more, join our Discord and check out our developer docs.