We build the voice and video input layer for people.
For the past few years almost all of this industry's money, and most of its best people, have gone into one thing: making models answer more accurately, more quickly, more like a person. That work has gone well.
But the hard part of using AI was never at that end. The hard part is this: when the thought arrives, you can't type.
Walking down the street, driving, just out of the shower, awake at three in the morning, standing in a queue, out on a run, on the way back from dropping the kids off. In those moments you have something to say, but there is no screen in front of you — or there is one and both hands are busy. By the time you sit down the sentence has left, or all that's still there is a blurred outline of it.
Input is the scarce half. And almost nobody is working on it.
Voice first, images close behind. Capture was never meant to have only one medium — sometimes the thing you want to keep is a sentence, sometimes it's something you only laid eyes on.
This is Ambient AI, and we also call it Aether AI: a layer that stays close to you, is there whenever you need it, and takes up none of your attention. It doesn't think for you. It only makes sure that when you have something to say, there is somewhere to say it.
The two get run together constantly, so it's worth being exact.
Edge AI is about moving the AI computation onto local hardware. That is not us. Our hardware does exactly two things locally — pick up sound and frame a shot. It doesn't read, doesn't understand, doesn't infer anything. Its whole job is to take your voice and what is in front of you cleanly off the world.
Turning that into text, turning a photo into something finished — all of it happens on your phone or your computer, on large models in the cloud. The hardware isn't clever. The hardware's job is to hear clearly and shoot clearly. That division of labour is deliberate: it lets the hardware be small, frugal with power and cheap, and it lets the model be swapped for a better one the moment a better one exists.
You speak, it captures, it turns that into text. The AI answers in text — you read the answer when you are back at a screen.
We don't do live voice conversation. Not for lack of ability. It's a choice.
In the second a thought arrives, what you want is not to lose it, not a chat. A conversation asks you to stop, spend attention, wait for a reply and then pick up again — while you are walking, driving, halfway through something else. A thing that demands your concentration, arriving exactly when you have none to spare.
Capture and conversation are two different things. We only do capture.
A photograph works the same way. One press and the frame is kept; what it turns into can wait until you are back at a screen. In the moment itself, the only thing you have to do is not miss it.
That choice makes the whole layer lighter: no wake word, nothing listening all day, nothing talking in your ear, and nothing that asks you to put down what you were doing.
This is the first principle behind every piece of hardware we build, and it is simple enough that it barely sounds like technology.
Sound pressure falls off with distance. Halve the distance and the pressure reaching the microphone roughly doubles — 6 dB. Background noise, on the other hand, doesn't get quieter because a microphone moved closer; it is spread evenly through the space. So every halving of the distance lifts more than a little signal. It lifts the whole signal-to-noise curve.
The microphone in an ordinary pair of earphones sits ten to fifteen centimetres from your mouth, which outdoors amounts to guessing. A phone in your hand can get to two or three centimetres, but first you have to pull it out, unlock it and find the right screen — and a thought does not survive three steps.
So for any piece of AI hardware the core question is the same one: how do you get the microphone closer to the mouth without asking for one extra movement?
The answer in a pair of AI sunglasses is to put the primary microphone at the bottom centre of the left rim, port facing down towards the mouth — worn, that is four to seven centimetres. If it's genuinely loud, or you would rather nobody heard you at all, take the glasses off and hold them up to your mouth: the same microphone is now two to three centimetres away. There is no mode to switch. You just shortened the distance with your hand.
We call this the Whisper Engine. What it means is this: at that distance you can speak very quietly, too quietly for someone half a metre away to hear, and it still gets every word.
It doesn't intrude on anyone else.A device worn against the body that records the whole room will always, in other people's eyes, be a problem. Our test is the level difference between the primary microphone and the reference one: when you speak the gap is 7 to 9 dB; someone half a metre away is almost equidistant from the two, so the gap is under 1 dB and never reaches the threshold. No voiceprint, no model — the port positions decide it.
It also protects you.This half gets raised far less often. In public, with a small piece of hardware held close to your mouth, you can speak at something close to a whisper. What you said stays between you and the AI.
The camera is another matter, and belongs in its own paragraph.The boundary on audio is settled by geometry; the boundary on photographs is settled by the person. It is for outdoors, for when you are on your own, not for recording other people. This is not a line we intend to fudge with technology.
Which is why it can't record a meeting.A few people brainstorming, an interview, a lecture — this layer can't do it, and that is what a voice recorder is for. It isn't a shortcoming. It is the other side of the same decision.
It doesn't think.It doesn't judge on your behalf, doesn't hand you advice, doesn't try to become one more AI. There are more than enough clever things on the market. What's missing is a quiet one.
It does one thing.It gets what you said out, whole. Fail at that one thing and nothing else stands up; do that one thing well and everything else can be left to somebody else.
Restraint isn't shipping fewer features. It is refusing to become something you have to learn.
The text that comes out can go into any coding agent, any framework, any box you can type into. You pick the destination; we put the words there.
There is a judgment about position buried in that: the giants will never unify each other's agents. Not one of them has a reason to hand a user smoothly over to a rival. Only someone who isn't on a side can do it.
So we deliberately don't build a model, and don't intend to. What we want is the wire, not either end of it.
Everything captured is stored as Markdown. An AI can read it, a person can read it, and the whole thing exports at any moment. Move to any other AI app and just take it with you.
We don't lock you in. A product that holds people through switching costs ends up spending its energy raising the walls instead of building anything worth staying for.
But the record itself turns into something else: an archive of how you think. Which thought arrived when, which questions keep coming back, which one you have been turning over for three years without an answer. What it is worth a year from now is nothing like what it is worth on day one.
The input layer isn't tied to any one form. The shell can change; the first principle doesn't. The microphone has to reach the mouth, and it records only the person wearing it.
This is the one we care most about right now, and not because it's cool. It's that when the sun is out you were going to wear sunglasses anyway — the only wearable nobody has to be talked into. A ring, a necklace, a pendant, a pet: each has to win a hard fight first, the fight to change what a person puts on their body. Sunglasses never have to fight it.
The main mic sits at the bottom centre of the left rim, its port facing down toward the mouth; the reference mic sits at the top of the right rim, port facing up and away, much further from the mouth. The camera is at the upper outer corner of the right rim. Playback, volume and calls all stay with the phone, the same as any pair of Bluetooth earbuds. No screen, no wake word.
Then add a prescription. Glasses with a prescription in them are worn roughly ten times the hours sunglasses are. And the property that matters most in an input layer isn't how clever it is — it's whether it's there when the thought arrives.
The camera makes this layer more than voice. And it follows the same principle the microphone does: either it's already on you, or one motion gets it into place. Pull out the phone, unlock it, find the camera — the instant worth photographing is gone somewhere inside those three steps.
The pendant module carries a microphone of its own. When you want it, you lift the pendant to your mouth — the closest this layer can get.
The rest of the time it is simply a pendant: at your chest, or hanging off a bag. Lift it, say the thing, put it back: three seconds in all. Its great advantage is that it changes nothing about what you wear.
Same principle in a friendlier shell. It sits on the desk, in a bag, in a child's hands; when you want to speak, you pick it up and bring it close.
Something you were already happy to pick up satisfies the “closest to the mouth” condition on its own.
The simplest one of them. It can be an ordinary pair of Bluetooth earbuds, or it can add local storage and pair with a transcription app. To catch a thought you take one out, hold it in your hand and bring it to your mouth; the rest of the day it stays in your ear; when you're not using it, it goes into a small cloth pouch and hangs at your chest as a pendant.
A true in-ear beats every other form in a loud environment, and when the world outside is noisy it is still the best choice. Not while you're moving, though — running or riding, the right pair of glasses fits better.
Forms will keep changing, and they should. The first principle doesn't.
AI wearables.They've reached their own iPhone-one moment. Every device sold is one more way in.
Speech to text.It crossed the usable line. Accuracy on casual speech, on dialects, in noise, has finally reached the point where you don't go back and fix it. Transcription moved from “roughly legible” to “usable as it stands”, which is what makes capture worth anything.
Agentic AI.It became reliable. So the words you capture finally have somewhere worth delivering them.
Position.What we mean to hold is the wire between a person and every AI, not any one of those AIs.
Stateful, not stateless.Anyone can copy a pipe. Nobody can copy a whole year of one person's memory — replacing it means starting from zero.
Scarce data.A structured record of how people actually drive AI in real life. That happens away from the screen, and only an entry point worn until it's second nature can reach it.
Neutrality.No model of our own, no side taken, so anything can plug in.
More use → deeper memory → higher switching cost, scarcer data → more people using it.
The voice interface between people and machines — plainly, the place where an ordinary person reaches AI — will not grow on a single object.
It's on the phone and on the desktop. It's also in AI glasses, in AI earbuds, in a voice ring, in a voice necklace, in a small voice pendant, in an AI pet that keeps you company.
The entry point doesn't belong to the smartest AI.
It belongs to whatever is nearest to you the moment you want to speak.