
Three and a Half Seconds
It is a Saturday evening and something in the other room is reading me an engineering report in a warm voice.
The report comes from a coding agent that has spent nine minutes going over my own website, and it has found something I could not have seen. Two places on the same page disagree with each other about the same component, three screens apart, and no human eye reading top to bottom was ever going to catch it. The agent has written the finding down and changed nothing, which is exactly the behaviour I want from it. Whether anything happens next is mine.
So I am hearing it out, and I pressed a button roughly three and a half seconds ago. That number is the only technical fact in this piece I want you to hold on to.
Why anything is being read to me at all
A working week now produces more material than a person can sit down with. Reports from agents. Diffs. Specifications. Long documents from other people who were also using machines to write them. The volume is not the interesting part. The interesting part is that it splits cleanly into two kinds, and almost nobody’s tooling respects the split.
Some of it has to be read. Read properly, at a desk, with a pen, because the words are load-bearing and I am going to be held to what I understood. Nothing gets to replace that.
But a great deal of it is material I need to have processed rather than material I need to commune with. I have to know what is in it. I do not need to give it the front of my mind. And reading that second category costs me the thing I am actually short of, which is not hours — it is focus. Every time I pull my attention over to a document I put down whatever I was holding, and picking it back up is the expensive part.
Being told something is a different posture from reading it. You can be told a thing while your hands are busy and your attention is elsewhere, and a surprising amount of that second category survives the treatment perfectly well.
So there are two tools. Calliope speaks, and Echo listens, turning my voice into text. Both run on machines in this house and nothing they handle leaves the building. Neither is a product. Nobody outside this address has used them and as far as I can tell nobody ever will. I want to talk about why I built them anyway, and about the thing I learned from it that turned out not to be about speech at all.
The controls are the confession
If you looked at the interface, the first thing you would notice is not the voice. It is the small controls underneath the text.
One of them says skip code. No company ships a skip code toggle. It would not survive a product meeting, because the people who need it are a vanishingly small share of anybody’s market. It exists in mine because on one particular afternoon I listened to a machine pronounce a hexadecimal colour value, then a file path, then a CSS selector, character by character, and decided I was never doing that again. It is not a feature. It is a scar.
The voices are the same tell. There are four, and they are not accents or genders. They are postures — a teacher, measured and instructional; a narrator; a critic; a companion. The one selected on a Saturday evening for a dense report about my own site was the companion, warm and unhurried, and that is not a capability decision. That is a man who has read enough long reports at the end of enough long days to know how he wants them delivered.
But the control I would point at, if you only got to look once, is the one that has been there since the very first version. As it speaks, the line it is speaking lights up on the page.
That is not a media player feature. Nobody following an audiobook needs it. It is there because I am very often not purely listening — I am half-listening and half-reading, and I want to be able to drop into the text at the exact place the voice has reached, take the two sentences that actually matter properly and slowly, and then let it carry me on again. It is built for a person crossing back and forth over that line all day, which is the entire reason the tool exists and is not a use case any general product is aimed at.
This is what software looks like when the whole user base is the person writing it. Controls nobody would fund, defaults nobody would defend, every one of them exactly right for one person. The interesting thing is not that such software can exist. It is that the cost of building it has fallen far enough that a single practitioner can now cross the distance between wanting a tool that fits and having one — and that the distance used to be much greater than most people have noticed it stopped being.
Not faster. Different.
Here is the part I did not expect, and I had it wrong myself for a while.
The first version of Calliope was good. I could hand it an entire article and get back a publishable, high-quality audio file, reliably, and I still think of it as a well-made thing. It shared a design language with Echo on purpose — the same proportions, the same restraint, the same feeling of having been made rather than assembled — so that together they looked like a practice rather than two weekends.
But I kept not reaching for it.
For a long time I assumed the problem was speed, and that what I needed was the same tool, quicker. That is not what happened, and the difference turns out to matter more than the speed does.
Hand the first version a long article and it takes as long as it takes. Minutes. Sometimes a lot of minutes. And that is completely fine, because it is a batch job — you give it the work, you go away, you come back to a finished file you could publish. Nobody standing there means nobody is waiting. For that job, speed is not a quality at all. Only the audio is.
What I actually needed most days was a different job. Read me this, now, while I am walking to the kitchen. And no amount of making the production tool faster produces that, because the production tool is answering a different question — one where the whole point is that you leave.
So the second attempt went after the other job. About three and a half seconds from pressing a button to hearing the first syllable, give or take; it moves around a little by run and by machine.
Somewhere under about four seconds something changes that is not a matter of degree. Above the line you invoke a tool: you decide to use it, you accept the wait as its price, and the decision is conscious every single time. Below the line you simply use it. The deliberation disappears. It stops being an event in your day and becomes a way you do things. The same threshold is visible in search, in autocomplete, and in the difference between a lift you press a button for and a staircase you just walk up.
And here is the detail that convinced me the threshold is real, because it is the one that costs something. The instant one is not as good. The audio is a step below the polished render, audibly, and I use it enormously more. For that job I would rather have a decent answer now than a beautiful one in four minutes. So latency was never a quality axis. It is a category axis. Below the threshold you do not get a better version of the same tool. You get a different tool, doing a different job, and the only way to find that out was to build the second one and watch which one I actually reached for.
I only got to find it out because I was the only customer. Anyone building for a market optimises audio quality, because quality is the axis that demonstrates well and sells. Latency is invisible in a demo and decisive in a life. Building for one person is not mainly about getting the features you want. It is about getting to choose which axis matters.
The proof of that is a thing I almost did not build. Partway through the second attempt I realised its fastest mode — the one that just reads you something, immediately — did not need the rest of the system around it at all. I lifted it out into a light version in an afternoon and put it on the house network, and now it answers on my phone, anywhere in the building. I use it constantly, including for the mundane case of reading back things my phone itself declines to read.
The names
Echo listens and returns what it hears. Calliope was the chief of the muses and the one associated with eloquence.
I chose both before I wrote either, and I am now fairly sure that was not decoration. Call a thing stt-service and you will build a service. Call it Echo and you find yourself asking what it means for it to return something faithfully — and that is a better question than the one you would otherwise have asked. Names set expectations, and you are the first person your own names happen to.
It changed the instructions, not just the workflow
Here is the consequence I did not design, and it is the one I would keep if I had to give the rest back.
Once you have a cheap, fast way of being told things, you start hearing what your collaborators actually write. And a great deal of what comes back from an AI assistant is formatted for a reader who is not there — bullet lists, tables, bold fragments, headings over three-sentence sections. Read aloud, it is unbearable. And the moment it is unbearable aloud you notice something worse: a lot of that structure was standing in for a point that had never been thought through in sentences.
So there is now a rule in my standing instructions, and it is one of the most useful things I have written. When we are thinking together — brainstorming, arguing, working out what the problem even is — write it as prose. Flowing sentences, the way you would say it out loud, no apparatus. Two people in a room. Brainstorming does not have charts in it. But when something has to be decided, analysed, compared, or held to later, that is a document, and a document gets everything a document should have: the lists, the tables, the numbers, the headings, the structure you can put a finger on.
The split is not cosmetic and it is not about my ears. It sorts by what I have to do with the thing. Talk is for building a shared understanding, and it can go past me while I am walking around the house. A decision is something I have to look at, weigh against another decision, and be accountable for afterwards, and it belongs on a page.
I did not sit down and design that. It fell out of having a tool that made listening cheap enough to do constantly, which made the mismatch obvious, which changed how I ask for things. It has since spread to how I work with more than one assistant. The tool was the smaller half; the working practice it exposed was the larger one, and it would still be true if both applications vanished tomorrow.
Where this argument costs me
Echo is the half that has not found its place, and I would rather be the one to say it.
The coding tool I use every day now ships a microphone. It is right there, somebody else maintains it, and it is good — and it does the ordinary dictation I used to open Echo for. That use is gone and I do not expect it back.
But I want to be exact about what was taken, because the easy story is that a large company beat me to it, and that is not what happened. Echo was never really aimed at general speech to text. Almost every model in that space cleans you up: it removes the ums, the false starts, the doubling back, and hands you a tidy sentence you did not say. Echo was built to do the opposite — to keep the disfluencies — because when what you are capturing is a person thinking out loud, the mess is the part worth keeping. The stumble is usually where the thought turned.
So the market arrived at speech to text and still has not arrived at Echo. What it took was the ordinary use I also happened to have.
Which leaves the honest version, and it is less tidy than a defeat. Echo is not dead and it is not fading. It is sitting still. It is waiting for the thing it ought to be the front door to — some way of catching what I think as I think it, and letting it accumulate somewhere I can walk back through later. I know roughly what that wants to be. I have not built it. So a tool I made and still believe in has been dormant for months while its author works out where it belongs, which is a far more common ending for personal software than either triumph or obsolescence, and almost nobody writes it down.
That is the cost of the argument, and I do not want to spend it and then take it back. The easy recovery here is I learned a lot, which is what people say when a project did not work, and I am not going to reach for it.
The honest claim is narrower than build your own tools. Most of the time you should not. A tool you build is a tool you maintain, forever, at your own expense, and the graveyard of abandoned personal software is enormous and largely deserved. The claim is this: when what you want is a different trade-off rather than a better version of the same one, nobody is coming to make it for you — and that has quietly become a solvable problem for one person with a few weekends. Wanting what everybody wants, only sooner, is not a reason to build. Wanting something on an axis the market does not price is.
And there is a corollary I did not expect and would pass on before any of the rest. Building the thing is the cheap part now. Knowing where it goes in your life is the expensive part. Calliope found its place the week it existed. Echo, which I think is the more interesting of the two, still has not.
Saturday evening
The report finished. The finding is on a card now, where findings go, and I will decide about it on Monday with the whole thing in front of me. My phone went quiet in the other room.
I did not build a text-to-speech engine. I built the specific experience of being told things — in a voice I chose, at a speed where I never think about the asking, with the line lit up in case I want to drop back into the words. And then, without meaning to, I rewrote how I ask a machine to talk to me.
The other one is upstairs, switched off, waiting for me to work out what it is for. I would build them both again tomorrow.