Live captions, as it is said
Speech turns into text the moment it is spoken, split into turns so a back-and-forth reads like a conversation rather than one long paragraph.
Put the phone on the table and Aloud writes out what the room is saying, as it is said, in text big enough to read from across it. Nobody else has to install anything, join anything, or hold anything. It is a hearing aid for a conversation you would otherwise have to keep asking people to repeat.
Speech recognition guesses. Every other app hands you the guess as a finished sentence and lets you believe all of it equally — which is how a misheard name, dose or time becomes something you find out about later, after acting on it.
Apple’s recogniser returns a confidence for every word, and sometimes the other readings it weighed and threw away. That is normally discarded somewhere between the framework and the screen. Aloud keeps it, and compares each word with the rest of its own turn — so a word it hesitated over is drawn as hesitated over.
It does not catch everything, and it would be worth less if we pretended it did: the recogniser can be wrong and sure of itself, and no amount of reading its confidence catches that. What this removes is the far more common case — the word it was never confident about in the first place, handed to you looking exactly like the words it was.
Not a meeting recorder you read afterwards. Everything here is pointed at the words landing on screen fast enough to reply to.
Speech turns into text the moment it is spoken, split into turns so a back-and-forth reads like a conversation rather than one long paragraph.
In Studio, each change of speaker takes its own colour, so a three-way conversation stays readable instead of blurring into one voice.
Your half of the conversation. Type a reply at full screen size and hold the phone up, or drop it straight into the transcript mid-session.
Attention drifted? One tap summarises the last minute. It reads the transcript already on your phone — nothing is sent anywhere to do it.
Set a session to stop after 5 to 90 minutes. When the clock runs down it pauses and asks — plus five, plus fifteen, or finish — rather than closing on you mid-sentence.
Sessions are filed automatically with a title and summary, searchable by any word anyone said, and stored on your phone alone.
Not marketing tiers — each one sets the speech recogniser up differently, and the right choice depends on how far away the voice is.
Runs the speech model on the phone itself. Best within arm's reach — on-device models start dropping words across a room.
Hands the work to Apple's server recogniser, which hears markedly further and copes better with a noisy room.
Everything Range does, plus a shorter gap before a turn is called finished — so a change of speaker lands on its own line.
There is no account to make and nothing for the other person to do.
There is no account and no server of ours. Saved sessions live in the app's own storage on your device, and Nearby never sends audio anywhere at all.
Read Privacy PolicyAnyone who has spent a conversation nodding along and hoping they guessed right.
Captions for the room, not just for a video. The text scales to whatever size you need, and Write Back gives you the other half of the conversation.
Cafés, hospital corridors, a relative who has grown softly spoken. Range hears across the table when your ears cannot.
A session saves itself with a title and a summary, so what the doctor or the client actually said is there to check afterwards.