/audiary/second-opinion.md
A second opinion on “read”
Every text-to-speech engine trips over the same handful of words. Audiary 2.0 asks the language model that already lives on your iPhone to settle them — without sending a single sentence anywhere.
In brief
- Words like “read,” “bass,” and “dove” are checked against the Apple Intelligence model already on the phone whenever Audiary's own rules can't tell which meaning is meant.
- On the standard public benchmark, mispronunciations of these words fall by about 44 percent.
- Entirely on device, works offline, on by default where Apple Intelligence is available, off with one switch.
# 1.0The words you wince at
English has a few hundred words that are spelled the same and said differently depending on what they mean. Linguists call them heteronyms. Listeners just call them mistakes, because that is how they land: “She read it yesterday” spoken with the present-tense “reed,” a “bass player” who apparently plays a fish, a dove that is somehow the past tense of dive.
They are rare on the page, maybe one every couple of hundred words in a novel, and you barely notice them when you read silently. Out loud they are impossible not to notice. A synthesizer can get ninety percent of them right and still sound like it is guessing, because the ten percent it misses are exactly the words a human never would. For someone who listens because reading is hard — dyslexia, low vision, a long day — a wrong “read” isn't a wince, it's a sentence that has to be replayed.
# 2.0Where rules stop
Audiary's neural voice runs entirely on the phone, and since 1.2 it has resolved heteronyms with a rules pass: a part-of-speech tagger decides whether “record” is a noun or a verb, and a small set of hand-written rules handles tense (“have read” is the past participle) and a few idioms (“wound up”). On the standard public benchmark for this problem, Google's WikipediaHomographData, that pass scores 89.7 percent. Google's own hand-written rules scored 89 percent on the same data in 2018, so the plateau is real and not ours alone.
The misses cluster in two places. Bare past tense, where nothing in the sentence marks the tense (“I read the letter”). And the words where part of speech says nothing at all: bass the fish and bass the instrument are both nouns, and so are both bows. On those sense-type words the rules manage about 60 percent, which is a coin flip dressed up as a system.
# 3.0The idea
The usual fix is to train a classifier. Google and Amazon both did, and their published systems reach 99 percent on this benchmark. That route needs training data per word, a model shipped with the app, and a retrain whenever the lexicon changes. For a one-person company shipping a $6.99 app, it is a lot of machinery for a problem that is, per sentence, very small.
The small version of the problem is a multiple-choice question. Here is a sentence. Here is one word in it. Which of these two meanings applies? That is a question a general language model answers well, and as of iOS 26 there is one already on the phone: the roughly three-billion-parameter model behind Apple Intelligence, exposed to apps through the Foundation Models framework. It runs on device, works offline, costs nothing per call, and adds nothing to Audiary's download.
So the design is a second opinion, not a replacement. The rules go first. Where they are confident, the model is never asked. Where they are silent, or on the specific words where the model has measurably better judgment, the sentence goes to the model with the target word marked and the choices written in plain English:
Two details do a lot of work here. The model never sees phonetic notation, because language models are bad at phonetics and good at meaning; the answer is mapped back to a pronunciation afterwards. And the answer is constrained: the model can only emit one of the listed options verbatim, so there is nothing to parse and nothing to hallucinate.
# 4.0The numbers
Everything below is scored on the benchmark's evaluation split, 1,314 sentences across the 132 heteronyms Audiary's lexicon covers, with exactly the prompt and routing that ship in the app. Which words go to the model was chosen on the separate training split, so the evaluation set never influenced its own score.
| System | Accuracy |
|---|---|
| Unaided synthesizer (always the default reading) | 67.6% |
| Google hand-written rules, 2018 | 89.0% |
| Audiary rules alone | 89.7% |
| Audiary rules + Apple Intelligence Assist | 94.1% |
| Fine-tuned T5, single model, 2024 | 97.9% |
| Google rules + trained classifiers, shipped in Google TTS | 99.0% |
| Amazon, BERT embeddings + per-word classifiers, 2021 | 99.1% |
Sources: Gorman et al. 2018; Nicolis & Klimkov 2021; Řezáčková et al. 2024. Audiary numbers measured 2026-09.
Read that honestly. The assist removes about 44 percent of the mispronunciations the rules leave behind, and it lands between Google's first rules and Google's first classifiers. The trained systems are still five points better. What they cost is training, a bundled model, and a retrain per lexicon change; what this costs is nothing new on the device and no training at all. It is the fast path, not the ceiling, and the ceiling is a project for another release.
The average also hides where the gain is. On the sense-type words, the ones you wince at, accuracy goes from about 60 percent to about 84. Bass went from zero of ten to ten of ten. On the noun-and-verb stress pairs the rules were already fine and stay in charge.
# 5.0The minefield
The first listening test was a short story written to be as hostile as possible: “He would wind the old clock while the wind chimes never moved… he wanted to present it to the museum as a present from the shop.” A dozen sentences, each carrying two senses of the same word. On the first run the model changed eleven readings and got eight of them wrong.
Every wrong one was in a sentence with the same word twice. The model was shown the whole passage, saw “wind chimes,” and decided that every “wind” in sight was a breeze. The fix was to show it less: only the sentence containing the target, with the target in brackets and every other occurrence of the same word blanked out, so “he would [[wind]] the old clock” arrives with no “wind chimes” to lean on. That took the minefield from three right to seven.
Then the more useful lesson. Sentences with two senses of one word are wordplay, not prose; real books almost never do it. The benchmark of natural Wikipedia sentences had never shown the failure because it never sets the trap. So the minefield stayed as a stress test, and a second script was written to read like an actual chapter, with one heteronym every couple of sentences and never two senses in one. That is the one the feature was judged on, and the one the routing was tuned against. Tuning to the trap would have optimized for a text nobody listens to.
# 6.0Getting the timing right
A model call takes about a third of a second on an iPhone 16 Pro. An utterance of audio lasts about fifteen. The arithmetic says the pass can stay far ahead of playback, and it does, but the interesting decisions were about what happens at the edges.
- Start early, never wait. The pass begins the moment you open something in the player, not when you press play, because people look at the screen for a second or two first and that is enough for the opening sentences. Synthesis takes whatever answers are ready and never delays audio for one that is not. An early version waited up to three seconds for an in-flight answer on prefetched utterances; it worked, and it was deleted, because a synthesis path that blocks on a model is the kind of cleverness that turns into a bug report.
- Look ahead, but only so far. The pass stays forty utterances ahead of the playhead, about ten minutes of audio, and moves with it. A whole-book pass up front would mean tens of minutes of model calls for chapters you might never reach; this way a novel costs calls only for what is actually listened to, and the worker goes idle when the window is answered rather than polling.
- Ask about heteronyms, not about text. Most utterances contain none, and a lowercase substring check dismisses them in microseconds. Of the words that remain, the routing table sends most to the rules without a call. In practice a novel generates around twenty model calls per hour of listening, roughly seven seconds of model time, against several minutes of neural synthesis for the same hour. The assist is a rounding error on the battery the voice already uses.
- Remember, and forget correctly. Answers are cached per utterance, so a second listen makes no calls. The cache is keyed to the prompt version and the app build, so an update that changes a rule or a gloss never replays a stale reading.
- Fail toward yesterday. About four percent of ordinary sentences hit the model's content guardrails and are refused; a few more hit an unavailable model or a decoding error. Every one of those falls back to the rules, which is exactly what Audiary did last month. The feature can be wrong in new ways, but it cannot be worse than the app it replaced.
# 7.0On your phone, off by a switch
The framework's default model runs locally and works in Airplane Mode; the SDK Audiary is built with has no cloud path at all. The only network involvement is Apple's own one-time model download when you turn on Apple Intelligence, which happens outside the app. Audiary's privacy policy has a paragraph on it, and it is short because there is little to say: nothing you listen to leaves the device.
The assist is on by default on iPhones and iPads with Apple Intelligence, hidden entirely on devices without it, and switchable off under Settings › Voice. On iOS 17 and 18 the app behaves exactly as before, plus the rule improvements for “read” and “bow” that came out of the same work and ship everywhere. None of it needs an account, a setting to discover, or a decision from the listener; the point was always that the person who depends on the voice most should have to think about it least.
# 8.0What's next
The honest ceiling for this approach is a couple of points above where it sits, and the honest route past it is the trained classifier the big vendors use, which the evaluation harness built for this work makes straightforward to measure. Whether that is worth doing depends on whether listeners notice the difference between 94 and 99 on words they hear once an hour. The plan is to listen for a while and find out.
Audiary is a one-time purchase, no subscription, no ads, on the App Store. Questions about the approach: andrew@workingmodel.cc.