Notes — Audiary 2.0 Rev. 2026-09-10 Sheet 1 / 1

/audiary/second-opinion.md

A second opinion on “read”

Audiary 2.0 · Apple Intelligence Assist · Andrew Pope, Working Model

Every text-to-speech engine trips over the same handful of words. Audiary 2.0 asks the language model that already lives on your iPhone to settle them — without sending a single sentence anywhere.

In brief

# 1.0The words you wince at

English has a few hundred words that are spelled the same and said differently depending on what they mean. Linguists call them heteronyms. Listeners just call them mistakes, because that is how they land: “She read it yesterday” spoken with the present-tense “reed,” a “bass player” who apparently plays a fish, a dove that is somehow the past tense of dive.

They are rare on the page, maybe one every couple of hundred words in a novel, and you barely notice them when you read silently. Out loud they are impossible not to notice. A synthesizer can get ninety percent of them right and still sound like it is guessing, because the ten percent it misses are exactly the words a human never would. For someone who listens because reading is hard — dyslexia, low vision, a long day — a wrong “read” isn't a wince, it's a sentence that has to be replayed.

# 2.0Where rules stop

Audiary's neural voice runs entirely on the phone, and since 1.2 it has resolved heteronyms with a rules pass: a part-of-speech tagger decides whether “record” is a noun or a verb, and a small set of hand-written rules handles tense (“have read” is the past participle) and a few idioms (“wound up”). On the standard public benchmark for this problem, Google's WikipediaHomographData, that pass scores 89.7 percent. Google's own hand-written rules scored 89 percent on the same data in 2018, so the plateau is real and not ours alone.

The misses cluster in two places. Bare past tense, where nothing in the sentence marks the tense (“I read the letter”). And the words where part of speech says nothing at all: bass the fish and bass the instrument are both nouns, and so are both bows. On those sense-type words the rules manage about 60 percent, which is a coin flip dressed up as a system.

# 3.0The idea

The usual fix is to train a classifier. Google and Amazon both did, and their published systems reach 99 percent on this benchmark. That route needs training data per word, a model shipped with the app, and a retrain whenever the lexicon changes. For a one-person company shipping a $6.99 app, it is a lot of machinery for a problem that is, per sentence, very small.

The small version of the problem is a multiple-choice question. Here is a sentence. Here is one word in it. Which of these two meanings applies? That is a question a general language model answers well, and as of iOS 26 there is one already on the phone: the roughly three-billion-parameter model behind Apple Intelligence, exposed to apps through the Foundation Models framework. It runs on device, works offline, costs nothing per call, and adds nothing to Audiary's download.

So the design is a second opinion, not a replacement. The rules go first. Where they are confident, the model is never asked. Where they are silent, or on the specific words where the model has measurably better judgment, the sentence goes to the model with the target word marked and the choices written in plain English:

Sentence: "But not before I sort out the [[bass]] section." Target word: [[bass]] Options: - the fish - 'caught a bass' (rhymes with mass) - music: the low range, voice, or instrument - 'bass guitar', 'bass player' (rhymes with base)

Two details do a lot of work here. The model never sees phonetic notation, because language models are bad at phonetics and good at meaning; the answer is mapped back to a pronunciation afterwards. And the answer is constrained: the model can only emit one of the listed options verbatim, so there is nothing to parse and nothing to hallucinate.

# 4.0The numbers

Everything below is scored on the benchmark's evaluation split, 1,314 sentences across the 132 heteronyms Audiary's lexicon covers, with exactly the prompt and routing that ship in the app. Which words go to the model was chosen on the separate training split, so the evaluation set never influenced its own score.

SystemAccuracy
Unaided synthesizer (always the default reading)67.6%
Google hand-written rules, 201889.0%
Audiary rules alone89.7%
Audiary rules + Apple Intelligence Assist94.1%
Fine-tuned T5, single model, 202497.9%
Google rules + trained classifiers, shipped in Google TTS99.0%
Amazon, BERT embeddings + per-word classifiers, 202199.1%

Sources: Gorman et al. 2018; Nicolis & Klimkov 2021; Řezáčková et al. 2024. Audiary numbers measured 2026-09.

Read that honestly. The assist removes about 44 percent of the mispronunciations the rules leave behind, and it lands between Google's first rules and Google's first classifiers. The trained systems are still five points better. What they cost is training, a bundled model, and a retrain per lexicon change; what this costs is nothing new on the device and no training at all. It is the fast path, not the ceiling, and the ceiling is a project for another release.

The average also hides where the gain is. On the sense-type words, the ones you wince at, accuracy goes from about 60 percent to about 84. Bass went from zero of ten to ten of ten. On the noun-and-verb stress pairs the rules were already fine and stay in charge.

# 5.0The minefield

The first listening test was a short story written to be as hostile as possible: “He would wind the old clock while the wind chimes never moved… he wanted to present it to the museum as a present from the shop.” A dozen sentences, each carrying two senses of the same word. On the first run the model changed eleven readings and got eight of them wrong.

Every wrong one was in a sentence with the same word twice. The model was shown the whole passage, saw “wind chimes,” and decided that every “wind” in sight was a breeze. The fix was to show it less: only the sentence containing the target, with the target in brackets and every other occurrence of the same word blanked out, so “he would [[wind]] the old clock” arrives with no “wind chimes” to lean on. That took the minefield from three right to seven.

Then the more useful lesson. Sentences with two senses of one word are wordplay, not prose; real books almost never do it. The benchmark of natural Wikipedia sentences had never shown the failure because it never sets the trap. So the minefield stayed as a stress test, and a second script was written to read like an actual chapter, with one heteronym every couple of sentences and never two senses in one. That is the one the feature was judged on, and the one the routing was tuned against. Tuning to the trap would have optimized for a text nobody listens to.

# 6.0Getting the timing right

A model call takes about a third of a second on an iPhone 16 Pro. An utterance of audio lasts about fifteen. The arithmetic says the pass can stay far ahead of playback, and it does, but the interesting decisions were about what happens at the edges.

# 7.0On your phone, off by a switch

The framework's default model runs locally and works in Airplane Mode; the SDK Audiary is built with has no cloud path at all. The only network involvement is Apple's own one-time model download when you turn on Apple Intelligence, which happens outside the app. Audiary's privacy policy has a paragraph on it, and it is short because there is little to say: nothing you listen to leaves the device.

The assist is on by default on iPhones and iPads with Apple Intelligence, hidden entirely on devices without it, and switchable off under Settings › Voice. On iOS 17 and 18 the app behaves exactly as before, plus the rule improvements for “read” and “bow” that came out of the same work and ship everywhere. None of it needs an account, a setting to discover, or a decision from the listener; the point was always that the person who depends on the voice most should have to think about it least.

# 8.0What's next

The honest ceiling for this approach is a couple of points above where it sits, and the honest route past it is the trained classifier the big vendors use, which the evaluation harness built for this work makes straightforward to measure. Whether that is worth doing depends on whether listeners notice the difference between 94 and 99 on words they hear once an hour. The plan is to listen for a while and find out.

Audiary is a one-time purchase, no subscription, no ads, on the App Store. Questions about the approach: andrew@workingmodel.cc.