How we test
How Meeting Cue’s On-Device AI Measures Up Against Claude
Meeting Cue writes its suggestions with Apple’s on-device language model. Because that model is much smaller than the ones running in data centres, we compared it directly with Anthropic’s Claude Opus and Claude Sonnet to see what the size difference actually costs.
At a glance.
What we tested.
We wrote 135 realistic meeting moments across nine industries: finance, healthcare, IT, hardware, construction, retail, business, entertainment and real estate. Each one has a few lines of conversation, a relevant passage from a public document, and a reference answer: what a well-prepared person in the room would say next.
Every model received exactly the same conversation and document passages. We tested five models:
- On your device: Apple’s built-in language model, with the instructions Meeting Cue ships today.
- On a desktop computer: gemma-4-12b and gpt-oss-20b, open models you can run on your own machine.
- In the cloud: Claude Opus 5 and Claude Sonnet 5, as our upper benchmark.
Each model wrote the same four lines Meeting Cue shows you: a plain-English summary of the conversation, one thing you could say right now, a broader question to raise, and a short friendly reaction.
Claude writes better suggestions.
We scored the answers in two ways.
Closeness to the reference answer. A text-similarity model measured how close each line came in meaning to the reference, from 0 (unrelated) to 1 (identical meaning).
| Model | Where it runs | Closeness |
|---|---|---|
| Claude Sonnet 5 | Cloud | 0.57 |
| Claude Opus 5 | Cloud | 0.57 |
| gemma-4-12b | Desktop | 0.54 |
| gpt-oss-20b | Desktop | 0.53 |
| Meeting Cue (Apple’s on-device model) | Your device | 0.50 |
On the line that matters most — what to say right now — Claude scored 0.63 to 0.64, and the on-device model 0.57.
A strict judge. A separate AI model marked every answer out of 8: is it supported by the document, does it commit to an answer, could you say it out loud, and does it stay on topic. To check the judge, we hid two answers of our own among the candidates: the written reference, which scored 7.3, and the canned fallback text Meeting Cue shows when no model is available, which scored 5.0.
Claude Sonnet (7.7) and Claude Opus (7.5) both scored above the written reference. The on-device model scored between 3.8 and 4.9 with earlier versions of our instructions, below the canned text, often turning the question back to the room instead of answering it. We have changed the instructions since then. The judge has not yet scored the version that ships today, so we are not quoting a number for it.
The gap is clear: a much larger model, running in a data centre, writes better lines.
Where the on-device model holds its own.
- It rarely invents figures. Across all 135 moments, the on-device model gave 4 amounts, percentages or quantities that appeared nowhere in the conversation or the document. Claude Sonnet did this 5 times, and Claude Opus 6 times.
- It answers almost everything. It answered 131 of 135. Apple’s safety filter sometimes declines finance and healthcare passages; when it does, Meeting Cue retries with the conversation alone rather than showing nothing.
- It is fast enough to use mid-conversation. It writes all four lines in about 3 seconds on an iPhone 16, and they appear as they are written. Claude took 9 to 11 seconds per moment, including the round trip over the network.
Fixing our biggest complaint without a bigger model.
Early feedback said the suggestions “seemed like a summary, not what I should say next.” That was fair. In a real meeting there often isn’t a relevant document, and then the on-device model tended to repeat the conversation back to you.
We tested 33 meeting moments with no documents at all. With our original instructions, the on-device model’s “what to say right now” line only repeated the conversation 63% of the time, and proposed something new just 3% of the time. After we rewrote the instructions, repetition fell to 9% and new proposals rose to 75%.
Meeting Cue now also checks every line before showing it and discards any line that only repeats what was just said, so you keep the previous suggestion instead. Even with documents, today’s on-device model still writes such a line 19% of the time; Claude never did in our tests. A bigger model is one way to fix this. Changing what we ask for, and checking the answer, fixed most of it while keeping everything on your device.
Finding the right passage is the real bottleneck.
A good suggestion needs the right passage from your documents, and that search is the same whichever model writes the lines. It runs on your device, using a small search model we trained on meeting speech.
People rarely speak the way documents are written. Someone asks, “Is six characters even allowed?”, and the document says, “Memorised secrets shall be at least 8 characters in length.”
- Typed questions: the right passage reaches the model in 87 of 135 moments.
- Real conversation: 54 of 135 with a standard search model, and 57 of 135 with ours.
If the passage isn’t found, no model — not even Claude — can use it. Closing this gap is what we are working on next.
What about meeting notes?
At the end of a meeting, Meeting Cue writes action items and key points, and links every note back to the lines of transcript it came from.
We haven’t compared these notes with Claude’s yet. We have compared them with summaries written by people, using 12 recorded meetings from the AMI Meeting Corpus that come with one.
- Working from the meeting’s reference transcript, 91% of the notes Meeting Cue showed were supported by the transcript lines quoted under them.
- Those notes covered about a third (0.32) of what the human note-takers wrote down.
- With a phone lying on the table and Meeting Cue’s own transcript, which is noisier, coverage falls to 0.23.
Every note shows its quote, one tap away, so you can check it. About one note in 25 to 50 misreads its quote, and the quote is how you catch it.
Why we still run on your device.
Claude is better at this task. But to use it, Meeting Cue would have to send every few seconds of your meeting, and your documents, to a server. With Meeting Cue, your audio, transcript and suggestions stay on your device, and nothing from a meeting is saved.
So the trade is this: suggestions that score 0.50 instead of 0.57, in exchange for a meeting that never leaves the room and a full set of lines in about 3 seconds rather than 9 to 11. We think that is the right trade for most meetings, and we will keep publishing the numbers as we close the gap.
How we tested: 135 meeting moments across 9 industries, each with a checked source passage from a public document and a written reference answer. Every model received identical conversation and passages. Closeness is the average text similarity (MiniLM) between each line a model wrote and the reference. The judge was Qwen3.5-9B running locally, one pass, with the written reference and the fallback text hidden among the answers as controls. The on-device scores were produced by Apple’s on-device model on a Mac, the same model an iPhone uses; timings are from an iPhone 16. The Claude models tested were Claude Opus 5 and Claude Sonnet 5; newer Claude models and Claude Haiku were not tested. The meeting-notes figures use 12 meetings from the AMI Meeting Corpus. Meeting Cue is not affiliated with Anthropic. More about Meeting Cue.