Meeting Cue 1.1
Meeting Cue Now Knows Who’s Speaking
Transcripts are only as useful as knowing who said what. With our latest update, Meeting Cue delivers a major leap in speaker recognition powered by NVIDIA’s Nemotron 3 Diarization model, running 100% locally on your device.
Whether your phone is lying flat on a noisy conference table or you’re running a fast-paced team huddle, your meeting transcripts are about to become significantly sharper, clearer, and more accurate.
What’s New at a Glance
Precision Performance Across Every Setup
We benchmarked our new model against real-world meeting recordings across identical transcript lines. The results show dramatic improvements in every microphone setup we tested:
- Phone on the table: accuracy jumped to 91.2% (up from 78.3%).
- Close microphones (iPhone and iPad): reached 92.6% accuracy (up from 85.6%).
- Close microphones (desktop): increased to 95.2% accuracy (up from 86.9%).
Transcript lines given the right speaker
With the phone on the table, the time attributed to the wrong speaker dropped by more than half, from 4.7% down to 2.0%.
Under the Hood: Real-Time Stream Processing
Traditional speaker tracking waited for a person to finish talking, generated a static “voiceprint,” and tried to group similar audio snippets together.
Meeting Cue now listens continuously. The new model checks voice activity every 10 milliseconds, tracking up to 8 distinct speakers in the order they first speak. As text is generated, the app maps each line to the dominant voice detected during that exact timeframe. If a line arrives before its last moment has been heard, Meeting Cue gives it a best guess and corrects the speaker label about a second later if needed.
Upgrade Overview
| Feature / metric | Before | Now |
|---|---|---|
| Phone-on-table accuracy | 78.3% | 91.2% |
| Close-microphone accuracy | 85.6% | 92.6% |
| Wrong-speaker audio time (phone on the table) | 4.7% | 2.0% |
| Number of speakers | Set by you | Automatic (up to 8 voices) |
| Processing location | 100% on-device | 100% on-device |
Efficient by Design
Meeting Cue runs speaker tracking on the iPhone 16’s Neural Engine, the part of Apple’s chip built for this kind of work. Processing a 0.72-second audio slice takes just 15.7 milliseconds, nearly three times faster than on the main CPU (44 milliseconds).
Time to process 0.72 seconds of audio on an iPhone 16 (milliseconds)
Limitations & Edge Cases
- Similar vocal traits: speakers with near-identical pitch and tone may occasionally share a label.
- Capacity limits: tracking supports up to 8 voices per session.
- Seamless name tags: natural introductions remain fully supported. When a participant says “Hi, I’m Sarah,” Meeting Cue changes their label to “Sarah” across the transcript.
Meeting Cue uses NVIDIA’s Nemotron 3 Diarization model under the OpenMDW license, optimized for local execution. Your audio is processed entirely on your device and never leaves it. Benchmarked on held-out meetings from the AMI Meeting Corpus; the model was trained on other AMI meetings, so these recordings may be a little more familiar to it than yours. More about Meeting Cue.