The Owlery Works™

Meeting Cue 1.1

Meeting Cue Now Knows Who’s Speaking

Transcripts are only as useful as knowing who said what. With our latest update, Meeting Cue delivers a major leap in speaker recognition powered by NVIDIA’s Nemotron 3 Diarization model, running 100% locally on your device.

Whether your phone is lying flat on a noisy conference table or you’re running a fast-paced team huddle, your meeting transcripts are about to become significantly sharper, clearer, and more accurate.

What’s New at a Glance

91.2%Accuracy on the table. Correctly attributes lines even when your phone is sitting flat in the middle of the room (up from 78.3%).
Up to 8Zero manual setup. Automatically identifies, numbers, and separates up to 8 distinct speakers as soon as they talk.
2%Built for privacy and speed. 100% on-device. On an iPhone 16 the Neural Engine is busy only about 2% of the time.

Precision Performance Across Every Setup

We benchmarked our new model against real-world meeting recordings across identical transcript lines. The results show dramatic improvements in every microphone setup we tested:

  • Phone on the table: accuracy jumped to 91.2% (up from 78.3%).
  • Close microphones (iPhone and iPad): reached 92.6% accuracy (up from 85.6%).
  • Close microphones (desktop): increased to 95.2% accuracy (up from 86.9%).

Transcript lines given the right speaker

Bar chart. Phone on the table: 78.3% before, 91.2% now. Close microphones on iPhone and iPad: 85.6% before, 92.6% now. Close microphones on the desktop: 86.9% before, 95.2% now.
AMI Meeting Corpus, first 10 minutes of each meeting: 12 meetings for the iPhone and iPad rows, 16 for the desktop row. The new model was better or level in every meeting except two with the phone on the table.

With the phone on the table, the time attributed to the wrong speaker dropped by more than half, from 4.7% down to 2.0%.

Under the Hood: Real-Time Stream Processing

Traditional speaker tracking waited for a person to finish talking, generated a static “voiceprint,” and tried to group similar audio snippets together.

Meeting Cue now listens continuously. The new model checks voice activity every 10 milliseconds, tracking up to 8 distinct speakers in the order they first speak. As text is generated, the app maps each line to the dominant voice detected during that exact timeframe. If a line arrives before its last moment has been heard, Meeting Cue gives it a best guess and corrects the speaker label about a second later if needed.

Diagram. The microphone feeds two things at once: speech recognition, which produces lines of text, and the speaker model, which reports every 10 milliseconds who is talking. Each line is then labelled with the voice heard most while it was spoken. Microphone Speech recognition lines of text Speaker model Labelled transcript Speaker 1 Can we ship Friday? Speaker 2 Only if QA signs off. Speaker 3 I can do that. who is talking, every 10 ms Microphone Speech recognition lines of text Speaker model who is talking, every 10 ms Labelled transcript Speaker 1 Can we ship Friday? Speaker 2 Only if QA signs off. Speaker 3 I can do that.
A line that arrives before its last moment has been heard gets a best guess at once, corrected about a second later if the guess was wrong.

Upgrade Overview

Feature / metricBeforeNow
Phone-on-table accuracy78.3%91.2%
Close-microphone accuracy85.6%92.6%
Wrong-speaker audio time (phone on the table)4.7%2.0%
Number of speakersSet by youAutomatic (up to 8 voices)
Processing location100% on-device100% on-device

Efficient by Design

Meeting Cue runs speaker tracking on the iPhone 16’s Neural Engine, the part of Apple’s chip built for this kind of work. Processing a 0.72-second audio slice takes just 15.7 milliseconds, nearly three times faster than on the main CPU (44 milliseconds).

Time to process 0.72 seconds of audio on an iPhone 16 (milliseconds)

Bar chart. Neural Engine: 15.7 milliseconds. Main processor: 44 milliseconds.
Measured on the device over two minutes of audio. Lower is better.

Limitations & Edge Cases

  • Similar vocal traits: speakers with near-identical pitch and tone may occasionally share a label.
  • Capacity limits: tracking supports up to 8 voices per session.
  • Seamless name tags: natural introductions remain fully supported. When a participant says “Hi, I’m Sarah,” Meeting Cue changes their label to “Sarah” across the transcript.

Meeting Cue uses NVIDIA’s Nemotron 3 Diarization model under the OpenMDW license, optimized for local execution. Your audio is processed entirely on your device and never leaves it. Benchmarked on held-out meetings from the AMI Meeting Corpus; the model was trained on other AMI meetings, so these recordings may be a little more familiar to it than yours. More about Meeting Cue.

← All posts