Skip to content

feat(diarizer): expose per-chunk embeddings from DiarizerManager - #992

Open
ertra wants to merge 1 commit into
FluidInference:mainfrom
ertra:feat/online-chunk-embeddings
Open

ertra wants to merge 1 commit into
FluidInference:mainfrom
ertra:feat/online-chunk-embeddings

Conversation

@ertra

@ertra ertra commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

Why is this change needed?

#633 added DiarizationResult.chunkEmbeddings so callers can post-process clustering without
re-running the embedding model, but only OfflineDiarizerManager fills it. DiarizerManager
(performCompleteDiarization) computes the same per-chunk, per-local-speaker embeddings and
always returns nil.

The online diarizer assigns speakers greedily, chunk by chunk, and never revisits a decision.
A caller that wants to re-cluster over the whole file has to reconstruct the (chunk, local
speaker) groups from TimedSpeakerSegment.embedding, relying on every segment of a local
speaker carrying a bit-identical array. That is an implementation detail rather than an API,
and it loses the chunk and slot numbers and never sees a speaker that got an ID in a chunk
without producing a segment there.

Motivation: re-clustering those embeddings with average linkage at the end of the file, in a
downstream macOS meeting recorder, took the mean DER on the 16-meeting AMI test set (headset
mix, collar 0.25, overlap included, DiarizerConfig.default with clusteringThreshold = 0.75)
from 36.1 % to 18.9 %. That result comes from the downstream code, not from this PR. This PR
only exposes the data; it changes no diarization output.

What changed

  • DiarizerConfig.exposeChunkEmbeddings (default false; same name and default as
    OfflineDiarizerConfig.exposeChunkEmbeddings), added as a trailing init parameter with a
    default.
  • performCompleteDiarization fills DiarizationResult.chunkEmbeddings when the flag is set,
    on both the debugMode and the normal return path. One entry per (chunk, local speaker) that
    received a speaker ID:
    • speakerId: the SpeakerManager ID that speaker's segments carry.
    • chunkIndex / speakerIndex: the chunk's position in the file and the local speaker slot.
    • startTimeSeconds / endTimeSeconds: first to last active frame of that speaker in the
      chunk, in file time (atTime included) — the same span definition as the offline pipeline.
    • embedding256: the same vector that speaker's segments carry.
    • rho128: empty; the online pipeline has no PLDA step.
  • Local speakers skipped for low activity or an invalid embedding get no entry.
  • The mapping is a pure internal static func buildChunkEmbeddings, unit-tested without
    models, mirroring OfflineDiarizerManager.buildPublicChunkEmbeddings.
  • Docs on ChunkEmbedding and DiarizationResult.chunkEmbeddings now cover both producers.

Backwards compatibility & performance

  • Opt-in. With the flag off (the default) the per-chunk helper is not called and the result is
    unchanged; verified on a real recording (segments identical, chunkEmbeddings == nil).
  • With the flag on, the helper is O(frames × local speakers) over data the chunk loop already
    holds: no model calls, no audio access.
  • All existing DiarizerConfig and DiarizationResult call sites compile unchanged.

Tests & lint

  • New DiarizerChunkEmbeddingTests (5 tests): ID, index and embedding propagation; slots with an
    empty ID or no active frame are skipped; the span runs from the first to the last active
    frame; empty input; the config default.
  • swift test --filter 'DiarizerChunkEmbeddingTests|ChunkEmbeddingExposureTests': 14 tests,
    0 failures.
  • swift format lint --configuration .swift-format clean on the three changed files; no new
    warnings.
  • End-to-end with the real pyannote_segmentation + wespeaker_v2 models on one 534 s
    multi-party call recording (clusteringThreshold = 0.75, Intel i9): 84 segments, 3 speakers,
    71 entries over 50 chunks. Every segment matched exactly one entry by speakerId with a
    bit-identical embedding and lay inside that entry's span; (chunkIndex, speakerIndex) was
    unique; every span lay inside its chunk; rho128 was empty throughout. 56 entries had at
    least one segment, equal to the number of distinct segment embeddings; the other 15 had none
    (spans 0.29–8.94 s, median 0.52 s). In 3 chunks two slots received the same speaker ID.

Notes for reviewers

  • Entries are a superset of segments. A speaker gets an ID with minActiveFramesCount
    active frames, and SpeakerManager.assignSpeaker applies no duration floor to a match with an
    existing speaker, while a segment needs one contiguous run of minSpeechDuration. A short
    backchannel from a known speaker is therefore an entry with no segment. That embedding did
    update the speaker model, so it is exposed; callers can filter by span.
  • Normalization. ChunkEmbedding's doc previously said embedding256 is L2-normalized.
    Neither extractor normalizes (OfflineEmbeddingExtractor.embedSpan documents its output as
    "not L2-normalized"; the online e2e run measured L2 norms of 1.11–6.28), so the doc now says
    the field carries the embedding as the extractor emitted it, and that the online vector is not
    unit-norm. Related, not changed here: DiarizerManager.extractSpeakerEmbedding's doc says
    "L2-normalized 256-dimensional embedding", but it returns EmbeddingExtractor.getEmbeddings
    output unchanged (we measured a norm of about 0.73 on real audio), and
    calculateEmbeddingQuality (magnitude / 10) assumes non-unit vectors. Happy to fix that
    doc here or separately.
  • The segment ↔ entry 1:1 invariant is not pinned by a unit test: it would need the models or a
    refactor of createTimedSegments. By inspection, buildChunkEmbeddings reads the same
    speakerIds[i] / embeddings[i] that createSegmentIfValid reads.

🤖 Generated with Claude Code

FluidInference#633 added DiarizationResult.chunkEmbeddings, filled only by the offline
pipeline. DiarizerManager.performCompleteDiarization computes the same
per-chunk, per-local-speaker embeddings and always returned nil, so a
caller that wants to re-cluster the online diarizer's output over the
whole file had to reconstruct the groups from identical segment
embeddings, losing the chunk and slot numbers.

Add DiarizerConfig.exposeChunkEmbeddings (default false, matching
OfflineDiarizerConfig). When set, performCompleteDiarization emits one
ChunkEmbedding per (chunk, local speaker) that received a speaker ID,
with the same speakerId and embedding its segments carry and a span from
the speaker's first to last active frame in the chunk. rho128 is empty:
the online pipeline has no PLDA step. Off by default, the output is
unchanged.

An entry can exist without a segment: an ID needs minActiveFramesCount
active frames, a segment one contiguous run of minSpeechDuration.

The mapping is a pure internal static helper, buildChunkEmbeddings,
covered by DiarizerChunkEmbeddingTests without loading models. The
ChunkEmbedding doc no longer claims embedding256 is L2-normalized;
neither extractor normalizes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant