Repository navigation
Conversation
FluidInference#633 added DiarizationResult.chunkEmbeddings, filled only by the offline pipeline. DiarizerManager.performCompleteDiarization computes the same per-chunk, per-local-speaker embeddings and always returned nil, so a caller that wants to re-cluster the online diarizer's output over the whole file had to reconstruct the groups from identical segment embeddings, losing the chunk and slot numbers. Add DiarizerConfig.exposeChunkEmbeddings (default false, matching OfflineDiarizerConfig). When set, performCompleteDiarization emits one ChunkEmbedding per (chunk, local speaker) that received a speaker ID, with the same speakerId and embedding its segments carry and a span from the speaker's first to last active frame in the chunk. rho128 is empty: the online pipeline has no PLDA step. Off by default, the output is unchanged. An entry can exist without a segment: an ID needs minActiveFramesCount active frames, a segment one contiguous run of minSpeechDuration. The mapping is a pure internal static helper, buildChunkEmbeddings, covered by DiarizerChunkEmbeddingTests without loading models. The ChunkEmbedding doc no longer claims embedding256 is L2-normalized; neither extractor normalizes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why is this change needed?
#633 added
DiarizationResult.chunkEmbeddingsso callers can post-process clustering withoutre-running the embedding model, but only
OfflineDiarizerManagerfills it.DiarizerManager(
performCompleteDiarization) computes the same per-chunk, per-local-speaker embeddings andalways returns
nil.The online diarizer assigns speakers greedily, chunk by chunk, and never revisits a decision.
A caller that wants to re-cluster over the whole file has to reconstruct the (chunk, local
speaker) groups from
TimedSpeakerSegment.embedding, relying on every segment of a localspeaker carrying a bit-identical array. That is an implementation detail rather than an API,
and it loses the chunk and slot numbers and never sees a speaker that got an ID in a chunk
without producing a segment there.
Motivation: re-clustering those embeddings with average linkage at the end of the file, in a
downstream macOS meeting recorder, took the mean DER on the 16-meeting AMI test set (headset
mix, collar 0.25, overlap included,
DiarizerConfig.defaultwithclusteringThreshold = 0.75)from 36.1 % to 18.9 %. That result comes from the downstream code, not from this PR. This PR
only exposes the data; it changes no diarization output.
What changed
DiarizerConfig.exposeChunkEmbeddings(defaultfalse; same name and default asOfflineDiarizerConfig.exposeChunkEmbeddings), added as a trailinginitparameter with adefault.
performCompleteDiarizationfillsDiarizationResult.chunkEmbeddingswhen the flag is set,on both the
debugModeand the normal return path. One entry per (chunk, local speaker) thatreceived a speaker ID:
speakerId: theSpeakerManagerID that speaker's segments carry.chunkIndex/speakerIndex: the chunk's position in the file and the local speaker slot.startTimeSeconds/endTimeSeconds: first to last active frame of that speaker in thechunk, in file time (
atTimeincluded) — the same span definition as the offline pipeline.embedding256: the same vector that speaker's segments carry.rho128: empty; the online pipeline has no PLDA step.internal static func buildChunkEmbeddings, unit-tested withoutmodels, mirroring
OfflineDiarizerManager.buildPublicChunkEmbeddings.ChunkEmbeddingandDiarizationResult.chunkEmbeddingsnow cover both producers.Backwards compatibility & performance
unchanged; verified on a real recording (segments identical,
chunkEmbeddings == nil).holds: no model calls, no audio access.
DiarizerConfigandDiarizationResultcall sites compile unchanged.Tests & lint
DiarizerChunkEmbeddingTests(5 tests): ID, index and embedding propagation; slots with anempty ID or no active frame are skipped; the span runs from the first to the last active
frame; empty input; the config default.
swift test --filter 'DiarizerChunkEmbeddingTests|ChunkEmbeddingExposureTests': 14 tests,0 failures.
swift format lint --configuration .swift-formatclean on the three changed files; no newwarnings.
pyannote_segmentation+wespeaker_v2models on one 534 smulti-party call recording (
clusteringThreshold = 0.75, Intel i9): 84 segments, 3 speakers,71 entries over 50 chunks. Every segment matched exactly one entry by
speakerIdwith abit-identical embedding and lay inside that entry's span;
(chunkIndex, speakerIndex)wasunique; every span lay inside its chunk;
rho128was empty throughout. 56 entries had atleast one segment, equal to the number of distinct segment embeddings; the other 15 had none
(spans 0.29–8.94 s, median 0.52 s). In 3 chunks two slots received the same speaker ID.
Notes for reviewers
minActiveFramesCountactive frames, and
SpeakerManager.assignSpeakerapplies no duration floor to a match with anexisting speaker, while a segment needs one contiguous run of
minSpeechDuration. A shortbackchannel from a known speaker is therefore an entry with no segment. That embedding did
update the speaker model, so it is exposed; callers can filter by span.
ChunkEmbedding's doc previously saidembedding256is L2-normalized.Neither extractor normalizes (
OfflineEmbeddingExtractor.embedSpandocuments its output as"not L2-normalized"; the online e2e run measured L2 norms of 1.11–6.28), so the doc now says
the field carries the embedding as the extractor emitted it, and that the online vector is not
unit-norm. Related, not changed here:
DiarizerManager.extractSpeakerEmbedding's doc says"L2-normalized 256-dimensional embedding", but it returns
EmbeddingExtractor.getEmbeddingsoutput unchanged (we measured a norm of about 0.73 on real audio), and
calculateEmbeddingQuality(magnitude / 10) assumes non-unit vectors. Happy to fix thatdoc here or separately.
refactor of
createTimedSegments. By inspection,buildChunkEmbeddingsreads the samespeakerIds[i]/embeddings[i]thatcreateSegmentIfValidreads.🤖 Generated with Claude Code