Repository navigation
fix(wordspotting): zero CTC keyword-spotter audio padding - #1003
Merged
Alex-Wengg merged 1 commit intoOct 9, 2026
Merged
Conversation
prepareAudioArray now builds the window through makePaddedAudioArray. That helper still allocates a rank-1 [maxSamples] or rank-2 [1, maxSamples] array in the mel input's float32 or float16 type, calls reset(to: 0) on the buffer, then writes the clamped leading samples. New tests check short clips, an empty clip, a clip that fills the window, a clip longer than the window, both dtypes, and both ranks, and require the copied prefix to match and every padding index to be finite and zero. Fixes FluidInference#991
Contributor
Author
|
@Alex-Wengg thanks for merging the keyword-spotter padding fix, so the CTC window always starts from zeroed audio. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why is this change needed?
prepareAudioArray now builds the window through makePaddedAudioArray. That helper still allocates a rank-1 [maxSamples] or rank-2 [1, maxSamples] array in the mel input's float32 or float16 type, calls reset(to: 0) on the buffer, then writes the clamped leading samples. New tests check short clips, an empty clip, a clip that fills the window, a clip longer than the window, both dtypes, and both ranks, and require the copied prefix to match and every padding index to be finite and zero.
CTC keyword spotting on short clips often returns log-probs that ignore the audio. Different sentences can greedy-decode to the same empty string and share one identical matrix, and the same clip can succeed or fail from run to run. prepareAudioArray allocates a fixed-length MLMultiArray for the mel window and copies only the clip into the front. The removed comment treated the tail as zero because MLMultiArray(shape:dataType:) was assumed to zero-fill; that initializer leaves storage uninitialized, and the mel model reads the whole window, so a non-zero or non-finite tail drives the log-probs.
Fixes #991