Summary
CtcKeywordSpotter.prepareAudioArray allocates a fixed-size MLMultiArray (240,000 samples for CTC-110M) and copies only the clip's samples into it. The rest is assumed to be zero:
// Copy actual samples (MLMultiArray is zero-initialized, so padding is implicit).
MLMultiArray(shape:dataType:) does not guarantee zeroed memory. In our runs the unfilled tail often held non-zero values, including NaN/Inf. The model reads the padding (it does not stop at the real audio length), so spotKeywordsWithLogProbs frequently returns log-probs that do not depend on the input audio at all.
Location: Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CustomVocabulary/WordSpotting/CtcKeywordSpotter+Inference.swift lines 236–259 at 0b1f4628; the same code is on main at 04e363c2.
Measurements (synthetic say audio only, CTC-110M, macOS, Apple Silicon)
- 4 different short sentences (2–6 s, 16 kHz mono), each spotted 5× in each of 3 processes, with an empty vocabulary. With
.all compute units, 6 of 60 log-prob matrices greedy-decoded to the sentence. The other 54 decoded to an empty string, and they were byte-identical across the four different sentences (one shared hash). With .cpuOnly, 17 of 60 were usable.
- We copied
prepareAudioArray line for line, ran it against the same loaded models, and counted the padding just before prediction (36 calls, 3 processes). Every bad call had padding full of non-zero/non-finite values. Good calls had zeros, or the tail of the previous clip's audio.
- A stride-aware deep copy of the mel and encoder outputs did not change anything, so output-buffer reuse does not explain it.
- Zero-padding the input ourselves to 240,000 samples: 60 of 60 correct, with one matrix hash per clip across 3 processes, under both
.all and .cpuOnly.
This may also explain item 2 of #967 (results changing between identical runs).
Repro idea
- Make two different short synthetic clips (
say -o a.aiff "…", convert to 16 kHz mono Float32).
- In one process, call
spotKeywordsWithLogProbs(audioSamples:customVocabulary:minScore:) with an empty vocabulary on each clip several times.
- Greedy-decode each matrix (argmax per frame, collapse repeats, drop blank) and hash the Float bits.
- Repeat with each clip extended to 240,000 samples with zeros first.
Note: comparing hashes alone is not enough, because different clips can produce identical (empty) evidence. Decode the matrices too.
Suggested fix
Zero the whole input array before copying the samples, for example with memset on dataPointer or by filling it in the same loop. Do not rely on allocation contents. The last partial chunk of long audio (computeLogProbsChunked) goes through the same function, so it is covered by the same fix.
Our workaround on the app side: pad every input to full windows ourselves (240,000 samples, or 240000 + ceil((n − 240000) / 208000) × 208000 for longer audio), then keep only the frames that cover the real audio.
Summary
CtcKeywordSpotter.prepareAudioArrayallocates a fixed-sizeMLMultiArray(240,000 samples for CTC-110M) and copies only the clip's samples into it. The rest is assumed to be zero:MLMultiArray(shape:dataType:)does not guarantee zeroed memory. In our runs the unfilled tail often held non-zero values, including NaN/Inf. The model reads the padding (it does not stop at the real audio length), sospotKeywordsWithLogProbsfrequently returns log-probs that do not depend on the input audio at all.Location:
Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CustomVocabulary/WordSpotting/CtcKeywordSpotter+Inference.swiftlines 236–259 at0b1f4628; the same code is onmainat04e363c2.Measurements (synthetic
sayaudio only, CTC-110M, macOS, Apple Silicon).allcompute units, 6 of 60 log-prob matrices greedy-decoded to the sentence. The other 54 decoded to an empty string, and they were byte-identical across the four different sentences (one shared hash). With.cpuOnly, 17 of 60 were usable.prepareAudioArrayline for line, ran it against the same loaded models, and counted the padding just beforeprediction(36 calls, 3 processes). Every bad call had padding full of non-zero/non-finite values. Good calls had zeros, or the tail of the previous clip's audio..alland.cpuOnly.This may also explain item 2 of #967 (results changing between identical runs).
Repro idea
say -o a.aiff "…", convert to 16 kHz mono Float32).spotKeywordsWithLogProbs(audioSamples:customVocabulary:minScore:)with an empty vocabulary on each clip several times.Note: comparing hashes alone is not enough, because different clips can produce identical (empty) evidence. Decode the matrices too.
Suggested fix
Zero the whole input array before copying the samples, for example with
memsetondataPointeror by filling it in the same loop. Do not rely on allocation contents. The last partial chunk of long audio (computeLogProbsChunked) goes through the same function, so it is covered by the same fix.Our workaround on the app side: pad every input to full windows ourselves (240,000 samples, or
240000 + ceil((n − 240000) / 208000) × 208000for longer audio), then keep only the frames that cover the real audio.