Skip to content

CtcKeywordSpotter: input padding is not zeroed, so CTC log-probs are often independent of the audio #991

Description

@boostedchaos

Summary

CtcKeywordSpotter.prepareAudioArray allocates a fixed-size MLMultiArray (240,000 samples for CTC-110M) and copies only the clip's samples into it. The rest is assumed to be zero:

// Copy actual samples (MLMultiArray is zero-initialized, so padding is implicit).

MLMultiArray(shape:dataType:) does not guarantee zeroed memory. In our runs the unfilled tail often held non-zero values, including NaN/Inf. The model reads the padding (it does not stop at the real audio length), so spotKeywordsWithLogProbs frequently returns log-probs that do not depend on the input audio at all.

Location: Sources/FluidAudio/ASR/Parakeet/SlidingWindow/CustomVocabulary/WordSpotting/CtcKeywordSpotter+Inference.swift lines 236–259 at 0b1f4628; the same code is on main at 04e363c2.

Measurements (synthetic say audio only, CTC-110M, macOS, Apple Silicon)

  • 4 different short sentences (2–6 s, 16 kHz mono), each spotted 5× in each of 3 processes, with an empty vocabulary. With .all compute units, 6 of 60 log-prob matrices greedy-decoded to the sentence. The other 54 decoded to an empty string, and they were byte-identical across the four different sentences (one shared hash). With .cpuOnly, 17 of 60 were usable.
  • We copied prepareAudioArray line for line, ran it against the same loaded models, and counted the padding just before prediction (36 calls, 3 processes). Every bad call had padding full of non-zero/non-finite values. Good calls had zeros, or the tail of the previous clip's audio.
  • A stride-aware deep copy of the mel and encoder outputs did not change anything, so output-buffer reuse does not explain it.
  • Zero-padding the input ourselves to 240,000 samples: 60 of 60 correct, with one matrix hash per clip across 3 processes, under both .all and .cpuOnly.

This may also explain item 2 of #967 (results changing between identical runs).

Repro idea

  1. Make two different short synthetic clips (say -o a.aiff "…", convert to 16 kHz mono Float32).
  2. In one process, call spotKeywordsWithLogProbs(audioSamples:customVocabulary:minScore:) with an empty vocabulary on each clip several times.
  3. Greedy-decode each matrix (argmax per frame, collapse repeats, drop blank) and hash the Float bits.
  4. Repeat with each clip extended to 240,000 samples with zeros first.

Note: comparing hashes alone is not enough, because different clips can produce identical (empty) evidence. Decode the matrices too.

Suggested fix

Zero the whole input array before copying the samples, for example with memset on dataPointer or by filling it in the same loop. Do not rely on allocation contents. The last partial chunk of long audio (computeLogProbsChunked) goes through the same function, so it is covered by the same fix.

Our workaround on the app side: pad every input to full windows ourselves (240,000 samples, or 240000 + ceil((n − 240000) / 208000) × 208000 for longer audio), then keep only the frames that cover the real audio.

Activity

  1. mvanhorn commented on Oct 8, 2026

    @mvanhorn
    Contributor

    prepareAudioArray allocates a fixed-length MLMultiArray for the mel window and copies only the clip into the front. The removed comment treated the tail as zero because MLMultiArray(shape:dataType:) was assumed to zero-fill; that initializer leaves storage uninitialized, and the mel model reads the whole window, so a non-zero or non-finite tail drives the log-probs.

    prepareAudioArray now builds the window through makePaddedAudioArray.

  2. added a commit that references this issue on Oct 9, 2026
    e05ab92
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions