You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Vocabulary rescoring currently computes CTC evidence even when no transcript
word is eligible for correction. This adds hasCTCRescoringCandidates to the
rescorer and prepared session so callers can skip that work. It uses the existing
matching rules, including aliases, compounds, multiword terms and thresholds.
Empty vocabulary or timings return false. Acoustic rescue conservatively returns
true for nonempty inputs. Callers that need standalone keyword detections must
still run rescoring. Term forms are also prepared once, and transcript words are
normalized once per request.
Initial measurements on one Apple Silicon Mac reduced unnecessary separate CTC
work from about 140–153 ms to 0.3 ms, and shared-head work from about 12 ms to
0.3 ms. These are component timings, not whole-dictation speedups.
Validation
144 selected checks pass, including parity against real CTC evidence across
12 candidate cases, threshold handling, empty inputs and acoustic rescue.
Contribution lint, strict lint for changed Swift files and formatting checks
pass.
The default Intel build passes, as does Intel Release compilation for macOS 14
with the optional text-processing trait disabled.
The streaming vocabulary test class is excluded: its first-word replacement
failure also reproduces on unchanged main. The full suite and hosted CI have not
been verified green.
Looks correct — early-stop placement matches all three candidate loops and the cached forms/set are equivalent to the old builders. One nit: hasCTCRescoringCandidates returns true whenever spotterRescueEnabled (default on), but rescue only runs on the term-centric path under largeVocabThreshold; gating it with !useBKTree && terms.count <= largeVocabThreshold would let large vocabularies skip CTC too.
Thanks for catching that. Fixed in 6c57258: the rescue shortcut now requires !useBKTree and terms.count <= largeVocabThreshold. Larger vocabularies can skip CTC when no text candidates match, even with rescue enabled. Added regression coverage at and above the threshold, including positive candidates and prepared-session parity, and updated the docs. All 103 selected vocabulary checks and formatting checks pass locally.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why is this change needed?
Vocabulary rescoring currently computes CTC evidence even when no transcript
word is eligible for correction. This adds
hasCTCRescoringCandidatesto therescorer and prepared session so callers can skip that work. It uses the existing
matching rules, including aliases, compounds, multiword terms and thresholds.
Empty vocabulary or timings return false. Acoustic rescue conservatively returns
true for nonempty inputs. Callers that need standalone keyword detections must
still run rescoring. Term forms are also prepared once, and transcript words are
normalized once per request.
Initial measurements on one Apple Silicon Mac reduced unnecessary separate CTC
work from about 140–153 ms to 0.3 ms, and shared-head work from about 12 ms to
0.3 ms. These are component timings, not whole-dictation speedups.
Validation
12 candidate cases, threshold handling, empty inputs and acoustic rescue.
pass.
with the optional text-processing trait disabled.
The streaming vocabulary test class is excluded: its first-word replacement
failure also reproduces on unchanged main. The full suite and hosted CI have not
been verified green.