Skip to content

Unicode case-insensitive pronunciation matches select the wrong replacement #2450

Description

@rudycelekli

Reproduction

The regex matcher accepts Unicode case variants such as dotless i, while replacement lookup uses a different casefold equivalence. Some matched words remain unchanged. Separately, distinct regex literals Straße and STRASSE collide in the casefold lookup and use the same replacement. Three failures and two controls reproduce this in the real pronunciation helper.

Expected behavior

Associate each literal regex alternative with its replacement. Preserve longest-key priority and last-entry precedence for equal-length case variants, including language-specific overrides.

The reproduction uses local files/localhost HTTP only; no model downloads or external services are involved. A focused patch and regression tests are prepared against current main.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions