Skip to content

Commit ea63ac3

Browse files
authored
fix(tts/kokoro-ane/zh): only fold an erhua 儿 into the previous syllable (#983) (#986)
Fixes #983. ## Problem `MandarinErhua.merge` folded every syllable whose pinyin base is `er` into the syllable before it. 二 / 而 / 耳 / 尔 share that pinyin with 儿, so after another syllable they were swallowed, and 儿 meaning "child" (女儿, 他儿子) lost its `ér`. Digits were affected as well, because `MandarinNumberNormalizer` turns `12` into `十二` before G2P. ## Fix - `Segment.pinyin` carries the Hanzi it was read from (`word:` replaces the unused `hanziCount:`). - `MandarinPinyinNormalizer.Syllable` gains `isErhuaSuffix` (defaults to `false`; the initializer stays source-compatible). - `MandarinG2P` sets the flag on a word's last syllable when `MandarinErhua.isSuffix(word:preceding:)` agrees, following misaki's `_merge_erhua`: - the character is 儿 or 兒; - it ends its word, or the segmenter split it off as a one-character word (这 | 儿 when 这儿 isn't a phrase), in which case it folds into the Hanzi before it in the buffer; - misaki's `not_erhua` / `must_erhua` word lists allow it (ported from `ZHFrontend`). - `MandarinErhua.merge` only folds flagged syllables. Not ported: misaki's POS gate (`a` / `j` / `nr`), since this pipeline has no POS tags. One behavior change outside the bug: user-lexicon `.syllables` tokens are taken literally, so an `er` written in a lexicon entry no longer folds. Erhua can still be written there with `@` bopomofo. ## Before / after Real `ANE-zh` assets, phonemes printed by `fluidaudiocli tts … --backend kokoro-ane --variant zh` (debug build): | Input | Before | After | |---|---|---| | 相隔二千里 | `ㄒ阳1ㄍㄜㄦ2ㄑ言1ㄌㄧ3` | `ㄒ阳1ㄍㄜ2ㄦ4ㄑ言1ㄌㄧ3` | | 别了二十年 | `ㄅㄝ2ㄌㄜㄦ5ㄕ十2ㄋ言2` | `ㄅㄝ2ㄌㄜ5ㄦ4ㄕ十2ㄋ言2` | | 第十二回 | `ㄉㄧ4ㄕ十ㄦ2ㄏ为2` | `ㄉㄧ4ㄕ十2ㄦ4ㄏ为2` | | 第12回 | `ㄉㄧ4ㄕ十ㄦ2ㄏ为2` | `ㄉㄧ4ㄕ十2ㄦ4ㄏ为2` | | 偶尔 | `ㄡㄦ3` | `ㄡ2ㄦ3` | | 然而他 | `ㄖㄢㄦ2ㄊㄚ1` | `ㄖㄢ2ㄦ2ㄊㄚ1` | | 木耳 | `ㄇㄨㄦ4` | `ㄇㄨ4ㄦ3` | | 女儿 | `ㄋㄩㄦ3` | `ㄋㄩ3ㄦ2` | | 他儿子 | `ㄊㄚㄦ1ㄗㄭ5` | `ㄊㄚ1ㄦ2ㄗㄭ5` | | 这儿 | `ㄓㄜㄦ4` | `ㄓㄜㄦ4` (unchanged) | | 小孩儿 | `ㄒ要3ㄏㄞㄦ2` | `ㄒ要3ㄏㄞㄦ2` (unchanged) | | 一会儿 | `ㄧ2ㄏ为ㄦ4` | `ㄧ2ㄏ为ㄦ4` (unchanged) | | 哪儿 | `ㄋㄚㄦ3` | `ㄋㄚㄦ3` (unchanged) | | 玩儿 | `万ㄦ2` | `万ㄦ2` (unchanged) | ## Tests - `MandarinErhuaTests`: the merge primitives now set `isErhuaSuffix`. Added cases for unflagged `er` (十二), `isSuffix` (other `er` characters, word-initial 儿, standalone 儿 with and without preceding Hanzi, `not_erhua` / `must_erhua`, traditional 兒), and end-to-end runs through `MandarinG2P` (十二, 相隔二千, 然而 / 偶尔 / 木耳, 女儿 as a phrase and split, 他儿子, phrase-final 小孩儿). - `swift test --filter 'MandarinErhuaTests|MandarinG2PTests|MandarinCustomLexiconTests|MandarinNumberNormalizerTests'`: 105 tests, 0 failures. - `swift format lint` is clean on the changed files.
1 parent 04e363c commit ea63ac3

4 files changed

Lines changed: 233 additions & 38 deletions

File tree

‎Sources/FluidAudio/TTS/KokoroAne/G2P/Mandarin/MandarinErhua.swift‎

Lines changed: 59 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -13,22 +13,28 @@ import Foundation
1313
///
1414
/// Boundary rules:
1515
///
16-
/// * `儿` at index 0 of the sandhi buffer is *never* merged — keeps
17-
/// standalone words (`儿子 érzi`, `儿童 értóng`) intact.
16+
/// * Only a syllable flagged `isErhuaSuffix` folds. The flag is set by
17+
/// `MandarinG2P` from the source text via `isSuffix(word:preceding:)`:
18+
/// the character must be `儿` / `兒` (not `二`, `而`, `耳`, `尔`, which
19+
/// share the pinyin `er`), it must end its word or stand alone, and
20+
/// the word must not be one misaki keeps as a full `ér` (`女儿`,
21+
/// `婴儿`, `孤儿`, …).
22+
/// * `儿` at index 0 of the sandhi buffer is *never* merged — there is
23+
/// nothing to fold into.
1824
/// * The preceding syllable must itself not be `er` — back-to-back
1925
/// `er er` is left as two syllables.
20-
/// * Any non-`er` base is mergeable. The whitelist is intentionally
21-
/// loose; cases that misaki blocks (e.g. `儿` at the start of a
22-
/// polysyllabic word) are already filtered out by the
23-
/// `dict.phrases` / single-char lookup happening upstream.
26+
///
27+
/// misaki also skips words tagged `a`, `j` or `nr` by jieba's POS tagger;
28+
/// this pipeline has no POS tags, so the word lists carry the common
29+
/// cases instead.
2430
///
2531
/// Operates in place on the same `pendingSyllables` buffer that
2632
/// `MandarinToneSandhi.apply` will see — invoke this *before* sandhi
2733
/// so 3+3 promotion considers the (now shorter) buffer.
2834
public enum MandarinErhua {
2935

30-
/// Fold trailing `er` syllables into their predecessors. Mutates
31-
/// `syllables` in place; the merged-into syllable gains
36+
/// Fold flagged trailing `er` syllables into their predecessors.
37+
/// Mutates `syllables` in place; the merged-into syllable gains
3238
/// `erhua = true`, the trailing `er` is removed.
3339
public static func merge(_ syllables: inout [MandarinPinyinNormalizer.Syllable]) {
3440
guard syllables.count >= 2 else { return }
@@ -40,7 +46,7 @@ public enum MandarinErhua {
4046
while i >= 1 {
4147
let cur = syllables[i]
4248
let prev = syllables[i - 1]
43-
if cur.base == "er" && shouldMergeInto(prev: prev) {
49+
if cur.base == "er" && cur.isErhuaSuffix && shouldMergeInto(prev: prev) {
4450
syllables[i - 1].erhua = true
4551
syllables.remove(at: i)
4652
// Advance past the now-merged anchor so an immediately
@@ -53,11 +59,55 @@ public enum MandarinErhua {
5359
}
5460
}
5561

62+
/// Whether the last character of `word` is an erhua suffix that may
63+
/// fold into the syllable before it. `preceding` is the Hanzi already
64+
/// in the sandhi buffer before `word`; it is only consulted when the
65+
/// segmenter split `儿` off as a word of its own (`这` + `儿` when
66+
/// `这儿` isn't in the phrase dict), which is the one case where the
67+
/// syllable to fold into belongs to an earlier word.
68+
///
69+
/// Mirrors misaki's rule: the character is `儿`, it is the last one of
70+
/// its word, and neither the word nor its last two characters are in
71+
/// `notErhua` unless the word is in `mustErhua`.
72+
public static func isSuffix(word: [Character], preceding: [Character]) -> Bool {
73+
guard let last = word.last, last == "儿" || last == "兒" else { return false }
74+
// The syllable `儿` folds into, and the text the word lists see.
75+
let context: [Character]
76+
if word.count >= 2 {
77+
context = word
78+
} else {
79+
guard !preceding.isEmpty else { return false }
80+
context = Array(preceding.suffix(2)) + word
81+
}
82+
let normalized = String(context).replacingOccurrences(of: "兒", with: "儿")
83+
if mustErhua.contains(normalized) { return true }
84+
let tail2 = String(normalized.suffix(2))
85+
let tail3 = String(normalized.suffix(3))
86+
return !notErhua.contains(normalized) && !notErhua.contains(tail2) && !notErhua.contains(tail3)
87+
}
88+
5689
/// Whitelist for the merge predicate. Conservative on purpose —
5790
/// any non-empty, non-`er` base is allowed.
5891
private static func shouldMergeInto(
5992
prev: MandarinPinyinNormalizer.Syllable
6093
) -> Bool {
6194
!prev.base.isEmpty && prev.base != "er"
6295
}
96+
97+
/// Words whose final `儿` is always read as erhua. From misaki's
98+
/// `ZHFrontend.must_erhua`.
99+
static let mustErhua: Set<String> = [
100+
"小院儿", "胡同儿", "范儿", "老汉儿", "撒欢儿", "寻老礼儿", "妥妥儿", "媳妇儿",
101+
]
102+
103+
/// Words whose final `儿` keeps its own `ér` syllable — mostly `儿`
104+
/// meaning "child" (`女儿`, `婴儿`) or part of a name (`红孩儿`).
105+
/// From misaki's `ZHFrontend.not_erhua`.
106+
static let notErhua: Set<String> = [
107+
"虐儿", "为儿", "护儿", "瞒儿", "救儿", "替儿", "有儿", "一儿", "我儿", "俺儿", "妻儿",
108+
"拐儿", "聋儿", "乞儿", "患儿", "幼儿", "孤儿", "婴儿", "婴幼儿", "连体儿", "脑瘫儿",
109+
"流浪儿", "体弱儿", "混血儿", "蜜雪儿", "舫儿", "祖儿", "美儿", "应采儿", "可儿", "侄儿",
110+
"孙儿", "侄孙儿", "女儿", "男儿", "红孩儿", "花儿", "虫儿", "马儿", "鸟儿", "猪儿", "猫儿",
111+
"狗儿", "少儿",
112+
]
63113
}

‎Sources/FluidAudio/TTS/KokoroAne/G2P/Mandarin/MandarinG2P.swift‎

Lines changed: 33 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -28,9 +28,10 @@ import Foundation
2828
/// encodes tone.
2929
/// 5. **Diacritic → digit** — each pinyin syllable is normalized to
3030
/// `(base, tone)` via `MandarinPinyinNormalizer`.
31-
/// 6. **Erhua merge** — `MandarinErhua.merge` folds trailing `儿`
31+
/// 6. **Erhua merge** — `MandarinErhua.merge` folds a word-final `儿`
3232
/// into the previous syllable so `小孩儿` emits a single
33-
/// r-coloured token (`ㄒㄧㄠ3ㄏㄞㄦ2`).
33+
/// r-coloured token (`ㄒㄧㄠ3ㄏㄞㄦ2`). Other characters read `er`
34+
/// (`二`, `而`, `耳`, `尔`) keep their own syllable.
3435
/// 7. **Tone sandhi** — 3+3 → 2+3, 不 / 一 contextual rules
3536
/// (`MandarinToneSandhi`).
3637
/// 8. **Pinyin → Bopomofo** — `MandarinBopomofoMap.encode` produces
@@ -118,8 +119,12 @@ public struct MandarinG2P: Sendable {
118119

119120
var output = ""
120121
var pendingSyllables: [MandarinPinyinNormalizer.Syllable] = []
122+
// The Hanzi behind `pendingSyllables`, so the erhua pass can tell
123+
// `儿` from `二` / `而` / `耳` / `尔`, which all read `er`.
124+
var pendingHanzi: [Character] = []
121125

122126
func flushPending() {
127+
defer { pendingHanzi.removeAll(keepingCapacity: true) }
123128
guard !pendingSyllables.isEmpty else { return }
124129
// Order: erhua first (it shrinks the buffer), then sandhi
125130
// operates on the merged result so 3+3 promotion sees the
@@ -141,17 +146,30 @@ public struct MandarinG2P: Sendable {
141146

142147
for seg in segments {
143148
switch seg {
144-
case .pinyin(let list, _):
145-
for py in list {
146-
pendingSyllables.append(MandarinPinyinNormalizer.normalize(py))
149+
case .pinyin(let list, let word):
150+
var syllables = list.map(MandarinPinyinNormalizer.normalize)
151+
let chars = Array(word)
152+
// Readings line up one per character; a phrase entry that
153+
// doesn't leaves no character to check, so nothing folds.
154+
if chars.count == syllables.count, let last = syllables.indices.last,
155+
syllables[last].base == "er",
156+
MandarinErhua.isSuffix(word: chars, preceding: pendingHanzi)
157+
{
158+
syllables[last].isErhuaSuffix = true
147159
}
160+
pendingSyllables.append(contentsOf: syllables)
161+
pendingHanzi.append(contentsOf: chars)
148162
case .syllables(let list):
149163
// User-lexicon pinyin tokens — already in
150164
// (base, tone) form. They join the same syllable buffer
151165
// so sandhi runs across user/dict boundaries naturally
152166
// (e.g. user word ending in tone-3 followed by dict
153-
// word starting with tone-3 → 3+3 promotion).
167+
// word starting with tone-3 → 3+3 promotion). They are
168+
// taken literally for erhua: their `er` tokens stay
169+
// unflagged, and a `儿` after them has no Hanzi context
170+
// to fold by.
154171
pendingSyllables.append(contentsOf: list)
172+
pendingHanzi.removeAll(keepingCapacity: true)
155173
case .punctuation(let s):
156174
// Sandhi never crosses punctuation; emit accumulated
157175
// syllables first.
@@ -204,12 +222,12 @@ public struct MandarinG2P: Sendable {
204222
// MARK: - Segmentation
205223

206224
enum Segment {
207-
/// Diacritic-form pinyin syllables. `hanziCount` is the number
208-
/// of Hanzi consumed from the input — needed by the polyphone
209-
/// pass to know whether a segment is a single-char fallback
210-
/// (eligible for g2pW override) or a phrase match (which the
211-
/// dict already context-disambiguated).
212-
case pinyin([String], hanziCount: Int)
225+
/// Diacritic-form pinyin syllables for `word`, the Hanzi consumed
226+
/// from the input. A one-character `word` is a single-char fallback
227+
/// (eligible for g2pW override); a longer one is a phrase match,
228+
/// which the dict already context-disambiguated. The text itself
229+
/// is what lets the erhua pass tell `儿` from other `er` readings.
230+
case pinyin([String], word: String)
213231
/// Pre-parsed syllables (user-lexicon source).
214232
case syllables([MandarinPinyinNormalizer.Syllable])
215233
case punctuation(String) // ASCII punctuation passthrough.
@@ -271,7 +289,7 @@ public struct MandarinG2P: Sendable {
271289
for word in words {
272290
let wordCharCount = word.count
273291
if wordCharCount >= 2, let pinyin = dict.phrases[word] {
274-
segments.append(.pinyin(pinyin, hanziCount: wordCharCount))
292+
segments.append(.pinyin(pinyin, word: word))
275293
offsetInRun += wordCharCount
276294
continue
277295
}
@@ -289,7 +307,7 @@ public struct MandarinG2P: Sendable {
289307
PolyphoneTarget(
290308
segmentIdx: segments.count, charPos: absPos))
291309
}
292-
segments.append(.pinyin([pinyin[0]], hanziCount: 1))
310+
segments.append(.pinyin([pinyin[0]], word: String(ch)))
293311
} else {
294312
// Unknown char — fall through as literal so
295313
// KokoroAneVocab can have a shot at it.
@@ -342,7 +360,7 @@ public struct MandarinG2P: Sendable {
342360
if let pinyin = dict.phrases[candidate] {
343361
flushHanziRun()
344362
flushLiteral()
345-
segments.append(.pinyin(pinyin, hanziCount: len))
363+
segments.append(.pinyin(pinyin, word: candidate))
346364
i += len
347365
matched = true
348366
break

‎Sources/FluidAudio/TTS/KokoroAne/G2P/Mandarin/MandarinPinyinNormalizer.swift‎

Lines changed: 8 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -22,11 +22,18 @@ public enum MandarinPinyinNormalizer {
2222
/// Set by `MandarinErhua.merge`; consumed by
2323
/// `MandarinBopomofoMap.encode` to append the `ㄦ` suffix.
2424
public var erhua: Bool
25+
/// Whether this `er` syllable is the erhua suffix `儿` and may fold
26+
/// into the syllable before it. Pinyin alone can't tell `儿` from
27+
/// `二` / `而` / `耳` / `尔`, so `MandarinG2P` sets this from the
28+
/// source text (see `MandarinErhua.isSuffix`); `MandarinErhua.merge`
29+
/// only folds syllables that carry it.
30+
public var isErhuaSuffix: Bool
2531

26-
public init(base: String, tone: Int, erhua: Bool = false) {
32+
public init(base: String, tone: Int, erhua: Bool = false, isErhuaSuffix: Bool = false) {
2733
self.base = base
2834
self.tone = tone
2935
self.erhua = erhua
36+
self.isErhuaSuffix = isErhuaSuffix
3037
}
3138
}
3239

0 commit comments

Comments
 (0)