diff --git a/Documentation/README.md b/Documentation/README.md index d6722324d..96829c42f 100644 --- a/Documentation/README.md +++ b/Documentation/README.md @@ -38,6 +38,7 @@ ## Text-to-Speech (TTS) - [Kokoro ANE v3 and legacy runtime](TTS/KokoroAne.md) +- [Kokoro English dictionary and G2P comparison](TTS/KokoroEnglishG2P.md) - [Kokoro v3 latency and compute-policy measurements](TTS/Benchmarks.md#kokoro-ane-v3-selected-m5-pro-measurements) - [Kokoro ANE placement and telemetry](ANE_Profiler.md#kokoro-ane-v3) - [PocketTTS](TTS/PocketTTS.md) diff --git a/Documentation/TTS/KokoroAne.md b/Documentation/TTS/KokoroAne.md index 0deb6916b..33d27a96d 100644 --- a/Documentation/TTS/KokoroAne.md +++ b/Documentation/TTS/KokoroAne.md @@ -136,6 +136,9 @@ For multi-voice / SSML / long-form, use `PocketTtsSynthesizer` or ## Variants +For English pronunciation lookup, fallback behavior and a bounded word-level +comparison, see [Kokoro English dictionary and G2P comparison](KokoroEnglishG2P.md). + The 7-stage chain is language-agnostic by construction (input ids, voice slices, and per-stage I/O contracts are identical across variants). Only the embedding vocab, HF subdirectory, voice-file layout, default voice, and the diff --git a/Documentation/TTS/KokoroEnglishG2P.md b/Documentation/TTS/KokoroEnglishG2P.md new file mode 100644 index 000000000..ff5379e62 --- /dev/null +++ b/Documentation/TTS/KokoroEnglishG2P.md @@ -0,0 +1,120 @@ +# Kokoro English pronunciation lookup and G2P fallback + +Kokoro's English frontend uses pronunciation overrides and the Misaki lexicon before +falling back to the BART grapheme-to-phoneme (G2P) model. A 40-word local comparison +illustrates why: the dictionary handles several irregular spellings that forced +BART inference mispronounces, while BART supplies pronunciations for dictionary misses. +Returning a pronunciation does not establish that it is correct. + +This is a selected word-level diagnostic, not an overall accuracy or TTS quality +benchmark. It compares **dictionary-only lookup with forced BART inference**. +It does not compare the current complete frontend against BART: current `main` +already adds initialism, compound and possessive rules before fallback. + +## Results + +| Measure | Result | +| --- | --- | +| Words selected before inference | 40 | +| Dictionary entries found | 34 of 40 | +| Nonempty BART outputs | 40 of 40 | +| Identical raw outputs on dictionary-covered words | 25 of 34 | +| Words with a CMU reference | 35 of 40 | +| BART inference failures | 0 | + +The six dictionary misses were `API`, `CPU`, `GPU`, `GitHub`, `Neuralink` and +`Kokoro`. These are misses in the tested lexicon cache, not necessarily in every +Misaki release. Full inputs, outputs, reference alternatives and artifact hashes +are in [the result snapshot](Results/KokoroEnglishG2P.json). + +Both methods returned the same phonemes for words such as `phone`, `world`, +`people`, `water`, `computer`, `yacht`, `choir`, `queue`, `island`, `debt` and +`algorithm`. Raw agreement is not accuracy; valid stress or reduced-vowel +differences can also produce different strings. + +### Selected differences + +These are actual output symbols, not listening-test transcriptions. Misaki uses +`A` for /eɪ/, `I` for /aɪ/ and `O` for /oʊ/. The `ˈ` mark denotes primary stress. + +| Word | Dictionary-only output | Forced BART output | Reference comparison | +| --- | --- | --- | --- | +| colonel | `kˈɜɹnᵊl` | `kˈɑlənᵊl` | Dictionary matches CMU's “kernel” pronunciation; BART introduces an extra syllable and changes the stressed vowel. | +| Wednesday | `wˈɛnzdˌA` | `wˈɛdnzdˌA` | BART inserts a `d` absent from both CMU alternatives. | +| inference | `ˈɪnfəɹəns` | `ɪnfˈɪɹəns` | BART shifts primary stress and changes the middle vowel relative to CMU. | +| NASA | `nˈæsə` | `nˈɑsə` | BART changes the stressed vowel relative to CMU. | +| API | Missing | `ˈæpi` | BART does not produce CMU's A-P-I initialism. | +| CPU | Missing | `spjˈu` | BART does not produce CMU's C-P-U initialism. | +| GitHub | Missing | `ɡˈɪθʌb` | BART produces /θ/ rather than CMU's /t h/ sequence. | + +References come from the [CMU Pronouncing Dictionary](https://github.com/cmusphinx/cmudict/tree/74790861f652b15e4ac49015a90074ad62a27690), +maintained by Carnegie Mellon University's Speech Group (see the retained +[CMUdict license](Results/CMUDict-LICENSE.txt)). All listed alternatives +are retained in the snapshot. Its lowercase `ai` entry includes both /aɪ/ and +the letter-name reading, so acronym intent must be considered separately. + +### What current callers receive + +The [current English frontend](../../Sources/FluidAudio/TTS/KokoroAne/G2P/English/KokoroAneEnglishPhonemizer.swift) +does more than raw dictionary lookup: + +- Custom pronunciations take precedence. +- Explicit letter-name overrides handle uppercase `AI` and `US`. +- Known words and acronyms such as `NASA` resolve from the lexicon. +- Unknown ASCII all-caps tokens of two to five letters are spelled using + per-letter lexicon entries, when those entries are available. +- Compound and possessive handling precedes BART fallback. + +Consequently, the forced-BART `API` and `CPU` rows are **not evidence that current +normal text synthesis mispronounces those initialisms**. With the required letter +entries loaded, the current rules also cover `GPU`. This behavior comes from code +inspection; the 40-word experiment did not run the current full frontend. + +For an uncovered application-specific name, use `setEnglishCustomLexicon(_:)` with +the intended pronunciation in Kokoro's supported phoneme vocabulary. Keep BART as +coverage for unresolved words rather than assuming every returned sequence is correct. + +## Method + +The frozen selection contains ten common words, ten irregular spellings, ten technical +words and ten names/acronyms. One real BART prediction was made per word; no TTS +waveforms were synthesized. Execution used an Apple M5 Pro, macOS 27.0, Swift 6.2.3 +debug build and the production CPU-only G2P configuration. + +The source checkout was `c9cf87a6e458cf0afb34aabdfa2b1faede1f7360`, with unrelated +working-tree changes. The three evaluated source files matched that commit: + +1. `LexiconAssetCache` loaded the actual cached `us_lexicon_cache.json`, filtering + entries to the English `ANE/vocab.json` token set. +2. That revision's `KokoroAneEnglishPhonemizer` performed dictionary lookup with no + custom overrides and a fallback returning `nil`. A failed lookup was recorded + as missing, rather than using a generated pronunciation as a reference. +3. `G2PModel.shared.phonemize(word:)` ran for every word, using the same lowercase + normalization as the deployed fallback, including for uppercase inputs. + +At comparison base `0b0fa2ad710843d3f40af885caf9a86d1fa49c26`, `G2PModel.swift` +is byte-identical to the evaluated file. The English frontend has changed since +the evaluated revision, as described above. Cached asset bytes are identified by +SHA-256; their upstream release revision was not established. + +To repeat the comparison, freeze a word manifest, load matching lexicon/vocabulary +and model assets, run dictionary-only lookup and forced BART on each word, and retain +both outputs. Use an internal debug test/probe to access these types; this is not a +new public benchmark command. The existing `g2p-benchmark` command evaluates the +separate Charsiu ByT5 model and does not reproduce this English BART comparison. + +## Limits and interpretation + +- Coverage and raw agreement are reported separately. No phoneme error rate or + overall pronunciation accuracy percentage was computed. +- Words were deliberately selected, not sampled from measured user traffic. + Training-data overlap with Misaki, BART or CMU was not ruled out. +- `quantization`, `spectrogram`, `GPU`, `Neuralink` and `Kokoro` have no reference + in the pinned CMU dictionary. Nonempty output does not validate their pronunciation. +- Word-level results do not measure sentence context, homographs, connected-speech + prosody, intelligibility of generated audio, naturalness or voice identity. +- Differences may originate in the trained model, conversion or decoding. The + original PyTorch checkpoint was not evaluated, so this is not a claim about + BART architectures in general or a demonstrated conversion defect. +- Single debug-call timings do not support a latency ranking. For text-to-audio + measurements, see [TTS Benchmarks](Benchmarks.md). diff --git a/Documentation/TTS/Results/CMUDict-LICENSE.txt b/Documentation/TTS/Results/CMUDict-LICENSE.txt new file mode 100644 index 000000000..81a0ae8ec --- /dev/null +++ b/Documentation/TTS/Results/CMUDict-LICENSE.txt @@ -0,0 +1,33 @@ +Copyright (C) 1993-2015 Carnegie Mellon University. All rights reserved. + +Redistribution and use in source and binary forms, with or without +modification, are permitted provided that the following conditions +are met: + +1. Redistributions of source code must retain the above copyright + notice, this list of conditions and the following disclaimer. + The contents of this file are deemed to be source code. + +2. Redistributions in binary form must reproduce the above copyright + notice, this list of conditions and the following disclaimer in + the documentation and/or other materials provided with the + distribution. + +This work was supported in part by funding from the Defense Advanced +Research Projects Agency, the Office of Naval Research and the National +Science Foundation of the United States of America, and by member +companies of the Carnegie Mellon Sphinx Speech Consortium. We acknowledge +the contributions of many volunteers to the expansion and improvement of +this dictionary. + +THIS SOFTWARE IS PROVIDED BY CARNEGIE MELLON UNIVERSITY ``AS IS'' AND +ANY EXPRESSED OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, +THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR +PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL CARNEGIE MELLON UNIVERSITY +NOR ITS EMPLOYEES BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, +SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT +LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, +DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY +THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT +(INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE +OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. diff --git a/Documentation/TTS/Results/KokoroEnglishG2P.json b/Documentation/TTS/Results/KokoroEnglishG2P.json new file mode 100644 index 000000000..66438358d --- /dev/null +++ b/Documentation/TTS/Results/KokoroEnglishG2P.json @@ -0,0 +1,502 @@ +{ + "scope": "40 selected English words; dictionary-only vs forced BART Core ML output; not the current full frontend or an accuracy benchmark", + "source_commit": "c9cf87a6e458cf0afb34aabdfa2b1faede1f7360", + "source_note": "The working checkout had unrelated edits. G2PModel.swift, KokoroAneEnglishPhonemizer.swift and LexiconAssetCache.swift matched this commit. G2PModel.swift also matches main at 0b0fa2ad710843d3f40af885caf9a86d1fa49c26; the English frontend has since gained initialism, compound and possessive handling.", + "host": { + "chip": "Apple M5 Pro", + "os": "macOS-27.0-arm64-arm-64bit", + "build": "Swift 6.2.3 debug", + "compute_units": "cpuOnly" + }, + "reference": { + "reference_url": "https://raw.githubusercontent.com/cmusphinx/cmudict/74790861f652b15e4ac49015a90074ad62a27690/cmudict.dict", + "reference_revision": "74790861f652b15e4ac49015a90074ad62a27690", + "reference_sha256": "81917843c7f44ce2b094ac63873c2c7a4cf802040792c455ba3ca406891c3d22", + "reference_policy": "Manual inspection of pinned CMU ARPABET alternatives; no automatic phoneme accuracy scoring. Raw dictionary/BART agreement uses exact strings, including stress.", + "license": "CMUDict-LICENSE.txt" + }, + "sha256": { + "G2PModel.swift": "988c4ec3e2aac2d9a1580c2ab88b15e87c696ed8c458a7df64bbbfcaaa803ee2", + "KokoroAneEnglishPhonemizer.swift": "c5499b83a81dca6241ed5c52b2401e9a811b13657ec78ec791cb8b88ac6b6d85", + "us_lexicon_cache.json": "6b36ba313202227d6914ad32cd684a0304bd2757e9ec4158ea7bc36ec40e224e", + "g2p_vocab.json": "295ed64b86c2820cd665b0602ae50c6947c0e82ac643082873e0be87dca282ce", + "LexiconAssetCache.swift": "649283e835dcaf5ddfa1e5d9a113e75fdb1db58cd5d4a1cf60e8de150427a79d", + "ANE/vocab.json": "8d65b0188b77eafc60751dac42bbac7ab5f5685074af44db91d1877b42dc1d7c" + }, + "model_file_sha256": { + "G2PEncoder.mlmodelc/analytics/coremldata.bin": "cf7fbd7e7a65529b2d2bf3941e458a0ab6dff7a298bf48e205a1727c81c26a99", + "G2PEncoder.mlmodelc/coremldata.bin": "0f14d46ca9fd06c68b4717294575b2b99449e67d40b7a2c56f926bf05cd90b11", + "G2PEncoder.mlmodelc/metadata.json": "c8e0cfd7f494ac1b3662ff8f1914b2b45f79ffb2791724cdc0576981996732e1", + "G2PEncoder.mlmodelc/model.mil": "8c617e569f37286b056dad800d862dc145be9a95fa9ed43857bb646ba199d7da", + "G2PEncoder.mlmodelc/weights/weight.bin": "6926bcd2827d21fec82839487b987e06f85fd8a6a5bb896bc4f6062461d014ec", + "G2PDecoder.mlmodelc/analytics/coremldata.bin": "dbf1767747fdc188222d467a45b04608b396c76c71db3abb18a5fb3680ef9827", + "G2PDecoder.mlmodelc/coremldata.bin": "607e960f19b4d9a30317a5a11869fcce84b300a909fcab2cc756c0d98e2dacd9", + "G2PDecoder.mlmodelc/metadata.json": "e54e98484fd60d26f22fd3c4e7fe87b0d92a5d2de1f958cc3c4bb36d4ae06a44", + "G2PDecoder.mlmodelc/model.mil": "fe647c598e0d9454d360b8ee49a59ae57ca147fc5330863ba84ccb90dce482ad", + "G2PDecoder.mlmodelc/weights/weight.bin": "cbaeb4e743359f607ab161af0c6d8a817462fdaec622ee788ef8ef952c5f8214" + }, + "summary": { + "words": 40, + "dictionary_covered": 34, + "dictionary_missing": [ + "API", + "CPU", + "GPU", + "GitHub", + "Neuralink", + "Kokoro" + ], + "bart_nonempty": 40, + "bart_errors": [], + "raw_identical_on_dictionary_covered": 25, + "metric_scope": "Coverage and literal output agreement, not accuracy or held-out generalization", + "with_cmu_reference": 35, + "no_cmu_reference": [ + "quantization", + "spectrogram", + "GPU", + "Neuralink", + "Kokoro" + ] + }, + "cases": [ + { + "word": "phone", + "category": "common", + "normalizedInput": "phone", + "dictionary": "fˈOn", + "bart": "fˈOn", + "error": null, + "cmu_arpabet_variants": [ + "F OW1 N" + ] + }, + { + "word": "hello", + "category": "common", + "normalizedInput": "hello", + "dictionary": "həlˈO", + "bart": "hˈɛlO", + "error": null, + "cmu_arpabet_variants": [ + "HH AH0 L OW1", + "HH EH0 L OW1" + ] + }, + { + "word": "world", + "category": "common", + "normalizedInput": "world", + "dictionary": "wˈɜɹld", + "bart": "wˈɜɹld", + "error": null, + "cmu_arpabet_variants": [ + "W ER1 L D" + ] + }, + { + "word": "people", + "category": "common", + "normalizedInput": "people", + "dictionary": "pˈipᵊl", + "bart": "pˈipᵊl", + "error": null, + "cmu_arpabet_variants": [ + "P IY1 P AH0 L" + ] + }, + { + "word": "water", + "category": "common", + "normalizedInput": "water", + "dictionary": "wˈɔɾəɹ", + "bart": "wˈɔɾəɹ", + "error": null, + "cmu_arpabet_variants": [ + "W AO1 T ER0" + ] + }, + { + "word": "computer", + "category": "common", + "normalizedInput": "computer", + "dictionary": "kəmpjˈuɾəɹ", + "bart": "kəmpjˈuɾəɹ", + "error": null, + "cmu_arpabet_variants": [ + "K AH0 M P Y UW1 T ER0" + ] + }, + { + "word": "window", + "category": "common", + "normalizedInput": "window", + "dictionary": "wˈɪndO", + "bart": "wˈɪndO", + "error": null, + "cmu_arpabet_variants": [ + "W IH1 N D OW0" + ] + }, + { + "word": "voice", + "category": "common", + "normalizedInput": "voice", + "dictionary": "vˈYs", + "bart": "vˈYs", + "error": null, + "cmu_arpabet_variants": [ + "V OY1 S" + ] + }, + { + "word": "speech", + "category": "common", + "normalizedInput": "speech", + "dictionary": "spˈiʧ", + "bart": "spˈiʧ", + "error": null, + "cmu_arpabet_variants": [ + "S P IY1 CH" + ] + }, + { + "word": "language", + "category": "common", + "normalizedInput": "language", + "dictionary": "lˈæŋɡwɪʤ", + "bart": "lˈæŋɡwɪʤ", + "error": null, + "cmu_arpabet_variants": [ + "L AE1 NG G W AH0 JH", + "L AE1 NG G W IH0 JH" + ] + }, + { + "word": "colonel", + "category": "irregular", + "normalizedInput": "colonel", + "dictionary": "kˈɜɹnᵊl", + "bart": "kˈɑlənᵊl", + "error": null, + "cmu_arpabet_variants": [ + "K ER1 N AH0 L" + ] + }, + { + "word": "yacht", + "category": "irregular", + "normalizedInput": "yacht", + "dictionary": "jˈɑt", + "bart": "jˈɑt", + "error": null, + "cmu_arpabet_variants": [ + "Y AA1 T" + ] + }, + { + "word": "choir", + "category": "irregular", + "normalizedInput": "choir", + "dictionary": "kwˈIəɹ", + "bart": "kwˈIəɹ", + "error": null, + "cmu_arpabet_variants": [ + "K W AY1 ER0" + ] + }, + { + "word": "queue", + "category": "irregular", + "normalizedInput": "queue", + "dictionary": "kjˈu", + "bart": "kjˈu", + "error": null, + "cmu_arpabet_variants": [ + "K Y UW1" + ] + }, + { + "word": "island", + "category": "irregular", + "normalizedInput": "island", + "dictionary": "ˈIlənd", + "bart": "ˈIlənd", + "error": null, + "cmu_arpabet_variants": [ + "AY1 L AH0 N D" + ] + }, + { + "word": "debt", + "category": "irregular", + "normalizedInput": "debt", + "dictionary": "dˈɛt", + "bart": "dˈɛt", + "error": null, + "cmu_arpabet_variants": [ + "D EH1 T" + ] + }, + { + "word": "subtle", + "category": "irregular", + "normalizedInput": "subtle", + "dictionary": "sˈʌɾᵊl", + "bart": "sˈʌɾᵊl", + "error": null, + "cmu_arpabet_variants": [ + "S AH1 T AH0 L" + ] + }, + { + "word": "receipt", + "category": "irregular", + "normalizedInput": "receipt", + "dictionary": "ɹəsˈit", + "bart": "ɹəsˈit", + "error": null, + "cmu_arpabet_variants": [ + "R IH0 S IY1 T", + "R IY0 S IY1 T" + ] + }, + { + "word": "Wednesday", + "category": "irregular", + "normalizedInput": "wednesday", + "dictionary": "wˈɛnzdˌA", + "bart": "wˈɛdnzdˌA", + "error": null, + "cmu_arpabet_variants": [ + "W EH1 N Z D IY0", + "W EH1 N Z D EY2" + ] + }, + { + "word": "enough", + "category": "irregular", + "normalizedInput": "enough", + "dictionary": "ɪnˈʌf", + "bart": "ɪnˈʌf", + "error": null, + "cmu_arpabet_variants": [ + "IH0 N AH1 F", + "IY0 N AH1 F" + ] + }, + { + "word": "neural", + "category": "technical", + "normalizedInput": "neural", + "dictionary": "nˈʊɹᵊl", + "bart": "nˈʊɹəl", + "error": null, + "cmu_arpabet_variants": [ + "N UH1 R AH0 L", + "N Y UH1 R AH0 L" + ] + }, + { + "word": "network", + "category": "technical", + "normalizedInput": "network", + "dictionary": "nˈɛtwˌɜɹk", + "bart": "nˈɛtwˌɜɹk", + "error": null, + "cmu_arpabet_variants": [ + "N EH1 T W ER2 K" + ] + }, + { + "word": "algorithm", + "category": "technical", + "normalizedInput": "algorithm", + "dictionary": "ˈælɡəɹˌɪðəm", + "bart": "ˈælɡəɹˌɪðəm", + "error": null, + "cmu_arpabet_variants": [ + "AE1 L G ER0 IH2 DH AH0 M" + ] + }, + { + "word": "inference", + "category": "technical", + "normalizedInput": "inference", + "dictionary": "ˈɪnfəɹəns", + "bart": "ɪnfˈɪɹəns", + "error": null, + "cmu_arpabet_variants": [ + "IH1 N F ER0 AH0 N S" + ] + }, + { + "word": "quantization", + "category": "technical", + "normalizedInput": "quantization", + "dictionary": "kwˌɑntəzˈAʃən", + "bart": "kwˌɑntəzˈAʃən", + "error": null, + "cmu_arpabet_variants": [] + }, + { + "word": "spectrogram", + "category": "technical", + "normalizedInput": "spectrogram", + "dictionary": "spˈɛktɹəɡɹˌæm", + "bart": "spˈɛktɹəɡɹˌæm", + "error": null, + "cmu_arpabet_variants": [] + }, + { + "word": "phoneme", + "category": "technical", + "normalizedInput": "phoneme", + "dictionary": "fˈOnˌim", + "bart": "fˈOnˌim", + "error": null, + "cmu_arpabet_variants": [ + "F OW1 N IY0 M" + ] + }, + { + "word": "decoder", + "category": "technical", + "normalizedInput": "decoder", + "dictionary": "dˌikˈOdəɹ", + "bart": "dᵻkˈOdəɹ", + "error": null, + "cmu_arpabet_variants": [ + "D IH0 K OW1 D ER0" + ] + }, + { + "word": "latency", + "category": "technical", + "normalizedInput": "latency", + "dictionary": "lˈAtᵊnsi", + "bart": "lˈAtᵊnsi", + "error": null, + "cmu_arpabet_variants": [ + "L EY1 T AH0 N S IY0" + ] + }, + { + "word": "synthesis", + "category": "technical", + "normalizedInput": "synthesis", + "dictionary": "sˈɪnθəsɪs", + "bart": "sˈɪnθəsɪs", + "error": null, + "cmu_arpabet_variants": [ + "S IH1 N TH AH0 S AH0 S" + ] + }, + { + "word": "AI", + "category": "names_acronyms", + "normalizedInput": "ai", + "dictionary": "ˈAˌI", + "bart": "ˈI", + "error": null, + "cmu_arpabet_variants": [ + "AY1", + "EY1 AY1" + ] + }, + { + "word": "API", + "category": "names_acronyms", + "normalizedInput": "api", + "dictionary": null, + "bart": "ˈæpi", + "error": null, + "cmu_arpabet_variants": [ + "EY2 P IY2 AY1" + ] + }, + { + "word": "CPU", + "category": "names_acronyms", + "normalizedInput": "cpu", + "dictionary": null, + "bart": "spjˈu", + "error": null, + "cmu_arpabet_variants": [ + "S IY2 P IY2 Y UW1" + ] + }, + { + "word": "GPU", + "category": "names_acronyms", + "normalizedInput": "gpu", + "dictionary": null, + "bart": "ʤˈipˌu", + "error": null, + "cmu_arpabet_variants": [] + }, + { + "word": "NASA", + "category": "names_acronyms", + "normalizedInput": "nasa", + "dictionary": "nˈæsə", + "bart": "nˈɑsə", + "error": null, + "cmu_arpabet_variants": [ + "N AE1 S AH0" + ] + }, + { + "word": "GitHub", + "category": "names_acronyms", + "normalizedInput": "github", + "dictionary": null, + "bart": "ɡˈɪθʌb", + "error": null, + "cmu_arpabet_variants": [ + "G IH1 T HH AH0 B" + ] + }, + { + "word": "Neuralink", + "category": "names_acronyms", + "normalizedInput": "neuralink", + "dictionary": null, + "bart": "nˈʊɹəlˌɪŋk", + "error": null, + "cmu_arpabet_variants": [] + }, + { + "word": "Kokoro", + "category": "names_acronyms", + "normalizedInput": "kokoro", + "dictionary": null, + "bart": "kəkˈɔɹO", + "error": null, + "cmu_arpabet_variants": [] + }, + { + "word": "Toronto", + "category": "names_acronyms", + "normalizedInput": "toronto", + "dictionary": "təɹˈɑntO", + "bart": "təɹˈɑntO", + "error": null, + "cmu_arpabet_variants": [ + "T ER0 AA1 N T OW0", + "T AO0 R AA1 N T OW0" + ] + }, + { + "word": "Alex", + "category": "names_acronyms", + "normalizedInput": "alex", + "dictionary": "ˈælɪks", + "bart": "ˈælˌɛks", + "error": null, + "cmu_arpabet_variants": [ + "AE1 L AH0 K S" + ] + } + ] +}