← 论文 18

PSP:面向 Indic 文本到语音的可解释逐维度口音基准

scored
↗ 原文 ↗ PDF · Hugging Face Daily
📋 摘要 ⭐ Indic TTS 口音基准,与 SE for AI、ML 测试、公平性测试方向几乎无关,且与已标记不感兴趣的 Praxy Voice 高度同源。 PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
中文
针对标准 TTS 评测指标(WER、CER、MOS、UTMOS)无法量化口音这一问题,本文提出 PSP (Phoneme Substitution Profile),一个面向 Indic 语言 TTS 的可解释、按音系维度划分的口音基准。研究问题聚焦于合成语音可能在可懂度与自然度上得分良好,却在目标语言的音位特征(如卷舌、送气、元音长短、泰米尔语 retroflex approximant 'zha')上呈现非母语口音。方法上,PSP 将口音分解为六个互补维度:retroflex collapse rate (RR)、aspiration fidelity (AF)、vowel-length fidelity (LF)、Tamil-zha fidelity (ZF)、Frechet Audio Distance (FAD) 与 prosodic signature divergence (PSD);前四项通过 forced alignment 结合基于 Wav2Vec2-XLS-R 第 9 层嵌入的母语者中心声学探针测量,后两项为语料级分布距离。作者在 Hindi、Telugu、Tamil 上对 ElevenLabs v3、Cartesia Sonic-3、Sarvam Bulbul、Indic Parler-TTS 四个商业及开源系统进行基准测试,并加入 Praxy Voice 及 Telugu 上的 R5→R6 案例研究。主要结果:(i) retroflex collapse 随音系难度单调增长(Hindi ~1%、Telugu ~40%、Tamil ~68%);(ii) PSP 排序与 WER 排序不一致,WER 领先的商业系统未必在卷舌或韵律保真度上领先;(iii) 没有单一系统在六个维度上达到 Pareto 最优。与现有评测的差异在于其以可解释的、按音位维度的方式量化口音,并公开母语参考中心、嵌入、韵律矩阵、golden sets 与评分代码。
English abstract
Standard text-to-speech (TTS) evaluation measures intelligibility (WER, CER) and overall naturalness (MOS, UTMOS) but does not quantify accent. A synthesiser may score well on all four yet sound non-native on features that are phonemic in the target language. For Indic languages, these features include retroflex articulation, aspiration, vowel length, and the Tamil retroflex approximant (letter zha). We present PSP, the Phoneme Substitution Profile, an interpretable, per-phonological-dimension accent benchmark for Indic TTS. PSP decomposes accent into six complementary dimensions: retroflex collapse rate (RR), aspiration fidelity (AF), vowel-length fidelity (LF), Tamil-zha fidelity (ZF), Frechet Audio Distance (FAD), and prosodic signature divergence (PSD). The first four are measured via forced alignment plus native-speaker-centroid acoustic probes over Wav2Vec2-XLS-R layer-9 embeddings; the latter two are corpus-level distributional distances. In this v1 we benchmark four commercial and open-source systems (ElevenLabs v3, Cartesia Sonic-3, Sarvam Bulbul, Indic Parler-TTS) on Hindi, Telugu, and Tamil pilot sets, with a fifth system (Praxy Voice) included on all three languages, plus an R5->R6 case study on Telugu. Three findings: (i) retroflex collapse grows monotonically with phonological difficulty Hindi < Telugu < Tamil (~1%, ~40%, ~68%); (ii) PSP ordering diverges from WER ordering -- commercial WER-leaders do not uniformly lead on retroflex or prosodic fidelity; (iii) no single system is Pareto-optimal across all six dimensions. We release native reference centroids (500 clips per language), 1000-clip embeddings for FAD, 500-clip prosodic feature matrices for PSD, 300-utterance golden sets per language, scoring code under MIT, and centroids under CC-BY. Formal MOS-correlation is deferred to v2; v1 reports five internal-consistency signals plus a native-audio sanity check.
加载中…
点文件 → 加为 tab;按 Esc 关闭
Esc
输入名称、URL、路径或标签...
选择 Enter 打开 Enter 新标签