A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese

	A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese
	Shiyu Zhou; Linhao Dong; Shuang Xu; Bo Xu
	2018
会议名称	ICONIP
会议录名称	ICONIP
期号	2018
会议日期	2018
会议地点	Siem Reap, Cambodia
摘要	The choice of modeling units is critical to automatic speech recognition (ASR) tasks. Conventional ASR systems typically choose context-dependent states (CD-states) or context-dependent phonemes (CD-phonemes) as their modeling units. However, it has been challenged by sequence-to-sequence attention-based models. On English ASR tasks, previous attempts have already shown that the modeling unit of graphemes can outperform that of phonemes by sequence-to-sequence attention-based model. In this paper, we are concerned with modeling units on Mandarin Chinese ASR tasks using sequence-to-sequence attention-based models with the Transformer. Five modeling units are explored including context-independent phonemes (CI-phonemes), syllables, words, sub-words and characters. Experiments on HKUST datasets demonstrate that the lexicon free modeling units can outperform lexicon related modeling units in terms of character error rate (CER). Among five modeling units, character based model performs best and establishes a new state-of-the-art CER of 26.64% on HKUST datasets.
关键词	Asr Multi-head Attention Modeling Units Sequence-to-sequence Transformer
学科门类	工学::计算机科学与技术（可授工学、理学学位）
收录类别	EI
文献类型	会议论文
条目标识符	http://ir.ia.ac.cn/handle/173211/41001
专题	复杂系统认知与决策实验室_听觉模型与认知计算
通讯作者	Shiyu Zhou
推荐引用方式 GB/T 7714	Shiyu Zhou,Linhao Dong,Shuang Xu,et al. A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese[C],2018.