This paper, do research on Chinese statistical language modeling based on Hidden Markov Trigram model. The main contents include gathering of single language corpus, model selection, training, smoothing and compression.
本文以隐马尔可夫Trigram模型为核心,研究中文的统计语言建模问题,包括单语语料库收集与整理、模型选择、训练、平滑、压缩等问题,并开发出一套通用的、面向对象的中文统计语言建模工具箱。
参考来源 - 中文统计自然语言处理隐马模型的研究·2,447,543篇论文数据,部分数据来源于NoteExpress
As a natural language processing tool, statistical language modeling is proved to be able to process large-scale real text.
统计语言模型作为一种自然语言处理的工具,已经被证明有能力处理大规模真实文本。
The current state of this practice employs the Unified Modeling Language (UML) as the primary modeling notation.
这些实践的当前情况是使用统一建模语言(UML)作为首选的建模符号。
The modeling language is a proprietary one, reducing intuitiveness.
建模语言也是厂商特有的,缺乏直观性。
应用推荐