In the word layer, used Statistics-based Bigram language model, it is suitable for Uyghur voice features.
在词层上,本文使用了适合于维吾尔语语音特征的语言模型——基于统计的二元文法语言模型。
参考来源 - 基于HTK的维吾尔语连续语音识别研究·2,447,543篇论文数据,部分数据来源于NoteExpress
In those texts, we select bigram as feature after Chinese word segmentation, deleting stop word and other process.
在筛选出的文本中,经过分词、去除停用词等处理后,选取二元词串作为特征;
Found a large number of high-degree overlapped bigrams and high-degree biased bigrams existing in bigram feature set.
发现特征集中存在大量高度重叠特征和高度偏差特征。
Secondly, this paper presents a hierarchical text filtering approach based on bigram in the off-line filtering module.
其次,针对离线过滤,本文提出了一种基于二元模型的分层文本过滤方法。
应用推荐