A word that is not included in a Chinese segmentation lexicon is called a new word.
新词指在进行词法切分时词典中未收录的词。
As a basic component of Chinese word segmentation system, the dictionary mechanism influences the speed and the efficiency of segmentation significantly.
词典是中文自动分词的基础,分词词典机制的优劣直接影响到中文分词的速度和效率。
This paper presented a syllable segmentation method for Chinese connected speech.
本文提出一种汉语语音连接词音节分割方法。
In this paper, we first construct the system architecture, improve the Chinese text segmentation algorithm, then, by making use of domain ontology base and sentence similarity, design the system.
文章首先构造了自动答疑系统架构,改进了中文分词算法,并利用领域本体库和语句相似度设计了该系统。
Aiming at the dissatisfied effect of Chinese word segmentation to Email texts, an improved Maximum Match Based Approach is presented.
针对邮件文本分词效果较差的特点,提出采用一种改进的最大匹配法来进行中文分词的方法。
Chinese word segmentation is always the first step of subject extraction. The quality of word segmentation is effective to the quality of text subject extraction.
而主题提取是以中文分词作为第一步,分词质量直接影响到文献主题提取的质量。
The method mentioned above has been applied to the segmentation of Chinese bank check amounts and get good results.
上述方法应用于银行支票手写体大写金额的分割,取得了很好的分割效果。
Here we explore SVM for a Chinese word segmentation task, use the context attributes and rule-based attributes as the features for a sample.
本文首次使用SVM方法来完成中文分词的任务,使用上下文窗体属性和基于规则的属性对样本进行刻画。
In order to solve the problem of automatic segmentation of handwritten Chinese text, a dynamic programming-based online handwritten Chinese character segmentation method is proposed.
为解决手写汉字文本的自动切分问题,提出了一种基于动态规划的联机手写汉字分割方法。
The paper introduces the design and implementation of Chinese word segmentation system, which is based on statistic the frequency of the word.
论文介绍了一个基于词频统计的中文分词系统的设计和实现。
Therefore, the primary issue of Chinese information processing, that is, to a sentence to separate words, this is the Chinese word segmentation problem.
因此中文信息处理的首要问题,就是要将句子中一个个词给分离出来,这就是中文分词问题。
Search engine technology related to natural language understanding, Chinese word segmentation, artificial intelligence, machine learning and so on.
搜索引擎的技术涉及到自然语言理解、中文分词、人工智能、机器学习等学科。
Chinese automatic segmentation is one of the most difficult problems in computer Chinese information disposal and the key problem that document content analysis must resolve.
汉语自动分词是计算机中文信息处理中的难题,也是文献内容分析中必须解决的关键问题之一。
This paper proposes a statistical method to solve overlapped ambiguity in Chinese words' segmentation.
该文利用一种统计的方法来解决交集型歧义字段的切分。
The segmentation of unconstrained handwritten Chinese words is the key problem of handwritten Chinese words recognition. It is also the difficulty of current segmentation-recognition truss.
非限定手写汉字分割问题是手写汉字识别的关键问题,也是目前的分割-识别框架中的难点。
The design and implementation of the Interface for Database Query in Chinese (IDCQ). The system includes regular word segmentation subsystem and object semantic analysis subsystem.
设计和实现了汉语数据库自然语言查询接口系统(IDCQ),系统包括正则分词子系统和对象语义解析子系统;
This paper presents a learning method to auto ma tically acquire segmentation knowledge from Chinese corpus.
文章描述了一种从熟语料中自动获取文本切分知识的机器学习的方法。
In this paper, the dictionary mechanism is dynamic TRIE tree, and we have designed the Chinese word segmentation dictionary. The dictionary USES less memory.
论文采用动态TRIE索引树的词典机制,设计并实现了汉语分词词典,有效地减少了词典空间。
Knowledge of Chinese words automatic segmentation can raise the precision of automatic segmentation, and it can satisfy high precision requirements.
使用自动分词知识可以进一步提高自动切分精度,满足高标准的需求。
Combinational ambiguity is a challenging issue in Chinese word segmentation in that its disambiguation depends on the contextual information.
组合型歧义切分字段一直是汉语自动分词的难点,难点在于消歧依赖其上下文语境信息。
A fast algorithm for generating Chinese word segmentation digraph was given.
给出了一种汉语分词有向图的快速生成算法。
This paper puts forward a new algorithm about automatic Chinese text segmentation based on Chinese characters string frequency and length descending.
提出了一种基于汉字串频度及串长度递减的中文文本自动切分算法。
To extend word segmentation repository and enhance word segmentation capacity, a Chinese word segmentation system based on automatic learning is proposed in this paper.
为扩展分词知识库,提高自动分词能力,本文提出了一种基于自学习机制的汉语自动分词系统。
Overlapping ambiguity is a major type of ambiguity in Chinese word segmentation.
交集型分词歧义是汉语自动分词中的主要歧义类型之一。
The Chinese words segmentation and labeling are basis of the Chinese language processing.
汉语的分词及词性标注是汉语语言处理的基础。
The former includes Chinese word segmentation, part - of - speech tagging, pinyin tagging, named entity recognition, new word detection, syntactic parsing, word sense disambiguation, etc.
前者涉及到词法、句法、语义分析,包括汉语分词、词性标注、注音、命名实体识别、新词发现、句法分析、词义消歧等。
The former includes Chinese word segmentation, part-of-speech tagging, pinyin tagging, named entity recognition, new word detection, syntactic parsing, word sense disambiguation, etc .
前者涉及到词法、句法、语义分析,包括汉语分词、词性标注、注音、命名实体识别、新词发现、句法分析、词义消歧等。
Automatic word segmentation of modern Chinese text is the base of Chinese information processing. So a general purpose application interface for word segmentation is important.
现代汉语文本自动分词是中文信息处理的重要基石,为此提供一个通用的分词接口是非常重要的。
Automatic word segmentation of modern Chinese text is the base of Chinese information processing. So a general purpose application interface for word segmentation is important.
现代汉语文本自动分词是中文信息处理的重要基石,为此提供一个通用的分词接口是非常重要的。
应用推荐