To overcome the incompleteness of modeling document characteristics and the lack of theory for current document similarity models,this paper puts forward to utilize mixture language model(MLM) to evaluate document-to-document similarity.
为了克服现有文档相似性模型对文档特性拟合的不完全性和缺乏理论根据的弱点,本文在统计语言模型的基础上,提出了一种基于混合语言模型(M ixture Language Model,MLM)文档相似性计算模型。
参考来源 - 基于混合语言模型的文档相似性计算模型 in C·2,447,543篇论文数据,部分数据来源于NoteExpress
Finally, an algorithm for computing document similarity is presented, which filters abnormal information more efficiently.
在信息匹配算法方面,通过计算文档向量之间的相似度,实现网络信息的有效过滤。
This paper proposes a new method for XML document similarity computation based on the synthetical features of XML documents.
该文提出了一个新的基于综合语义的可扩展标记语言文档相似度计算方法。
In respect to the limitation of document similarity measuring based on VSM, this paper put forward an algorithm based on public substring of strings.
针对向量空间模型在文档相似度量方面的局限,提出了基于计算公共子串的文档相似度量算法。
应用推荐