基于后缀数组的分词技术①

微信公众号

网站二维码

首页 > 过刊浏览>2010年第19卷第8期 >229-230

基于后缀数组的分词技术①增强出版
DOI:
                        
                    
作者:
                        
                        
                    
作者单位:
作者简介:
通讯作者:
中图分类号:
基金项目:曲靖师范学院基金(2008QN007);云南省教育厅研究课题(09C0188)

Word Segment Based on Suffix Array

Author:

Affiliation:

Fund Project:

摘要

图/表

访问统计

参考文献

相似文献

引证文献

增强出版

文章评论

摘要:

中文分词技术是机器翻译、分类、搜索引擎以及信息检索的基础，但是，互联网上不断出现的新词严重影响了分词的性能，为了提高新词的识别率，建立待分词内容的后缀数组，然后计算其公共前缀共同出现的次数，采用阈值对其进行过滤筛选出候选词语，实验结果表明，该方法在新词识别方面有一定的优势。

Abstract:

Chinese word segmentation technology is the basis of machine translation, classification, search engines, as well as information retrieval. But the Internet emerging new words have seriously affected the performance of word segmentation. To improve the recognition rate of new words, suffix array is used in this paper, and the number of length of common prefix is calculated. The candidates on their words are filtered out by the threshold. Experimental results show that the new word recognition method has advantages.

参考文献

相似文献

引证文献

引用本文

任雪利,代余彪.基于后缀数组的分词技术①.计算机系统应用,2010,19(8):229-230

复制

文章指标

点击次数:
下载次数:
HTML阅读次数:
引用次数:

历史

收稿日期:2009-12-04
最后修改日期:2010-01-18
录用日期:
在线发布日期:
出版日期:

微信公众号

网站二维码

引用本文

分享

文章指标

历史