###
计算机系统应用:2018,27(9):210-214
本文二维码信息
码上扫一扫!
改进特征权重的短文本聚类算法
马存1,2, 郭锐锋2, 高岑2, 孙咏2
(1.中国科学院大学, 北京 100049;2.中国科学院 沈阳计算技术研究所, 沈阳 110168)
Short Text Clustering Algorithm with Improved Feature Weight
MA Cun1,2, GUO Rui-Feng2, GAO Cen2, SUN Yong2
(1.University of Chinese Academy of Sciences, Beijing 100049, China;2.Shenyang Institute of Computing Technology, Chinese Academy of Sciences, Shenyang 110168, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 179次   下载 212
投稿时间:2018-01-27    修订日期:2018-03-07
中文摘要: 短文本的研究一直是自然语言处理领域的热门话题,由于短文本特征稀疏、用语口语化严重的特点,它的聚类模型存在维度高、主题聚焦性差、语义信息不明显的问题.针对对上述问题的研究,本文提出了一种改进特征权重的短文本聚类算法.首先,定义多因子权重规则,基于词性和符号情感分析构造综合评估函数,结合词项和文本内容相关度进行特征词选择;接着,使用Skip-gram模型(Continuous Skip-gram Model)在大规模语料中训练得到表示特征词语义的词向量;最后,利用RWMD算法计算短文本之间的相似度并将其应用K-Means算法中进行聚类.最后在3个测试集上的聚类效果表明,该算法有效提高了短文本聚类的准确率.
中文关键词: 特征权重  情感分析  词向量  RWMD距离
Abstract:Short text research has been a hot topic in the field of natural language processing. Due to the sparseness of short texts and serious colloquialisms, its clustering model has the problems of high dimensionality, poor focus of theme, and unclear semantic information. In view of the above problems, this study proposes a short text clustering algorithm with improving the feature weight. Firstly, the rules of multi-factor weight are defined, the comprehensive evaluation function is constructed based on part-of-speech and symbolic sentiment analysis, and the feature words are selected according to the relevancy between the term and the text content. Then, a word skip vector model (continuous skip-gram model) trained in large-scale corpus to obtain a word vector representing the semantic meaning of the feature words. Finally, the RWMD algorithm is used to calculate the similarity between short texts and the K-means algorithm is used to cluster them. The clustering results on the three test sets show that the algorithm effectively improves the accuracy of short text clustering.
文章编号:     中图分类号:    文献标志码:
基金项目:
引用文本:
马存,郭锐锋,高岑,孙咏.改进特征权重的短文本聚类算法.计算机系统应用,2018,27(9):210-214
MA Cun,GUO Rui-Feng,GAO Cen,SUN Yong.Short Text Clustering Algorithm with Improved Feature Weight.COMPUTER SYSTEMS APPLICATIONS,2018,27(9):210-214

用微信扫一扫

用微信扫一扫