###
DOI:
计算机系统应用英文版:2013,22(11):195-199
本文二维码信息
码上扫一扫!
基于Hadoop平台的XML文档重复数据检测
(暨南大学 信息科学技术学院学院, 广州 510632)
XML Data Duplicate Detection Based on Hadoop Platform
(College of Information Science and Technology, Jinan University, Guangzhou 510632, China)
摘要
图/表
参考文献
相似文献
本文已被:浏览 1535次   下载 2819
Received:April 22, 2013    Revised:May 28, 2013
中文摘要: XML数据越来越广泛地被用于信息交换与集成中, 其数据质量问题引起了人们的关注. 解决由数据质量引发的问题, 实体识别技术非常关键. 为了克服现有方法的不足, 在海量XML数据上进行高效的重复对象检测, 以实体识别技术为基础提出了基于Hadoop平台的XML文档重复检测算法, 它将所有标签节点统称为属性, 用实体来描述属性, 通过属性的比较, 快速地找到在某些属性上相同的所有实体对象, 并利用Hadoop应用框架处理海量数据的优势实现并行处理. 经过试验验证该方法良好的扩展性, 伸缩性和高效性.
中文关键词: XML  数据质量  重复检测  Hadoop  分布式
Abstract:As being more and more widely used for data exchange and integration, the XML data quality issues cause more concern. In order to overcome the problems caused by data quality, Entity Resolution(ER) is critical. To overcome the drawbacks of current methods's deficiency and perform entity resolution efficiently and effectively on massive XML data set, under the basis of Entity Resolution, an XML data duplicate detection based on hadoop platform algorithm is presented in this paper. The method uses entities to describe their atrributes. By the comparing of the attributes,we can find all the objects that have the same attributes quickly. Meanwhile, taking the advantage of the Hadoop platform which can process massive data parallel. From the experiments, the method has excellent performance in scalability, flexibility and efficiency.
文章编号:     中图分类号:    文献标志码:
基金项目:
引用文本:
李振兴,刘波.基于Hadoop平台的XML文档重复数据检测.计算机系统应用,2013,22(11):195-199
LI Zhen-Xing,LIU Bo.XML Data Duplicate Detection Based on Hadoop Platform.COMPUTER SYSTEMS APPLICATIONS,2013,22(11):195-199