###

DOI:

计算机系统应用英文版:2011,20(11):50-54,19

View/Add Comment 过刊浏览高级检索 HTML

←前一篇 | 后一篇→

码上扫一扫！

下载全文

基于多维语义的互联网药品信息提取方法

顾轶灵

(复旦大学软件学院,上海 201203)

Multidimensional-Semantics-Based Web Medicine Information Extraction

GU Yi-Ling

(Software School, Fudan University, Shanghai 201203, China)

摘要

图/表

参考文献

相似文献

本文已被：浏览 2143次下载 3197次
Received:March 10, 2011 Revised:April 18, 2011

中文摘要: 提出了基于多维语义的互联网药品信息提取方法,构建语义词典通过从多个维度对互联网药品知识进行描述,克服了不同来源网页之间的异构性并找出了其隐藏的共性。同时,采用了基于结构语义熵的方法对目标网页信息聚集区域进行定位,从中提取感兴趣的药品信息。最后再通过语义词典对提取的信息进行验证并自动生成XPath 提取规则进行补充。该方法能够自动有效地从互联网的多个信息来源获取药品信息,实验证明其具有较高的准确性与召回率,可以为政府相关部门加强互联网药品市场监管提供足够的信息依据。

中文关键词: Web 信息提取多维语义词典互联网药品信息结构语义熵 XPath

Abstract:A multidimensional-semantics based Web information extraction method is proposed in this article to extract medicine information on the Web. The method overcomes the heterogeneity of Web pages from different sources and finds the common characteristics among them by building up a semantic dictionary and describes the knowledge of medicine information over the Web. At the same time, it utilizes a structural-semantic-entropy-based approach to detect data-rich sections on Web pages, then extract information of interest from them and finally verify and supplement the extracted information by generating extraction rules using XPath. The method is able to obtain information from heterogeneous sources both automatically and effectively. Experiments shown that it has high precision and recall, thus can provide sufficient information for the government to enhance supervision of medicine market on the Web.

keywords: Web information extraction multidimensional semantic dictionary Web medicine information Structuralsemantic entropy Xpath

文章编号： 中图分类号： 文献标志码：

基金项目:

Author Name	Affiliation
GU Yi-Ling	Software School, Fudan University, Shanghai 201203, China

Author Name	Affiliation
GU Yi-Ling	Software School, Fudan University, Shanghai 201203, China

引用文本：
顾轶灵.基于多维语义的互联网药品信息提取方法.计算机系统应用,2011,20(11):50-54,19
GU Yi-Ling.Multidimensional-Semantics-Based Web Medicine Information Extraction.COMPUTER SYSTEMS APPLICATIONS,2011,20(11):50-54,19