基于多种策略的页面内容提取算法

高琰; 谷士文; 谭立球

基于多种策略的页面内容提取算法

详细信息

作者简介:
高琰(1973- ),女,讲师,博士,研究领域为智能信息处理,E-mail:gaoyan@mail.csu.edu.cn

计量
- 文章访问数: 1406
- HTML全文浏览量: 89
- PDF下载量: 422
- 被引次数: 0
出版历程
- 收稿日期: 2006-06-14
- 刊出日期: 2007-08-25

Web Content Extraction Based on Multiple Strategies

摘要

摘要: 针对W eb页面存在与主题无关的噪音的问题,提出了基于页面结构与页面内容相结合的多策略页面内容提取算法.该算法根据改进的VIPS(基于视觉信息的页面分割算法)生成页面的块结构树,通过定义内聚度阈值和块结构树的最大深度,实现了块结构树中不同区域内不同分块粒度的要求;根据W eb页面提供的结构信息和内容信息提取块结构树叶子节点中的"主题"块和"主题相关"块;最后,对主题块和主题相关块的内容进行合并,提取页面的主要内容.实验表明,对任意下载、不同内容类型的页面,该算法都能有效地提取页面内容.
- VIPS(基于视觉信息的页面分割算法) /
- 内聚度 /
- 最大深度 /
- 内容信息 /
- 结构信息
Abstract: In order to filter the noise in a web page,a new multi-strategy algorithm to extract the contents of a web page was proposed.With this algorithm,the granularity in different areas of the block tree of a web page established by the improved VIPS(visual based page segment) algorithm is controlled by defining the permitted degree of coherence and the maximum depth of the block tree.In addition,"topic" or "topic-relevant" blocks among the leaves of the block tree can be extracted from the blocks’ content information and structure information.Finally,the main content of a web page can be extracted by merging these blocks’ contents.Experiments on the web pages of three sites indicates that the proposed algorithm is effective for extracting the contents of any type of web pages.
- VIPS(visual based page segment) /
- degree of coherence /
- maximum depth /
- content information /
- structure information

HTML全文

参考文献(0)

施引文献

附加材料(0)

访问统计

点击查看大图

计量

文章访问数: 1406
HTML全文浏览量: 89
PDF下载量: 422
被引次数: 0

基于多种策略的页面内容提取算法

作者简介: 高琰(1973- ),女,讲师,博士,研究领域为智能信息处理,E-mail:gaoyan@mail.csu.edu.cn

计量

出版历程

Web Content Extraction Based on Multiple Strategies

计量

出版历程

目录

作者简介:
高琰(1973- ),女,讲师,博士,研究领域为智能信息处理,E-mail:gaoyan@mail.csu.edu.cn