Implementation of a weblog extraction system with an improved template extraction technique
CHANG E; Chang E (E-mail:chang_e@seu.edu.cn)
2013-03-25
发表期刊Chinese Journal of Library and Information Science
ISSN1674-3393
卷号6期号:1页码:52-63
摘要

Purpose: The objectives of this study are to explore an effective technique to extract information from weblogs and develop an experimental system to extract structured information as much as possible with this technique. The system will lay a foundation for evaluation, analysis, retrieval, and utilization of the extracted information.

Design/methodology/approach: An improved template extraction technique was proposed. Separate templates designed for extracting blog entry titles, posts and their comments were established, and structured information was extracted online step by step. A dozen of data items, such as the entry titles, posts and their commenters and comments, the numbers of views, and the numbers of citations were extracted from eight major Chinese blog websites, including Sina, Sohu and Bokee.

Findings: Results showed that the average accuracy of the experimental extraction system reached 94.6%. Because the online and multi-threading extraction technique was adopted, the speed of extraction was improved with the average speed of 15 pages per second without considering the network delay. In addition, entries posted by Ajax technology can be extracted successfully.

Research limitations: As the templates need to be established in advance, this extraction technique can be effectively applied to a limited range of blog websites. In addition, the stability of the extraction templates was affected by the source code of the blog pages.

Practical implications: This paper has studied and established a blog page extraction system, which can be used to extract structured data, preserve and update the data, and facilitate the collection, study and utilization of the blog resources, especially academic blog resources.

Originality/value: This modified template extraction technique outperforms the Web page downloaders and the specialized blog page downloaders with structured and comprehensive data extraction.

;

Purpose: The objectives of this study are to explore an effective technique to extract information from weblogs and develop an experimental system to extract structured information as much as possible with this technique. The system will lay a foundation for evaluation, analysis, retrieval, and utilization of the extracted information.

Design/methodology/approach: An improved template extraction technique was proposed. Separate templates designed for extracting blog entry titles, posts and their comments were established, and structured information was extracted online step by step. A dozen of data items, such as the entry titles, posts and their commenters and comments, the numbers of views, and the numbers of citations were extracted from eight major Chinese blog websites, including Sina, Sohu and Bokee.

Findings: Results showed that the average accuracy of the experimental extraction system reached 94.6%. Because the online and multi-threading extraction technique was adopted, the speed of extraction was improved with the average speed of 15 pages per second without considering the network delay. In addition, entries posted by Ajax technology can be extracted successfully.

Research limitations: As the templates need to be established in advance, this extraction technique can be effectively applied to a limited range of blog websites. In addition, the stability of the extraction templates was affected by the source code of the blog pages.

Practical implications: This paper has studied and established a blog page extraction system, which can be used to extract structured data, preserve and update the data, and facilitate the collection, study and utilization of the blog resources, especially academic blog resources.

Originality/value: This modified template extraction technique outperforms the Web page downloaders and the specialized blog page downloaders with structured and comprehensive data extraction.

关键词WeBlog (Blog) Web Information Extraction Extraction Template
学科领域编辑出版
URL查看原文
项目资助者This work is supported by the Foundation for Humanities and Social Sciences of the Chinese Ministry of Education (Grant No.: 08JC870002).
文献类型期刊论文
条目标识符http://ir.las.ac.cn/handle/12502/6150
专题Journal of Data and Information Science_Chinese Journal of Library and Information Science-2013
通讯作者Chang E (E-mail:chang_e@seu.edu.cn)
推荐引用方式
GB/T 7714
CHANG E,Chang E . Implementation of a weblog extraction system with an improved template extraction technique[J]. Chinese Journal of Library and Information Science,2013,6(1):52-63.
APA CHANG E,&Chang E .(2013).Implementation of a weblog extraction system with an improved template extraction technique.Chinese Journal of Library and Information Science,6(1),52-63.
MLA CHANG E,et al."Implementation of a weblog extraction system with an improved template extraction technique".Chinese Journal of Library and Information Science 6.1(2013):52-63.
条目包含的文件
文件名称/大小 文献类型 版本类型 开放类型 使用许可
E CHANG.pdf(11544KB) 开放获取使用许可请求全文
个性服务
推荐该条目
保存到收藏夹
查看访问统计
导出为Endnote文件
谷歌学术
谷歌学术中相似的文章
[CHANG E]的文章
[Chang E (E-mail:chang_e@seu.edu.cn)]的文章
百度学术
百度学术中相似的文章
[CHANG E]的文章
[Chang E (E-mail:chang_e@seu.edu.cn)]的文章
必应学术
必应学术中相似的文章
[CHANG E]的文章
[Chang E (E-mail:chang_e@seu.edu.cn)]的文章
相关权益政策
暂无数据
收藏/分享
所有评论 (0)
暂无评论
 

除非特别说明,本系统中所有内容都受版权保护,并保留所有权利。