Unsupervised learning of mDTD extraction patterns for Web text mining
SCIE
SCOPUS
- Title
- Unsupervised learning of mDTD extraction patterns for Web text mining
- Authors
- Kim, D; Jung, HM; Lee, GG
- Date Issued
- 2003-07
- Publisher
- PERGAMON-ELSEVIER SCIENCE LTD
- Abstract
- This paper presents a new extraction pattern, called modified Document Type Definition (mDTD), which relies on analytical interpretation to identify extraction target from the contents of the Web documents. From conventional DTD in XML documents, we develop two major extensions: first, we introduce an extended content model with type-specific operators and keywords, and second, we refine the way to interpret the conventional DTD rules. As the result of the two, bur mDTD becomes freely represent HTML structures and extraction targets. The goal of mDTD is to overcome the current major barriers, that is, domain portability (with minimal human intervention) and high performance, on information extraction. The human experts compose an mDTD as seed rules, and then our system automatically extracts a set of instances by the mDTD from structured documents on the Web. We use the extracted instances as Sequential mDTD Learner (SmL) inputs to generate new mDTD rules based on part-of-speech tags and features for lexical similarity. This process does not require any hand-annotated corpus. We have experimented with 330 Korean and 220 English Web documents on audio and video shopping sites. The average extraction precision is 91.3% for Korean and 81.9% for English. (C) 2003 Elsevier Science Ltd. All rights reserved.
- Keywords
- Web text mining; information extraction; extraction pattern; document type definition; sequential covering algorithm; INFORMATION EXTRACTION
- URI
- https://oasis.postech.ac.kr/handle/2014.oak/18507
- DOI
- 10.1016/S0306-4573(03)00004-9
- ISSN
- 0306-4573
- Article Type
- Article
- Citation
- INFORMATION PROCESSING & MANAGEMENT, vol. 39, no. 4, page. 623 - 637, 2003-07
- Files in This Item:
- There are no files associated with this item.
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.