STAPI: An Automatic Scraper for Extracting Iterative Title-Text Structure fromWeb Documents

被引：0

作者：

Zhang, Nan ^{[1
]}

Wilson, Shomir ^{[1
]}

Mitra, Prasenjit ^{[1
,2
]}

机构：

[1] Penn State Univ, Coll Informat Sci & Technol, University Pk, PA 16802 USA

[2] L3S Res Ctr, Hannover, Germany

来源：

LREC 2022: THIRTEEN INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION | 2022年

关键词：

Iterative Title-Text Dataset; Title-Text Identification; Automatic Scraping; Information Extraction; WRAPPER INDUCTION;

D O I：

暂无

中图分类号：

TP39 [计算机的应用];

学科分类号：

081203 ; 0835 ;

摘要：

Formal documents often are organized into sections of text, each with a title, and extracting this structure remains an under-explored aspect of natural language processing. This iterative title-text structure is valuable data for building models for headline generation and section title generation, but there is no corpus that contains web documents annotated with titles and prose texts. Therefore, we propose the first title-text dataset on web documents that incorporates a wide variety of domains to facilitate downstream training. We also introduce STAPI (Section Title And Prose text Identifier), a two-step system for labeling section titles and prose text in HTML documents. To filter out unrelated content like document footers, its first step involves a filter that reads HTML documents and proposes a set of textual candidates. In the second step, a typographic classifier takes the candidates from the filter and categorizes each one into one of the three pre-defined classes (title, prose text, and miscellany). We show that STAPI significantly outperforms two baseline models in terms of title-text identification. We release our dataset along with a web application to facilitate supervised and semi-supervised training in this domain.

引用

页码：3461 / 3470

页数：10