Quantifying the Systematic Bias in the Accessibility and Inaccessibility of Web Scraping Content From URL-Logged Web-Browsing Digital Trace Data

被引：1

作者：

Dahlke, Ross ^{[1
,3
]}

Kumar, Deepak ^{[2
]}

Durumeric, Zakir ^{[2
]}

Hancock, Jeffrey T. ^{[1
]}

机构：

[1] Stanford Univ, Stanford, CA 94305 USA

[2] Stanford Univ, Stanford, CA 94305 USA

[3] Stanford Univ, Dept Commun, 450 Jane Stanford Way, Stanford, CA 94305 USA

来源：

SOCIAL SCIENCE COMPUTER REVIEW | 2023年

关键词：

digital trace data; internet measurement; misinformation; web-log data; web scraping; news; news consumption; BIG DATA; NEWS; PAYWALL; INTERNET; FUTURE; MEDIA; ROT;

D O I：

10.1177/08944393231218214

中图分类号：

TP39 [计算机的应用];

学科分类号：

081203 ; 0835 ;

摘要：

Social scientists and computer scientists are increasingly using observational digital trace data and analyzing these data post hoc to understand the content people are exposed to online. However, these content collection efforts may be systematically biased when the entirety of the data cannot be captured retroactively. We call this often unstated assumption the problematic assumption of accessibility. To examine the extent to which this assumption may be problematic, we identify 107k hard news and misinformation web pages visited by a representative panel of 1,238 American adults and record the degree to which the web pages individuals visited were accessible via successful web scrapes or inaccessible via unsuccessful scrapes. While we find that the URLs collected are largely accessible and with unrestricted content, we find there are systematic biases in which URLs are restricted, return an error, or are inaccessible. For example, conservative misinformation URLs are more likely to be inaccessible than other types of misinformation. We suggest how social scientists should capture and report digital trace and web scraping data.

引用

页数：16