Outlier Ranking for Large-Scale Public Health Data

被引：0

作者：

Joshi, Ananya ^{[1
]}

Townes, Tina ^{[1
]}

Gormley, Nolan ^{[1
]}

Neureiter, Luke ^{[1
]}

Rosenfeld, Roni ^{[1
]}

Wilder, Bryan ^{[1
]}

机构：

[1] Carnegie Mellon Univ, 5000 Forbes Rd, Pittsburgh, PA 15213 USA

来源：

THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 20 | 2024年

基金：

美国国家科学基金会;

关键词：

ALGORITHMS;

D O I：

暂无

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

Disease control experts inspect public health data streams daily for outliers worth investigating, like those corresponding to data quality issues or disease outbreaks. However, they can only examine a few of the thousands of maximally-tied outliers returned by univariate outlier detection methods applied to large-scale public health data streams. To help experts distinguish the most important outliers from these thousands of tied outliers, we propose a new task for algorithms to rank the outputs of any univariate method applied to each of many streams. Our novel algorithm for this task, which leverages hierarchical networks and extreme value analysis, performed the best across traditional outlier detection metrics in a human-expert evaluation using public health data streams. Most importantly, experts have used our open-source Python implementation since April 2023 and report identifying outliers worth investigating 9.1x faster than their prior baseline. Other organizations can readily adapt this implementation to create rankings from the outputs of their tailored univariate methods across large-scale streams.

引用

页码：22176 / 22184

页数：9

共 50 条

[1] Outlier Detection Forest for Large-Scale Categorical Data Sets
Sun, Zhipeng
Du, Hongwei
Ye, Qiang
Liu, Chuang
Kibenge, Patricia Lilian
Huang, Hui
Li, Yuying
[J]. COMPUTATIONAL DATA AND SOCIAL NETWORKS, 2019, 11917 : 45 - 56
[2] Big Data, Large-Scale Text Analysis, and Public Health Research
Chowkwanyun, Merlin
[J]. AMERICAN JOURNAL OF PUBLIC HEALTH, 2019, 109 : 5126 - 5127
[3] Information-Theoretic Outlier Detection for Large-Scale Categorical Data
Wu, Shu
Wang, Shengrui
[J]. IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2013, 25 (03) : 589 - 602
[4] The Steganographer is the Outlier: Realistic Large-Scale Steganalysis
Ker, Andrew D.
Pevny, Tomas
[J]. IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, 2014, 9 (09) : 1424 - 1435
[5] Dialogue Response Ranking Training with Large-Scale Human Feedback Data
Gao, Xiang
Zhang, Yizhe
Galley, Michel
Brockett, Chris
Dolan, Bill
[J]. PROCEEDINGS OF THE 2020 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP), 2020, : 386 - 395
[6] Heuristic Ranking Classification Method for Complex Large-Scale Survival Data
Fard, Nasser
Sadeghzadeh, Keivan
[J]. MODELLING, COMPUTATION AND OPTIMIZATION IN INFORMATION SYSTEMS AND MANAGEMENT SCIENCES - MCO 2015 - PT II, 2015, 360 : 47 - 56
[7] An outlier detection approach in large-scale data stream using rough set
Manmohan Singh
Rajendra Pamula
[J]. Neural Computing and Applications, 2020, 32 : 9113 - 9127
[8] Outlier Detection in Large-Scale Sensor Network Data Using Shrinkage Estimators
Wu, Ming-Chun
Chen, Kwang-Cheng
[J]. 2015 IEEE GLOBAL COMMUNICATIONS CONFERENCE (GLOBECOM), 2015,
[9] An outlier detection approach in large-scale data stream using rough set
Singh, Manmohan
Pamula, Rajendra
[J]. NEURAL COMPUTING & APPLICATIONS, 2020, 32 (13): : 9113 - 9127
[10] A fast algorithm for learning a ranking function from large-scale data sets
Raykar, Vikas C.
Duraiswami, Ramani
Krishnapuram, Balaji
[J]. IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, 2008, 30 (07) : 1158 - 1170

← 1 2 3 4 5 →