Calpric: Inclusive and Fine-grained Labeling of Privacy Policies with Crowdsourcing and Active Learning

被引：0

作者：

Qiu, Wenjun ^{[1
]}

Lie, David ^{[1
]}

Austin, Lisa ^{[1
]}

机构：

[1] Univ Toronto, Toronto, ON, Canada

来源：

PROCEEDINGS OF THE 32ND USENIX SECURITY SYMPOSIUM | 2023年

基金：

加拿大自然科学与工程研究理事会;

关键词：

D O I：

暂无

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and comprehensive set of training data. To address these challenges, we present Calpric, which combines automatic text selection and segmentation, active learning and the use of crowdsourced annotators to generate a large, balanced training set for Android privacy policies at low cost. Automated text selection and segmentation simplifies the labeling task, enabling untrained annotators from crowdsourcing platforms, like Amazon's Mechanical Turk, to be competitive with trained annotators, such as law students, and also reduces inter-annotator agreement, which decreases labeling cost. Having reliable labels for training enables the use of active learning, which uses fewer training samples to efficiently cover the input space, further reducing cost and improving class and data category balance in the data set. The combination of these techniques allows Calpric to produce models that are accurate over a wider range of data categories, and provide more detailed, fine-grained labels than previous work. Our crowdsourcing process enables Calpric to attain reliable labeled data at a cost of roughly $0.92-$1.71 per labeled text segment. Calpric's training process also generates a labeled data set of 16K privacy policy text segments across 9 data categories with balanced assertions and denials.

引用

页码：1055 / 1072

页数：18

共 50 条

[21] Fine-grained Delta Privacy Preservation for Hierarchical Contexts
Jiang, Xue
Huang, Yu
Wei, Hengfeng
[J]. 2016 INT IEEE CONFERENCES ON UBIQUITOUS INTELLIGENCE & COMPUTING, ADVANCED & TRUSTED COMPUTING, SCALABLE COMPUTING AND COMMUNICATIONS, CLOUD AND BIG DATA COMPUTING, INTERNET OF PEOPLE, AND SMART WORLD CONGRESS (UIC/ATC/SCALCOM/CBDCOM/IOP/SMARTWORLD), 2016, : 261 - 268
[22] Learning to Navigate for Fine-Grained Classification
Yang, Ze
Luo, Tiange
Wang, Dong
Hu, Zhiqiang
Gao, Jun
Wang, Liwei
[J]. COMPUTER VISION - ECCV 2018, PT XIV, 2018, 11218 : 438 - 454
[23] The Fine-Grained Impact of Gaming (?) on Learning
Gong, Yue
Beck, Joseph E.
Heffernan, Neil T.
Forbes-Summers, Elijah
[J]. INTELLIGENT TUTORING SYSTEMS, PT 1, PROCEEDINGS, 2010, 6094 : 194 - 203
[24] Fine-Grained Privacy Detection with Graph-Regularized Hierarchical Attentive Representation Learning
Chen, Xiaolin
Song, Xuemeng
Ren, Ruiyang
Zhu, Lei
Cheng, Zhiyong
Nie, Liqiang
[J]. ACM TRANSACTIONS ON INFORMATION SYSTEMS, 2020, 38 (04)
[25] Enforcing Fine-grained Constant-time Policies
Ammanaghatta Shivakumar, Basavesh
Barthe, Gilles
Grégoire, Benjamin
Laporte, Vincent
Priya, Swarn
[J]. Proceedings of the ACM Conference on Computer and Communications Security, 2022, : 83 - 96
[26] Mechanisms and policies for supporting fine-grained cycle stealing
Ryu, Kyung Dong
Hollingsworth, Jeffrey K.
Keleher, Peter J.
[J]. Proceedings of the International Conference on Supercomputing, 1999, : 93 - 100
[27] Modelling Fine-Grained Access Control Policies in Grids
Benjamin Aziz
[J]. Journal of Grid Computing, 2016, 14 : 477 - 493
[28] Modelling Fine-Grained Access Control Policies in Grids
Aziz, Benjamin
[J]. JOURNAL OF GRID COMPUTING, 2016, 14 (03) : 477 - 493
[29] Blockchain-based mechanism for fine-grained authorization in data crowdsourcing
Ma, Haiying
Huang, Elmo X.
Lam, Kwok-Yan
[J]. FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2020, 106 : 121 - 134
[30] Active Microservice Fine-Grained Scaling Algorithm
Peng, Kai
Ma, Fangling
Xu, Bo
Guo, Jialu
Hu, Menglan
[J]. Computer Engineering and Applications, 2024, 60 (08) : 274 - 286

← 1 2 3 4 5 →