Calpric: Inclusive and Fine-grained Labeling of Privacy Policies with Crowdsourcing and Active Learning

被引:0
|
作者
Qiu, Wenjun [1 ]
Lie, David [1 ]
Austin, Lisa [1 ]
机构
[1] Univ Toronto, Toronto, ON, Canada
基金
加拿大自然科学与工程研究理事会;
关键词
D O I
暂无
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and comprehensive set of training data. To address these challenges, we present Calpric, which combines automatic text selection and segmentation, active learning and the use of crowdsourced annotators to generate a large, balanced training set for Android privacy policies at low cost. Automated text selection and segmentation simplifies the labeling task, enabling untrained annotators from crowdsourcing platforms, like Amazon's Mechanical Turk, to be competitive with trained annotators, such as law students, and also reduces inter-annotator agreement, which decreases labeling cost. Having reliable labels for training enables the use of active learning, which uses fewer training samples to efficiently cover the input space, further reducing cost and improving class and data category balance in the data set. The combination of these techniques allows Calpric to produce models that are accurate over a wider range of data categories, and provide more detailed, fine-grained labels than previous work. Our crowdsourcing process enables Calpric to attain reliable labeled data at a cost of roughly $0.92-$1.71 per labeled text segment. Calpric's training process also generates a labeled data set of 16K privacy policy text segments across 9 data categories with balanced assertions and denials.
引用
收藏
页码:1055 / 1072
页数:18
相关论文
共 50 条
  • [21] Fine-grained Delta Privacy Preservation for Hierarchical Contexts
    Jiang, Xue
    Huang, Yu
    Wei, Hengfeng
    [J]. 2016 INT IEEE CONFERENCES ON UBIQUITOUS INTELLIGENCE & COMPUTING, ADVANCED & TRUSTED COMPUTING, SCALABLE COMPUTING AND COMMUNICATIONS, CLOUD AND BIG DATA COMPUTING, INTERNET OF PEOPLE, AND SMART WORLD CONGRESS (UIC/ATC/SCALCOM/CBDCOM/IOP/SMARTWORLD), 2016, : 261 - 268
  • [22] Learning to Navigate for Fine-Grained Classification
    Yang, Ze
    Luo, Tiange
    Wang, Dong
    Hu, Zhiqiang
    Gao, Jun
    Wang, Liwei
    [J]. COMPUTER VISION - ECCV 2018, PT XIV, 2018, 11218 : 438 - 454
  • [23] The Fine-Grained Impact of Gaming (?) on Learning
    Gong, Yue
    Beck, Joseph E.
    Heffernan, Neil T.
    Forbes-Summers, Elijah
    [J]. INTELLIGENT TUTORING SYSTEMS, PT 1, PROCEEDINGS, 2010, 6094 : 194 - 203
  • [24] Fine-Grained Privacy Detection with Graph-Regularized Hierarchical Attentive Representation Learning
    Chen, Xiaolin
    Song, Xuemeng
    Ren, Ruiyang
    Zhu, Lei
    Cheng, Zhiyong
    Nie, Liqiang
    [J]. ACM TRANSACTIONS ON INFORMATION SYSTEMS, 2020, 38 (04)
  • [25] Enforcing Fine-grained Constant-time Policies
    Ammanaghatta Shivakumar, Basavesh
    Barthe, Gilles
    Grégoire, Benjamin
    Laporte, Vincent
    Priya, Swarn
    [J]. Proceedings of the ACM Conference on Computer and Communications Security, 2022, : 83 - 96
  • [26] Mechanisms and policies for supporting fine-grained cycle stealing
    Ryu, Kyung Dong
    Hollingsworth, Jeffrey K.
    Keleher, Peter J.
    [J]. Proceedings of the International Conference on Supercomputing, 1999, : 93 - 100
  • [27] Modelling Fine-Grained Access Control Policies in Grids
    Benjamin Aziz
    [J]. Journal of Grid Computing, 2016, 14 : 477 - 493
  • [28] Modelling Fine-Grained Access Control Policies in Grids
    Aziz, Benjamin
    [J]. JOURNAL OF GRID COMPUTING, 2016, 14 (03) : 477 - 493
  • [29] Blockchain-based mechanism for fine-grained authorization in data crowdsourcing
    Ma, Haiying
    Huang, Elmo X.
    Lam, Kwok-Yan
    [J]. FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2020, 106 : 121 - 134
  • [30] Active Microservice Fine-Grained Scaling Algorithm
    Peng, Kai
    Ma, Fangling
    Xu, Bo
    Guo, Jialu
    Hu, Menglan
    [J]. Computer Engineering and Applications, 2024, 60 (08) : 274 - 286