Calpric: Inclusive and Fine-grained Labeling of Privacy Policies with Crowdsourcing and Active Learning

被引:0
|
作者
Qiu, Wenjun [1 ]
Lie, David [1 ]
Austin, Lisa [1 ]
机构
[1] Univ Toronto, Toronto, ON, Canada
基金
加拿大自然科学与工程研究理事会;
关键词
D O I
暂无
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
A significant challenge to training accurate deep learning models on privacy policies is the cost and difficulty of obtaining a large and comprehensive set of training data. To address these challenges, we present Calpric, which combines automatic text selection and segmentation, active learning and the use of crowdsourced annotators to generate a large, balanced training set for Android privacy policies at low cost. Automated text selection and segmentation simplifies the labeling task, enabling untrained annotators from crowdsourcing platforms, like Amazon's Mechanical Turk, to be competitive with trained annotators, such as law students, and also reduces inter-annotator agreement, which decreases labeling cost. Having reliable labels for training enables the use of active learning, which uses fewer training samples to efficiently cover the input space, further reducing cost and improving class and data category balance in the data set. The combination of these techniques allows Calpric to produce models that are accurate over a wider range of data categories, and provide more detailed, fine-grained labels than previous work. Our crowdsourcing process enables Calpric to attain reliable labeled data at a cost of roughly $0.92-$1.71 per labeled text segment. Calpric's training process also generates a labeled data set of 16K privacy policy text segments across 9 data categories with balanced assertions and denials.
引用
收藏
页码:1055 / 1072
页数:18
相关论文
共 50 条
  • [1] Fine-Grained Crowdsourcing for Fine-Grained Recognition
    Jia Deng
    Krause, Jonathan
    Li Fei-Fei
    [J]. 2013 IEEE CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION (CVPR), 2013, : 580 - 587
  • [2] Privacy-preserving and Fine-grained Data Aggregation Framework for Crowdsourcing
    Zhuo, Gaoqiang
    [J]. 2017 TENTH INTERNATIONAL CONFERENCE ON MOBILE COMPUTING AND UBIQUITOUS NETWORK (ICMU), 2017, : 93 - 98
  • [3] Privacy-Preserving and Fair Crowdsourcing Framework With Fine-Grained Reuse Based on Blockchain
    Jiang, Shunrong
    Zhang, Xiao
    Chen, Jingwei
    Li, Jinpeng
    Wu, Haiqin
    Liu, Yiliang
    Zhou, Yong
    [J]. IEEE TRANSACTIONS ON NETWORK AND SERVICE MANAGEMENT, 2024, 21 (04): : 4061 - 4075
  • [4] ABCrowdMed: A Fine-Grained Worker Selection Scheme for Crowdsourcing Healthcare With Privacy-Preserving
    Li, Jiani
    Wang, Tao
    Yang, Bo
    Yang, Qiliang
    Zhang, Wenzheng
    Hong, Keyong
    [J]. IEEE TRANSACTIONS ON SERVICES COMPUTING, 2023, 16 (05) : 3182 - 3195
  • [5] Fine-Grained Disclosure of Access Policies
    Ardagna, Claudio Agostino
    di Vimercati, Sabrina De Capitani
    Foresti, Sara
    Neven, Gregory
    Paraboschi, Stefano
    Preiss, Franz-Stefan
    Samarati, Pierangela
    Verdicchio, Mario
    [J]. INFORMATION AND COMMUNICATIONS SECURITY, 2010, 6476 : 16 - +
  • [6] Fine-Grained Privacy Setting Prediction Using a Privacy Attitude Questionnaire and Machine Learning
    Raber, Frederic
    Kosmalla, Felix
    Krueger, Antonio
    [J]. HUMAN-COMPUTER INTERACTION - INTERACT 2017, PT IV, 2017, 10516 : 445 - 449
  • [7] PATIENT AWARE ACTIVE LEARNING FOR FINE-GRAINED OCT CLASSIFICATION
    Logan, Yash-yee
    Benkert, Ryan
    Mustafa, Ahmad
    AlRegib, Ghassan
    [J]. 2022 IEEE INTERNATIONAL CONFERENCE ON IMAGE PROCESSING, ICIP, 2022, : 3908 - 3912
  • [8] Towards Fine-grained Sampling for Active Learning in Object Detection
    Desai, Sai Vikas
    Balasubramanian, Vineeth N.
    [J]. 2020 IEEE/CVF CONFERENCE ON COMPUTER VISION AND PATTERN RECOGNITION WORKSHOPS (CVPRW 2020), 2020, : 4010 - 4014
  • [9] Labeling Workflow Views with Fine-Grained Dependencies
    Bao, Zhuowei
    Davidson, Susan B.
    Milo, Tova
    [J]. PROCEEDINGS OF THE VLDB ENDOWMENT, 2012, 5 (11): : 1208 - 1219
  • [10] Towards Fine-Grained Localization of Privacy Behaviors
    Jain, Vijayanta
    Ghanavati, Sepideh
    Peddinti, Sai Teja
    McMillan, Collin
    [J]. 2023 IEEE 8TH EUROPEAN SYMPOSIUM ON SECURITY AND PRIVACY, EUROS&P, 2023, : 258 - 277