Impact of model settings on the text-based Rao diversity index

被引:0
|
作者
Andrea Zielinski
机构
[1] Fraunhofer Institute for Systems and Innovation Research,
来源
Scientometrics | 2022年 / 127卷
关键词
LDA topic model; Rao Stirling; Interdisciplinarity;
D O I
暂无
中图分类号
学科分类号
摘要
Policymakers and funding agencies tend to support scientific work across disciplines, thereby relying on indicators for interdisciplinarity. Recently, text-based quantitative methods have been proposed for the computation of interdisciplinarity that hold promise to have several advantages over the bibliometric approach. In this paper, we provide a systematic analysis of the computation of the text-based Rao index, based on probabilistic topic models, comparing a classical LDA model versus a neural network topic model. We provide a systematic analysis of model parameters that affect the diversity scores and make the interaction between its different components explicit. We present an empirical study on a real data set, upon which we quantify the diversity of the research within several departments of Fraunhofer and Max Planck Society by means of scientific abstracts published in Scopus between 2008 and 2018. Our experiments show that parameter variations, i.e. the choice of the Number of topics, hyper-parameters, and size and balance of the underlying data used for training the model, have a strong effect on the topic model-based Rao metrics. In particular, we could observe that the quality of the topic models impacts on the downstream task of computing the Rao index. Topic models that yield semantically cohesive topics are less affected by fluctuations when varying over the number of topics, and result in more stable measurements of the Rao index.
引用
下载
收藏
页码:7751 / 7768
页数:17
相关论文
共 50 条
  • [1] Impact of model settings on the text-based Rao diversity index
    Zielinski, Andrea
    SCIENTOMETRICS, 2022, 127 (12) : 7751 - 7768
  • [2] Impact of Model Settings on the Text-based Rao Diversity Index
    Zielinski, Andrea
    18TH INTERNATIONAL CONFERENCE ON SCIENTOMETRICS & INFORMETRICS (ISSI2021), 2021, : 1405 - 1416
  • [3] Text-Based Measures of Document Diversity
    Bache, Kevin
    Newman, David
    Smyth, Padhraic
    19TH ACM SIGKDD INTERNATIONAL CONFERENCE ON KNOWLEDGE DISCOVERY AND DATA MINING (KDD'13), 2013, : 23 - 31
  • [4] Influence diffusion model in text-based communication
    Matsumura, Naohiro
    Ohsawa, Yukio
    Ishizuka, Mitsuru
    Transactions of the Japanese Society for Artificial Intelligence, 2002, 17 (03) : 259 - 267
  • [5] Uncover Impact Factors of Text-based CAPTCHA Identification
    Tamang, Tsheten
    Bhattarakosol, Pattarasinee
    2012 7TH INTERNATIONAL CONFERENCE ON COMPUTING AND CONVERGENCE TECHNOLOGY (ICCCT2012), 2012, : 556 - 560
  • [6] Impact of Users' Beliefs in Text-Based Linguistic Interaction
    Catania, Vincenzo
    Monteleone, Salvatore
    Palesi, Maurizio
    Patti, Davide
    IEEE ACCESS, 2020, 8 : 46861 - 46867
  • [7] Impact of Glaucoma and Dry Eye on Text-Based Searching
    Sun, Michelle J.
    Rubin, Gary S.
    Akpek, Esen K.
    Ramulu, Pradeep Y.
    TRANSLATIONAL VISION SCIENCE & TECHNOLOGY, 2017, 6 (03):
  • [8] Impact of glaucoma and dry eye on text-based search
    Sun, Michelle
    Rubin, Gary S.
    Ramulu, Pradeep Y.
    INVESTIGATIVE OPHTHALMOLOGY & VISUAL SCIENCE, 2016, 57 (12)
  • [9] Text-based informatics
    Valdes-Perez, RE
    SCIENTIST, 1998, 12 (14): : 10 - 10
  • [10] What’s in the Text?—A Text-Based Financial Constraint Index for Indian Manufacturing Firms
    Amal Jose
    Saumitra Bhaduri
    Journal of Quantitative Economics, 2025, 23 (1) : 285 - 314