Impact of model settings on the text-based Rao diversity index

被引：0

作者：

Zielinski, Andrea ^{[1
]}

机构：

[1] Fraunhofer Inst Syst & Innovat Res, Breslauer Str 48, D-76131 Karlsruhe, Germany

来源：

SCIENTOMETRICS | 2022年 / 127卷 / 12期

关键词：

LDA topic model; Rao Stirling; Interdisciplinarity; SCIENCE; INTERDISCIPLINARITY; MAP;

D O I：

10.1007/s11192-022-04312-x

中图分类号：

TP39 [计算机的应用];

学科分类号：

081203 ; 0835 ;

摘要：

Policymakers and funding agencies tend to support scientific work across disciplines, thereby relying on indicators for interdisciplinarity. Recently, text-based quantitative methods have been proposed for the computation of interdisciplinarity that hold promise to have several advantages over the bibliometric approach. In this paper, we provide a systematic analysis of the computation of the text-based Rao index, based on probabilistic topic models, comparing a classical LDA model versus a neural network topic model. We provide a systematic analysis of model parameters that affect the diversity scores and make the interaction between its different components explicit. We present an empirical study on a real data set, upon which we quantify the diversity of the research within several departments of Fraunhofer and Max Planck Society by means of scientific abstracts published in Scopus between 2008 and 2018. Our experiments show that parameter variations, i.e. the choice of the Number of topics, hyper-parameters, and size and balance of the underlying data used for training the model, have a strong effect on the topic model-based Rao metrics. In particular, we could observe that the quality of the topic models impacts on the downstream task of computing the Rao index. Topic models that yield semantically cohesive topics are less affected by fluctuations when varying over the number of topics, and result in more stable measurements of the Rao index.

引用

页码：7751 / 7768

页数：18

共 50 条

[21] Text-based interfaces and text-based bibliographic enhancements: Thinking beyond standard bibliographic information (and text)
Wall, TB
[J]. PROCEEDINGS OF THE ASIS ANNUAL MEETING, 1996, 33 : 278 - 278
[22] A Multiplicative Model for Learning Distributed Text-Based Attribute Representations
Kiros, Ryan
Zemel, Richard S.
Salakhutdinov, Ruslan
[J]. ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 27 (NIPS 2014), 2014, 27
[23] IZE - TEXT-BASED POWER
OMALLEY, C
[J]. PERSONAL COMPUTING, 1988, 12 (11): : 262 - 262
[24] Tracking with text-based messages
Alberola, C
Cybenko, GV
[J]. IEEE INTELLIGENT SYSTEMS & THEIR APPLICATIONS, 1999, 14 (04): : 70 - 78
[25] Text-based NP Enrichment
Elazar, Yanai
Basmov, Victoria
Goldberg, Yoav
Tsarfaty, Reut
[J]. TRANSACTIONS OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, 2022, 10 : 764 - 784
[26] Noticing and text-based chat
Lai, Chun
Zhao, Yong
[J]. LANGUAGE LEARNING & TECHNOLOGY, 2006, 10 (03): : 102 - 120
[27] Text-Based Industry Momentum
Hoberg, Gerard
Phillips, Gordon M.
[J]. JOURNAL OF FINANCIAL AND QUANTITATIVE ANALYSIS, 2018, 53 (06) : 2355 - 2388
[28] DEAFNESS AND TEXT-BASED LITERACY
PAUL, PV
[J]. AMERICAN ANNALS OF THE DEAF, 1993, 138 (02) : 72 - 75
[29] A deep learning model for recognition of complex Text-based CAPTCHAs
Arain, Rafaqat Hussain
Shaikh, Riaz Ahmed
Maitlo, Abdullah
Kumar, Kamlesh
Shah, Syed Safdar Ali
[J]. INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND NETWORK SECURITY, 2018, 18 (02): : 103 - 107
[30] Text-Based Recession Probabilities
Massimo Ferrari Minesso
Laura Lebastard
Helena Le Mezo
[J]. IMF Economic Review, 2023, 71 : 415 - 438

← 1 2 3 4 5 →