A large dataset of scientific text reuse in Open-Access publications

被引:0
|
作者
Lukas Gienapp
Wolfgang Kircheis
Bjarne Sievers
Benno Stein
Martin Potthast
机构
[1] Leipzig University,Text Mining and Retrieval Group
[2] Bauhaus-Universität Weimar,Web Technology and Information Systems Group
[3] Center for Scalable Data Analytics and Artificial Intelligence,ScaDS.AI
来源
关键词
D O I
暂无
中图分类号
学科分类号
摘要
We present the Webis-STEREO-21 dataset, a massive collection of Scientific Text Reuse in Open-access publications. It contains 91 million cases of reused text passages found in 4.2 million unique open-access publications. Cases range from overlap of as few as eight words to near-duplicate publications and include a variety of reuse types, ranging from boilerplate text to verbatim copying to quotations and paraphrases. Featuring a high coverage of scientific disciplines and varieties of reuse, as well as comprehensive metadata to contextualize each case, our dataset addresses the most salient shortcomings of previous ones on scientific writing. The Webis-STEREO-21 does not indicate if a reuse case is legitimate or not, as its focus is on the general study of text reuse in science, which is legitimate in the vast majority of cases. It allows for tackling a wide range of research questions from different scientific backgrounds, facilitating both qualitative and quantitative analysis of the phenomenon as well as a first-time grounding on the base rate of text reuse in scientific publications.
引用
收藏
相关论文
共 50 条
  • [1] A large dataset of scientific text reuse in Open-Access publications
    Gienapp, Lukas
    Kircheis, Wolfgang
    Sievers, Bjarne
    Stein, Benno
    Potthast, Martin
    SCIENTIFIC DATA, 2023, 10 (01)
  • [2] SCIENTIFIC PUBLICATIONS MARKET AND OPEN-ACCESS POLICIES
    Munteanu, R. A.
    Apetroae, M.
    Munteanu, M.
    Iudean, D.
    QUALITY MANAGEMENT IN HIGHER EDUCATION, PROCEEDINGS, 2008, : 553 - 557
  • [3] Why are Latin American countries in the limbo of open-access scientific publications?
    Chavarria, Max
    Lomonte, Bruno
    Gutierrez, Jose Maria
    Moreno, Edgardo
    Perez, Alice L.
    SCIENCE AND PUBLIC POLICY, 2024,
  • [4] Open Access to Scientific Publications
    Beaudouin-Lafon, Michel
    COMMUNICATIONS OF THE ACM, 2010, 53 (02) : 32 - 34
  • [5] Open Access in scientific publications
    Perez Rodrigo, Carmen
    REVISTA ESPANOLA DE NUTRICION COMUNITARIA-SPANISH JOURNAL OF COMMUNITY NUTRITION, 2010, 16 (04): : 203 - 203
  • [6] Open access to scientific publications
    Verma, IM
    MOLECULAR THERAPY, 2004, 10 (04) : 609 - 609
  • [7] Marketing of acquisition for open-access platforms of publications
    Hilse, Stefan
    Depping, Ralf
    BIBLIOTHEK FORSCHUNG UND PRAXIS, 2008, 32 (03) : 334 - 347
  • [8] Open-access Journals - A Scientific Thriller
    Roulet, J. F.
    van Meerbeek, Bart
    JOURNAL OF ADHESIVE DENTISTRY, 2013, 15 (06): : 503 - 504
  • [9] An open-access dataset of emergency department admissions at a large teaching hospital in Iran
    Hassan, Zohreh Bani
    Kazemi, Mohammad Reza
    Jangi, Majid
    Tabesh, Hamed
    DATA IN BRIEF, 2024, 56
  • [10] An Analysis on the Open-Access Publications in "Religious Education" Journal
    Kuscuoglu, Aslihan
    DINBILIMLERI AKADEMIK ARASTIRMA DERGISI-JOURNAL OF ACADEMIC RESEARCH IN RELIGIOUS SCIENCES, 2021, 21 (02): : 1097 - 1128