




































Hypothesis, Vol. 37, No. 1, 2025

Authors’ Reply to Rife et al: Clarifying Definitions and
Methods in Assessing the Accuracy of Scite

Caitlin Bakker, MLIS, AHIP-Da, Nicole Theis-Mahon, MLIS, AHIPb, Sarah Jane Brown,
MScc

aDiscovery Technologies Librarian, Dr. John Archer Library, University of Regina, Regina,
Saskatchewan, Canada, https://www.orcid.org/0000-0003-4154-8382 ,
caitlin.bakker@uregina.ca

bEvidence Synthesis Librarian & Health Sciences Collections Coordinator, Health Sciences
Library, University of Minnesota, Minneapolis, Minnesota,
https://www.orcid.org/0000-0002-6913-5195 , theis025@umn.edu

cLibrarian Liaison to the College of Pharmacy, Health Sciences Library, University of
Minnesota, Minneapolis, Minnesota, https://www.orcid.org/0000-0001-7699-4417 ,
sjbrown@umn.edu

Cite as: Bakker C, Theis-Mahon N, and Brown SJ. Authors’ Reply to Rife et al: Clarifying
Definitions and Methods in Assessing the Accuracy of Scite. Hypothesis. 2025;37(1). doi:
10.18060/28651

Bakker, Theis-Mahon, Brown. All works in Hypothesis are licensed under
a CC BY-NC 4.0 DEED Attribution-NonCommercial 4.0 International. Authors own
copyright of their articles appearing in Hypothesis. Readers may copy articles without
permission of the copyright owner(s), as long as the author(s) are acknowledged in the copy,
and the copy is used for educational, not-for-profit purposes. For any other use of articles,
please contact the copyright owner(s).

1

https://www.orcid.org/0000-0003-4154-8382
https://orcid.org/
mailto:Caitlin.Bakker@uregina.ca
https://www.orcid.org/0000-0002-6913-5195
https://orcid.org/
mailto:theis025@umn.edu
https://www.orcid.org/0000-0001-7699-4417
https://orcid.org/
mailto:sjbrown@umn.edu
https://creativecommons.org/licenses/by-nc/4.0/


Hypothesis, Vol. 37, No. 1, 2025

We appreciate the time and attention Rife et al. have spent considering our original paper, and
in crafting their response1. We note with appreciation their encouragement of other
researchers to perform external validations of their tool, and their acknowledgement of the
potential biases of their own evaluation. In reflecting upon Rife et al.’s statements1, we noted
that there appear to be several misunderstandings and misstatements, and we appreciate the
opportunity to provide clarification of our research and aims.

Rife et al. argue that we “restricted [our] analyses to citations of retracted works in systematic
literature reviews, which artificially limits the types of statements that could be considered
supporting or contrasting a specific claim.1” We agree that the decision to focus on systematic
reviews is a notable limitation of our study, as is the decision to focus on retracted
publications. These limitations, and their implications, were described in depth in the original
article. In the limitations section of our article, we note the following:

“Our research has several notable limitations. First, we focus on a sample of
publications both within a discipline and using a specific study design. While
scite is not programmed to perform differently based on study design or
discipline, it is possible that studies using different publications would reveal
different levels of accuracy. We also chose to focus on systematic reviews that
had cited at least one retracted publication. Retraction is a relatively rare and
extreme publication state. As such, the specific citations included in our sample
are, by nature of their retracted status, not indicative of the majority of
publications. Although we chose to focus on these outliers under the assumption
that these publications would warrant stronger critique, this does limit the
generalizability of our findings. Finally, this work was done in conjunction with a
larger research project, rather than as a standalone project. As such, the sample
was derived through this larger work, rather than being selected for the sole
purpose of assessing scite. Other sampling methods may lead to different
results.2”

We agree that focusing on a specific sample is a limitation of our study that could limit
generalizability; however, we do not agree with the assessment of this as “artificially
limit[ed].1” Rather, we argue that our approach is indicative of one instance of
information-seeking behavior. Information-seeking behavior is defined as “the purposive
seeking for information as a consequence of a need to satisfy some goal.3” It is inextricably
tied to the specific needs and context of the user. As Kuhlthau notes, “an information search is
a learning process in which choices along the way are dependent on personal constructs rather
than on one universal predictable search for everyone.4” We engaged with a tool based on our
specific information need, which was grounded in an interest in a specific subject area and
study design. Rife et al.’s positioning of our approach as “artificially limited” is seemingly at
odds with well-established principles and practices of information-seeking behavior.

We believe that Rife et al.1 misunderstood and mischaracterized our research aims and intent
in several ways. In response to our findings, they present three examples of citations in
systematic reviews outside our original set and state their assumption of how we would have
classified the citation. First, we note that one of their examples appears to be included in error.
The first example they present is Gatto5 citing Li et al6. We believe that Rife et al. may be
conflating Li et al.’s 2014 paper6 with their 2013 paper of the same name7, which was
retracted. The retraction notice for Li et al.’s 2013 paper8 states that the authors requested the

2



Hypothesis, Vol. 37, No. 1, 2025

retraction due to errors they discovered in the paper following its publication. In this notice,
the authors state that they would be publishing a corrected version of the manuscript, which is
presumably what was subsequently published in 20148. As we have found no evidence that
the 2014 paper6 was also retracted, it seems that Gatto5 was not citing a retracted publication,
and as such this example is not relevant.

Rife et al.’s language throughout their manuscript may lead readers to misinterpret their
assumptions as fact, as they repeatedly state what we “would have” done1. These statements
are problematic for several reasons. We are of the opinion that researchers should attempt to
avoid positioning their hypotheses, assumptions, or beliefs as proven fact, and that to state
unproven beliefs as fact is contrary to and undermines the principles of the scientific method.
We are reminded of Faraday’s assertion of the importance of “distinguish[ing] that knowledge
which consists of assumption, by which I mean theory and hypothesis, from that which is the
knowledge of facts and laws; never raising the former to the dignity or authority of the
latter. . . 9”

Rife et al.’s positioning of their assumptions of our classification scheme as fact is especially
problematic as their assumptions are incorrect. In their second example, they assert that we
would have classified the following passage in Zhang et al.10 as supporting:

“Some cross-sectional studies have shown a close association between thyroid
disease and metabolic disorders (i.e., metabolic syndrome [MetS] and its
components) (11–13). MetS is characterized by a cluster of abnormal metabolic
parameters consisting of insulin resistance, central obesity, type 2 diabetes,
impaired glucose tolerance, hyperinsulinemia, and dyslipidemia (14). The global
prevalence of MetS is between 11.6% and 62.5% (15).” (p. 2) (emphasis added by
Rife et al.1)

Rife et al. state that “[a]gain, because Kaur does not mention the fact that Zhang et al. was
retracted, Bakker et al. would classify this citation as supporting.1" We would like to clarify
that Zhang10 is not the retracted publication; Kaur11 is the retracted publication. This is a
passage from Zhang’s 2021 paper citing Kaur’s 2014 retracted publication, not Kaur citing
Zhang.

In addition to clarifying which publication is retracted, we would also like to clarify that Rife
et al.’s statement is wholly inaccurate1. This citation was not included in our original study,
since it did not meet our inclusion criteria. However, after consulting Zhang10, we confirmed
that we would have classified this citation as mentioning, as it is by Scite. This is based on our
holistic evaluation of how the citation is functioning within the paper overall. In this case, the
citation occurred in the Introduction and functioned to provide context and background, rather
than furthering the findings of that retracted publication. While one could argue that such a
citation may be inappropriate, appropriateness of citation was not in the scope of our work.
Rife et al.’s misstatement leads us to believe that they may have fundamentally misunderstood
our classification system1.

Rife et al. appear to consider our classification scheme to be a binary one, where supporting
citations are determined to be such by the mere absence of the word retracted1. Our original
manuscript describes a three-part classification system rather than a binary classification
system2 . We classified citations as supporting, contrasting, or mentioning, depending on how

3



Hypothesis, Vol. 37, No. 1, 2025

they functioned within the context of the citing article. We attempted to align this
classification with what was proposed by Nicholson et al. in their QSS article12.

In that article, Nicholson et al. state that “scite focuses on the authors’ reasons for citing a
paper" and that “[e]xtracted citation statements are classified into supporting, contrasting, or
mentioning, to identify studies that have tested the claim and to evaluate how a scientific
claim has been evaluated in the literature by subsequent research.12" In the context of
evidence synthesis, our area of study, such evaluation may manifest in a number of ways,
including the systematic assessment of potential biases and methodological issues through
formalized risk of bias assessments, or through the identification of outliers through statistical
tests, such as sensitivity and leave-one-out analyses, among others.

The decision to include or exclude a publication, and the ways in which those publications are
included, is a multi-faceted one. Although authors consider whether the study meets
eligibility criteria, they also may consider whether the methodological quality of the study is
sufficient to include it in particular analyses, or if additional analyses should be conducted to
account for the influence of any one study. In addition to applying inclusion and exclusion
criteria, authors of evidence syntheses must evaluate and apply their scientific judgment to
each individual report and its claims. To be included or excluded in an evidence synthesis is,
in some cases, evidence of “how a scientific claim has been evaluated in the literature by
subsequent research.12”

Our description of our three-part classification system was tailored to the core readership of
Hypothesis: health information professionals who are likely familiar with principles of
evidence-based practice. We classified publications as supporting if they were included as
reports in the evidence synthesis–that is to say, if they were included as component studies
that had undergone the rigorous selection, appraisal, extraction and syntheses processes and
were subsequently included as component data within the larger study. Of note, not all
included reports were considered supporting citations, as an included report whose
methodology and findings were heavily critiqued, for example, would not have been classified
as supporting. We classified citations as contrasting if they were treated as retracted, that is to
say if the publication was treated as “contain[ing] such seriously flawed or erroneous content
or data that their findings and conclusions cannot be relied upon,13” whether that was due to a
concern with the report, such as findings that were inexplicably divergent from other
comparable studies, or the underlying research, such as concerns raised about methodological
issues or suspicions of misconduct. The majority of citations in our study were classified as
mentioning, as they were by Scite2.

We understand that to those unfamiliar with evidence synthesis, our reference to “included
reports” in evidence syntheses could have led to the misunderstanding that this simply meant
any cited paper, as opposed to a study that had gone through an intensive review, appraisal
and evaluation process to assess its quality and fit for purpose before being incorporated into
the body of evidence being presented. We also recognize that our brief description of
supporting, contrasting and mentioning citations omitted various nuances that we considered
when assessing citations. While it had been our aim to be as succinct as possible rather than
describing the extensive decision-making process behind each classification decision, we
believe that such brevity may have posed challenges for some readers, for which we
apologize.

4



Hypothesis, Vol. 37, No. 1, 2025

The third example Rife et al. present, Stavale14 citing Tsukumo15, is a systematic review that
“aimed to investigate the profile of medical and life sciences research retractions from authors
affiliated with Brazilian academic institutions.14" Given that the scope of this project is
focused specifically on identifying and analyzing retracted publications, we cannot speak to
how we would have classified this citation as it is an entirely different situation and use case
than those in which we were primarily interested.

Like many information professionals, we are intrigued and excited by technology and the
opportunities for identifying, accessing and synthesizing an ever-increasing amount of
information. Tools such as Scite, Perplexity, Research Rabbit, Primo Research Assistant and
others have the potential to aid researchers in completing complex tasks, such as identifying
gaps in the literature, synthesizing publication findings, and generating text. As noted in
Information Literacy in the Age of Algorithms, “these mysterious black boxes can answer in
seconds a question that formerly required hours in a library (though the answer may not
necessarily be entirely accurate).16”

Such tools have the potential to–and arguably already are–revolutionizing the ways in which
students and researchers access, use, and create information. The need for algorithmic literacy
and critical discourse in the context of evidence synthesis and evidence-based practice has
never been greater. As Bender et al. note in their seminal paper on large language models:
“the risks associated with synthetic but seemingly coherent text are deeply connected to the
fact that such synthetic text can enter into conversations without any person or entity being
accountable for it.17” In other words, an algorithm functions to take an input and produce an
output but has no understanding of or accountability for the meaning of what is being ingested
or produced. As a growing number of tools enter the marketplace, and existing tools are
further developed, refined and integrated into everyday life, it is essential that information
professionals actively engage in critical evaluation to ensure that those using the tool have full
knowledge of its capabilities and its limitations.

CRediT
Caitlin Bakker: Conceptualization, Writing – original draft, Writing - review & editing;
Nicole Theis-Mahon: Conceptualization, Writing -original draft, Writing - review & editing;
Sarah Jane Brown: Conceptualization, Writing – original draft, Writing - review & editing

References
1. Rife SC, Nicholson JM, Uppala A, Rosati D. Reply to Bakker et al.: Assessing the
accuracy of the scite citation classification system requires the same definitions to be used for
training as for testing. Hypothesis. 2025; 37(1). doi: 10.18060/28018

2. Bakker C, Theis-Mahon N, Brown SJ. Evaluating the accuracy of scite, a smart citation
index. Hypothesis. 2023;35(2). doi: 10.18060/26528

3. Wilson TD. Human Information Behavior. Informing Sci. 2000;2(2):049–56. doi:
10.28945/576

4. Kuhlthau CC. Seeking meaning: a process approach to library and information services.

5



Hypothesis, Vol. 37, No. 1, 2025

2nd ed. Westport, Connecticut: Libraries Unlimited; 2004. 247 p.

5. Gatto RG. Molecular and microstructural biomarkers of neuroplasticity in
neurodegenerative disorders through preclinical and diffusion magnetic resonance imaging
studies. J Integr Neurosci. 2020;19(3):571–92. doi: 10.31083/j.jin.2020.03.165

6. Li W, Yu J, Liu Y, Huang X, Abumaria N, Zhu Y, et al. Elevation of brain magnesium
prevents synaptic loss and reverses cognitive deficits in Alzheimer’s disease mouse model.
Mol Brain. 2014;7(1):65–65. doi: 10.1186/s13041-014-0065-y

7. Li W, Yu J, Liu Y, Huang X, Abumaria N, Zhu Y, et al. Elevation of brain magnesium
prevents and reverses cognitive deficits and synaptic loss in Alzheimer’s disease mouse
model. J Neurosci Off J Soc Neurosci. 2013 May 8;33(19):8423–41. doi:
10.1523/JNEUROSCI.4610-12.2013

8. Author-Initiated Retraction: Li et al., Elevation of brain magnesium prevents and reverses
cognitive deficits and synaptic loss in Alzheimer’s disease mouse model. J Neurosci. 2014
Apr 16;34(16):5733–5733. doi: 10.1523/JNEUROSCI.1265-14.2014

9. Faraday M. XXIII. A speculation touching electric conduction and the nature of matter.
Lond Edinb Dublin Philos Mag J Sci. 1844 Feb;24(157):136–44. doi:
10.1080/14786444408644817

10. Zhang C, Gao X, Han Y, Teng W, Shan Z. Correlation between thyroid nodules and
metabolic syndrome: a systematic review and meta-analysis. Front Endocrinol.
2021;12:730279. doi: 10.3389/fendo.2021.730279

11. Kaur J. A comprehensive review on metabolic syndrome. Cardiol Res Pract.
2014;2014:943162. doi: 10.1155/2014/943162

12. Nicholson JM, Mordaunt M, Lopez P, Uppala A, Rosati D, Rodrigues NP, et al. scite: a
smart citation index that displays the context of citations and classifies their intent using deep
learning. Quant Sci Stud. 2021;2(3):882–98. doi: 10.1162/qss_a_00146

13. Barbour V, Kleinert S, Wager E, Yentis S. Guidelines for retracting articles [Internet].
Committee on Publication Ethics; 2009 Sep [cited 2024 Oct 16]. Available from:
https://publicationethics.org/node/19896

14. Stavale R, Ferreira GI, Galvão JAM, Zicker F, Novaes MRCG, Oliveira CM de, et al.
Research misconduct in health and life sciences research: a systematic review of retracted
literature from Brazilian institutions. PloS One. 2019;14(4):e0214272. doi:
10.1371/journal.pone.0214272

15. Tsukumo DML, Carvalho-Filho MA, Carvalheira JBC, Prada PO, Hirabara SM, Schenka
AA, et al. Loss-of-function mutation in Toll-like receptor 4 prevents diet-induced obesity and
insulin resistance. Diabetes. 2007 Aug;56(8):1986–98. doi: 10.2337/db06-1595

16. Head AJ, Fister B, MacMillan M. Information literacy in the age of algorithms [Internet].

6

https://publicationethics.org/node/19896


Hypothesis, Vol. 37, No. 1, 2025

United States: Project Information Research Institute; 2020 Jan [cited 2024 Oct 24].
Available from: https://projectinfolit.org/publications/algorithm-study

17. Bender EM, Gebru T, McMillan-Major A, Shmitchell S. On the dangers of stochastic
parrots: can language models be too big?. In: Proceedings of the 2021 ACM Conference on
Fairness, Accountability, and Transparency. Virtual Event Canada: ACM; 2021. p. 610–23.
doi: 10.1145/3442188.3445922

7

https://projectinfolit.org/publications/algorithm-study

