Large language models (LLMs) are increasingly considered for scientific evaluation workflows, including editorial triage, screening, and ranking. This raises a fairness and reliability question: are LLM-based evaluators counterfactually invariant to extrinsic author-context metadata when manuscript content is held fixed? We conduct a controlled counterfactual audit of 80 recent English-language arXiv manuscripts from computer science, mathematics, physics, and quantitative biology. For each manuscript, we append a corresponding-author profile in a 2 × 2 × 2 factorial design varying name-coded identity, institution prestige tier, and bibliometric profile, alongside a no-metadata BLIND baseline. Six heterogeneous LLM evaluators assign ratings on 14 review-style criteria and three summary outcomes: content quality, perceived prestige/impact, and acceptance recommendation. Estimates are computed at the manuscript level after averaging repeated calls within model and pooling evaluator models with equal weight; uncertainty is assessed with bootstrap confidence intervals, sign-flip tests, false-discovery-rate control, and model heterogeneity analyses. Evaluator outputs are not invariant to author metadata. Bibliometric and institution-tier cues systematically increase perceived prestige/impact and acceptance-related judgments, while effects on contentquality judgments and name-coded effects are smaller. Ranking analyses show that metadata perturbations can displace manuscripts and alter top-K shortlist membership. These results identify a concrete reliability risk for LLM-assisted scientific evaluation pipelines.

Author metadata affects Large Language Model scores in scientific peer review

Marco Rospocher
2026-01-01

Abstract

Large language models (LLMs) are increasingly considered for scientific evaluation workflows, including editorial triage, screening, and ranking. This raises a fairness and reliability question: are LLM-based evaluators counterfactually invariant to extrinsic author-context metadata when manuscript content is held fixed? We conduct a controlled counterfactual audit of 80 recent English-language arXiv manuscripts from computer science, mathematics, physics, and quantitative biology. For each manuscript, we append a corresponding-author profile in a 2 × 2 × 2 factorial design varying name-coded identity, institution prestige tier, and bibliometric profile, alongside a no-metadata BLIND baseline. Six heterogeneous LLM evaluators assign ratings on 14 review-style criteria and three summary outcomes: content quality, perceived prestige/impact, and acceptance recommendation. Estimates are computed at the manuscript level after averaging repeated calls within model and pooling evaluator models with equal weight; uncertainty is assessed with bootstrap confidence intervals, sign-flip tests, false-discovery-rate control, and model heterogeneity analyses. Evaluator outputs are not invariant to author metadata. Bibliometric and institution-tier cues systematically increase perceived prestige/impact and acceptance-related judgments, while effects on contentquality judgments and name-coded effects are smaller. Ranking analyses show that metadata perturbations can displace manuscripts and alter top-K shortlist membership. These results identify a concrete reliability risk for LLM-assisted scientific evaluation pipelines.
2026
Large language models, Scientific evaluation, Peer review, Algorithmicauditing, Prestige bias, Bibliometric bias
File in questo prodotto:
File Dimensione Formato  
s44163-026-02268-y.pdf

accesso aperto

Licenza: Creative commons
Dimensione 4.12 MB
Formato Adobe PDF
4.12 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11562/1205647
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact