Large language models (LLMs) are increasingly considered for scientific evaluation workflows, including editorial triage, screening, and ranking. This raises a fairness and reliability question: are LLM-based evaluators counterfactually invariant to extrinsic author-context metadata when manuscript content is held fixed? We conduct a controlled counterfactual audit of 80 recent English-language arXiv manuscripts from computer science, mathematics, physics, and quantitative biology. For each manuscript, we append a corresponding-author profile in a 2 × 2 × 2 factorial design varying name-coded identity, institution prestige tier, and bibliometric profile, alongside a no-metadata BLIND baseline. Six heterogeneous LLM evaluators assign ratings on 14 review-style criteria and three summary outcomes: content quality, perceived prestige/impact, and acceptance recommendation. Estimates are computed at the manuscript level after averaging repeated calls within model and pooling evaluator models with equal weight; uncertainty is assessed with bootstrap confidence intervals, sign-flip tests, false-discovery-rate control, and model heterogeneity analyses. Evaluator outputs are not invariant to author metadata. Bibliometric and institution-tier cues systematically increase perceived prestige/impact and acceptance-related judgments, while effects on contentquality judgments and name-coded effects are smaller. Ranking analyses show that metadata perturbations can displace manuscripts and alter top-K shortlist membership. These results identify a concrete reliability risk for LLM-assisted scientific evaluation pipelines.
Author metadata affects Large Language Model scores in scientific peer review
Marco Rospocher
2026-01-01
Abstract
Large language models (LLMs) are increasingly considered for scientific evaluation workflows, including editorial triage, screening, and ranking. This raises a fairness and reliability question: are LLM-based evaluators counterfactually invariant to extrinsic author-context metadata when manuscript content is held fixed? We conduct a controlled counterfactual audit of 80 recent English-language arXiv manuscripts from computer science, mathematics, physics, and quantitative biology. For each manuscript, we append a corresponding-author profile in a 2 × 2 × 2 factorial design varying name-coded identity, institution prestige tier, and bibliometric profile, alongside a no-metadata BLIND baseline. Six heterogeneous LLM evaluators assign ratings on 14 review-style criteria and three summary outcomes: content quality, perceived prestige/impact, and acceptance recommendation. Estimates are computed at the manuscript level after averaging repeated calls within model and pooling evaluator models with equal weight; uncertainty is assessed with bootstrap confidence intervals, sign-flip tests, false-discovery-rate control, and model heterogeneity analyses. Evaluator outputs are not invariant to author metadata. Bibliometric and institution-tier cues systematically increase perceived prestige/impact and acceptance-related judgments, while effects on contentquality judgments and name-coded effects are smaller. Ranking analyses show that metadata perturbations can displace manuscripts and alter top-K shortlist membership. These results identify a concrete reliability risk for LLM-assisted scientific evaluation pipelines.| File | Dimensione | Formato | |
|---|---|---|---|
|
s44163-026-02268-y.pdf
accesso aperto
Licenza:
Creative commons
Dimensione
4.12 MB
Formato
Adobe PDF
|
4.12 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



