OpenFindings
Human-ledAutomated checks: pending

What Makes a Scientific Question Succeed? Predicting Future Attention, Resolution, and Premise Revision from Question Structure and Evidence Context

Hui Mao · Version 1 · CC-BY-4.0

Abstract

Large language models are increasingly asked to propose “important open questions”, and are typically evaluated by asking another model how good those questions feel. We replace taste with history. We construct a temporally grounded dataset of 980 astronomy research questions frozen at five historical cutoffs (2012–2020), spanning eight subfields and four distinct generation sources — evidence-tension mining, direct LLM elicitation, author-stated future work, and weak/negative controls — on top of the complete arXiv astro-ph metadata record (313,189 papers, 2005–2025). Each question carries a six-group cutoff-time feature record (text structure, evidence context, novelty, falsifiability, tractability, field environment) and a tiered outcome label derived from the five years of literature that actually followed. Under a strict, source-blinded outcome judge, 52.7% of questions were substantively addressed; negative-control questions were addressed least (32%) and direct-LLM questions most (88%), the latter carrying a documented parametric-leakage caveat. Interpretable models trained on 2012–2016 cutoffs predict 2020 outcomes at AUC 0.71 and transfer to entirely held-out subfields at AUC 0.65–0.71. Measured label noise implies a perfect-oracle ceiling of AUC 0.76 on these labels, so the model captures 83% of the attainable range above chance. Across univariate, controlled, ablation, and stability analyses, the robust predictors of future attention are prior community recognition (overlap with pre-cutoff review articles, β = +0.34, p < 10−4), the density of directly related prior evidence (β = +0.26), and explicitly stated competing hypotheses (β = +0.17) — while high entity-specificity predicts less and slower engagement under literature-bounded labels. Five of six pre-stated mechanism hypotheses, including “evidence tension beats novelty”, are not supported at scale — a result only visible because the dataset is large, multi-cutoff, and control-laden. Answering the structural-prior question posed in [23], we further show that within the tension-mined subset the generating evidence-graph structure itself is predictive: breadth of the paper cluster in tension predicts uptake (β = +0.45/SD), explicit contradiction predicts engagement, and quantified, cutoff-persistent tensions carry every resolution and premise refutation — a learnable, mining-time prior for structure-first question generation. Finally, a complete replication of the protocol in a maximally distant second domain — anti-aging skincare-ingredient research (19,193 PubMed abstracts, 1,043 backtested questions, a deterministic template generator and a checkable-condition judge) — reproduces the predictability (AUC 0.75–0.84), the control ordering, the breadth effect, the hypothesis failures, and the reliability ceiling (κ = 0.22 vs. 0.21), while the sign of specificity reverses exactly as the engagement-cost account predicts. All data, code, labels, and judge rationales for both domains are released.

Transparency note. Automated integrity checks are separate from independent scientific peer review.

Research object

**Methods:** We construct a temporally grounded benchmark of 980 astronomy research questions across five historical cutoffs and eight subfields, using 313,189 arXiv records. Questions are generated from evidence tensions, direct LLM elicitation, author-stated future work, and negative controls. We extract cutoff-time structural and evidence-context features, evaluate subsequent five-year outcomes, and train interpretable predictive models using temporal and cross-subfield holdouts. We further replicate the framework on 1,043 questions derived from 19,193 PubMed abstracts concerning anti-aging skincare ingredients.

**Results:** The primary outcome judge identifies substantive future engagement for 52.7% of astronomy questions. Structured features achieve an out-of-time AUC of 0.713, with cross-subfield generalization. Prior recognition in review literature, density of related evidence, and explicit competing hypotheses are positive predictors of future attention. Five of six pre-stated mechanism hypotheses are not supported at scale. Evidence-cluster breadth and explicit contradictions provide additional predictive signals within tension-mined questions. Cross-domain replication achieves AUC values of approximately 0.75–0.84. Outcome-label reliability remains limited, with inter-judge agreement of approximately κ = 0.21–0.22.

**Conclusions:** Future scientific attention can be predicted to a meaningful degree from measurable properties of research questions and their existing evidence contexts. Community recognition and evidence density are more reliable predictors than intuitive assessments of novelty or tension alone. However, predicted attention should not be equated with intrinsic scientific value, and outcome measurement remains sensitive to retrieval coverage, judge reliability, and potential temporal knowledge leakage.