A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases
Decision letter
Decision Letter
VERDICT: major
Summary of Evaluation
This manuscript benchmarks 17 large language models (12 locally executed open-weight conditions, 5 provider-hosted closed models) on a single, clearly specified metadata-screening task: deciding whether a public RNA-seq project contains exogenous ABA-treated Arabidopsis thaliana samples together with matched untreated controls. Against a keyword-retrieval reference (F1 = 0.59), LLM classification substantially improves precision, and several 2025 open-weight models (gpt-oss-120b_low, F1 = 0.992; qwen3-next-80b-a3b-thinking, F1 = 0.984) match the best closed models tested. The paper also reports runtime, an F1–runtime trade-off for local execution, a five-run reproducibility check on a 50-project subset, and an AUPRC analysis of self-reported confidence scores.
The panel was consistent and broadly positive. The workflow is genuinely reusable: code is on GitHub under MIT licence, prompts, integrated metadata inputs, per-model outputs, runtime tables and the full ground-truth label table are deposited on Figshare, and the hardware, inference stack and version numbers are all stated. The limitations section is unusually candid — task narrowness, single-curator labels, discreteness of the self-reported probabilities, cross-session probability drift, and the unevaluated extraction feature are all disclosed by the authors rather than discovered by reviewers. That candour is a strength and I have not treated any author-flagged weakness as a new defect.
Two things keep this from a minor revision.
First, the validity of the label set. Every number in the paper — including the headline "F1 > 0.98" and the "F1 = 1.00 under HIGH confidence" result — is measured against labels assigned by one curator, from the same integrated metadata text supplied to the models, with no independent agreement estimate. The advocate's defence in the debate (that the task is not reducible to keyword matching, given the baseline's 0.42 precision and gpt-3.5-turbo's F1 = 0.630 on identical input) is a fair rebuttal to the charge of triviality, but it is not evidence about label reliability, which is the actual gap. Because the labels are deposited and the criteria explicit, a reader can inspect them — this mitigates but does not substitute for a measured agreement estimate on a task the authors themselves describe as operating on "incomplete, ambiguous, or inconsistent" metadata. A bounded second-curator exercise on a subset would settle it.
Second, the confidence-filtering conclusion rests on cut-offs (p < 0.25, p > 0.75) that the authors correctly describe as predefined rather than optimised, but which fall between the seven discrete values the models actually emit (0.05, 0.40, 0.60, 0.75, 0.80, 0.90, 0.95). Whether the "F1 = 1.00 in the HIGH subset" result survives at 0.20/0.80 or 0.40/0.60 is unknown and is a reanalysis whose outcome could change the claim. The deposited outputs make this cheap to run, but its result is not predictable from what is currently reported — which is what makes this a major rather than a minor revision.
Both items are bounded and neither threatens the design. I expect a revision to be straightforward, and I would be glad to see this paper listed.
Two corrections to the referee record, since these reviews are published in full. (i) The contribution reviewer states that Ikeda et al. (2025), "Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database," is uncited; it is in fact reference 13 (GigaScience 14, giaf070). The framing criticism about acknowledging LLM metadata curation as an existing line of work still stands, but the specific omission does not. (ii) The citation-integrity audit reports the reference list as truncated to five fragments. Reading the manuscript directly, this appears to be a text-extraction artefact of the two-column/line-numbered layout — the author names and titles for references 1–5 are present but split from their journal strings, and references 6–19 are complete with DOIs. I have therefore not treated "14 missing references" as an author-facing defect, and have retained only the specific, checkable citation items below.
Required Revisions
-
Provide a label-reliability estimate, or bound the claims to the label set as published. Have a second annotator, working independently from the same integrated metadata text and the same four criteria (lines 800–813), label a random subset of at least 30 of the 150 projects. Report percent agreement and Cohen's κ, list the disagreements with the adjudicated outcome, and state whether any top-ranked model's F1 ordering changes if the disputed items are relabelled. If a second annotator is genuinely unavailable, say so explicitly and instead (a) publish the curator's list of borderline projects with the reasoning applied to each, and (b) restate all performance claims in the abstract and results as being relative to a single-curator label set, so that no reader takes "near-perfect classification" as agreement with an adjudicated standard.
-
Report threshold sensitivity for the confidence-based grouping. Using the already-deposited self-reported probabilities, recompute HIGH-condition precision, recall, F1 and n′ across a grid of cut-offs (e.g. lower bounds 0.20/0.25/0.30/0.40 crossed with upper bounds 0.60/0.70/0.75/0.80) for all 17 models under prompt 2. State which conclusions are robust — in particular whether gpt-oss-120b_high, qwen3-30b-a3b-thinking-2507 and qwen3-next-80b-a3b-thinking retain F1 = 1.00, and how n′ moves. Because the outputs are discrete, also make explicit which discrete values each cut-off admits or excludes. If the perfect-F1 result is threshold-fragile, the corresponding claim in the abstract and discussion must be weakened accordingly.
-
Qualify the two abstract-level claims that currently outrun the evidence.
- Confidence: the abstract sentence "self-reported confidence scores may help identify high-confidence cases that can be prioritized for automated processing" must carry the model-dependence qualifier that the body already supplies. State in the abstract that this held for the higher-performing models tested but failed for gpt-3.5-turbo-0125 (HIGH-condition F1 = 0.286) and gpt-4o-mini-2024-07-18 (HIGH-condition F1 = 0.000), and that per-model validation on labelled data is a precondition for operational use.
- Parity: "comparable to that of closed models" should specify which closed models. gpt-5.1-2025-11-13 (F1 = 0.984) and gemini-2.5-pro (F1 = 1.000) match or exceed the best open-weight condition; the clean result is that 2025 open-weight models exceed the 2023–2024 closed models tested and approach the current ones on this task. Similarly, the title/abstract framing "for screening RNA-seq metadata in public databases" should be narrowed to the evaluated setting (one organism, one treatment, bulk RNA-seq, one retrieval strategy, n = 150).
-
Fix the AUPRC reporting to match the data. Given that scores take seven discrete values, state the tie-handling/interpolation used by your implementation, and either (a) report AUPRC together with a rank-agreement statistic against the all-sample F1 ranking (e.g. Spearman ρ with a CI) rather than asserting "strong agreement" from a count of rank differences within ±1, or (b) drop the numerical AUPRC values and present the analysis purely as a ranking check. The existing caveat in the discussion is appropriate but does not by itself license reporting AUPRC values to three decimals.
-
State the design of the evaluation explicitly, and the provenance of the prompts. Say in the Methods that all 150 projects served as a single evaluation set with no train/validation/test split and no held-out data. Then state whether prompts 1 and 2 were fixed a priori from the task definition or refined after inspecting outputs on these same 150 projects. If they were refined, disclose it and discuss the resulting overfitting risk for the reported F1 values; this is the single most consequential unstated methodological choice in the paper.
-
Flag the quantization confound in the open-versus-closed comparison. Table 4 records the quantization scheme per model (MXFP4, MLX 4bit, Q4_K_M), but the text nowhere notes that quantized local models are being compared against full-precision provider-hosted models, and that no quantized-versus-unquantized check was performed. Add this as an explicit caveat where the open/closed comparison is drawn (Results and Discussion). A single-model ablation would be welcome but is not required.
-
Reframe the keyword baseline honestly. State plainly that "Keyword Search Only (Baseline)" is a recall-upper-bound reference computed by treating every retrieved candidate as positive, not a method any practitioner would use, and that its recall of 1.00 is true only within the retrieved pool (you say this once at line 141 — it needs to be equally clear in the abstract and Table 1 caption). If it can be done from the deposited metadata text without new inference, add one cheap deterministic comparator — for example a rule requiring co-occurrence of an ABA-treatment term and a control/mock term at sample level — so the improvement attributable to semantic judgement, rather than to any filtering at all, can be seen.
-
Repair the specific citation defects. (a) Reference 11 (Vaswani et al., "Attention is all you need") carries
doi:10.65215/2q58a426, which is not the arXiv DOI prefix; verify and correct. While doing so, check that this reference is the right support for the claim at lines 88–90 that "LLMs learn from large text corpora and can capture complex patterns in natural languages" — the transformer paper supports the architecture, not that statement as written. (b) Reference 16 ("Arena Leaderboard", https://arena.ai/ja/leaderboard) needs an access date in the reference entry, and references 17 and 18 need arXiv identifiers. (c) Confirm that the reference list renders completely in the deposited PDF; our automated extraction returned entries 1–5 in fragments, which we believe is a layout artefact, but please check. -
Position the contribution against the existing LLM-metadata-curation literature in the introduction, not only the discussion. State plainly that LLM-assisted extraction and classification of repository metadata is an established and active line of work (references 13 and 14 already cover part of it), and that the delta here is a controlled open-weight-versus-closed benchmark under local execution on a defined plant-stress screening task. If concurrent work on LLM classification of sequencing-record sample origin or on automated ChIP-seq metadata extraction is relevant, cite it; if you judge it out of scope, that is an acceptable answer, but the framing should not leave a reader with the impression that the approach itself is new.
-
Add dispersion to the runtime results. Supplementary Table 5 and Figure 2(e) report mean per-project runtime only. Add SD or median and range, and note that closed-model latencies (provider-managed hardware, network round-trip) are not comparable to local wall-clock times, so the two should not be read off the same axis.
Minor Suggestions
- Spot-check the sample-attribute extraction feature on 10–20 projects against the integrated metadata text and report a rough error rate. You are right that full evaluation is hard, but a bounded spot-check converts an unevaluated feature into a weakly evaluated one, which is considerably more useful to a reader.
- Report the coverage cost of HIGH-confidence filtering as a fraction (e.g. qwen3-30b-a3b-thinking-2507 retains 88/150 = 59% at perfect F1), since an operational reader needs coverage and reliability side by side.
- Add binomial or bootstrap confidence intervals to the F1 values in Table 1. At n = 150 with 63 positives, several of the top-ranked differences (0.992 vs 0.984 vs 0.968) are almost certainly within sampling noise, and stating this would strengthen rather than weaken your argument.
- Extend the five-run reproducibility check to at least one dense model and one instruct model, and to the full 150 projects for one model, so that the drift you honestly report for gpt-oss-120b_low can be placed in context.
- Report input token counts after metadata integration (distribution across the 150 projects) and note the compression achieved by consolidating redundant sample fields; this bears on whether context length constrained any model.
- Give concrete cost figures: total API spend per model condition, and local wall-clock hours for the full 150-project sweep. The cost argument in the discussion is currently qualitative.
- Consider reporting whether performance differed between the 104 projects present in both GEO and BioProject and the 46 unique to one source, as a cheap probe of metadata-quality sensitivity.
- A second organism/treatment pair (e.g. rice/drought, given reference 6) would be the single most valuable extension for a follow-up paper. It is not required here — the pilot framing is honest and the claims are scaled to it — but the manuscript would be more persuasive if it named this as the concrete next validation step rather than a general call for "additional validation."