Review v2/r1 · round 1 · manuscript v2
A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases
Panel readout
5 specialists · scored 1–5
Legacy scaled score
Range 4.0–5.0, a spread of 1.0.
- ethics5.0Confidence 4 of 5
- scientific validity4.0Confidence 4 of 5
- reporting reproducibility4.0Confidence 4 of 5
- data analysis4.0Confidence 4 of 5
- contribution context4.0Confidence 5 of 5
Score is the referee's assessment of the work.Confidence is how sure that referee was of its own reading, recorded separately and never combined. The editor's verdict is its own judgment of the reports, not a threshold applied to this mean. The legacy aggregate is the historical panel mean multiplied by 20. It is not comparable to the Editor-in-Chief's publication readiness score.
Abstract
as posted by the authors
Although the Gene Expression Omnibus and other public repositories are expanding rapidly, curation across these databases has not kept pace. Data reuse is often hindered by unstandardized metadata comprising unstructured text. To address this, we developed a workflow that combines retrieval via application programming interfaces with semantic filtering using large language models (LLMs) to support metadata screening as an initial step in broader curation workflows. As a focused pilot evaluation, we benchmarked multiple LLMs using metadata from 150 candidate Arabidopsis RNA-seq projects to classify projects containing exogenous ABA-treated samples and matched untreated controls. Simple keyword searches yielded many false positives (F1=0.59); classification using LLMs significantly improved performance. Several open-weight models achieved near-perfect classification performance in this defined task (F1>0.98), comparable to that of closed models. We also found that, for some high-performing models, self-reported confidence scores may help identify high-confidence cases that can be prioritized for automated processing. These results suggest that open-weight LLMs can support scalable metadata screening in local environments as an initial step in broader curation workflows, providing a foundation for accelerating public dataset reuse.
The review
- SummaryThe panel's assessment in brief.
- Decision letterThe editor's verdict and what it requires.
- Desk screenWhether the submission cleared the bar for full review.
- Advocate / skeptic debateThe case for and against, in full.
- Debate synthesisThe condensed account of the debate the editor read.
- Venue suggestionsWhere this might be submitted.
- Manuscript statisticsDeterministic counts over the text the panel read.
Specialist reports
Editorial audits
Factual checklists, not opinions. They skip the debate and go straight to the editor.
The text the panel read
counted, not judged
Counted at ingest, no model involved. These describe theconverted text the referees read, not your PDF.
Size
- Words
- 11,678
- Main text
- 11,335
- Sentences
- 478
- Display equations
- 0
excluding references
Sentences
- Median sentence
- 23 words
- Longest tenth
- 40 words
- Over 40 words
- 9%
- Passive
- ~0.3515/sentence
regex approximation
Evidence on the page
- Citations
- not countable
- Numbers
- 191.47
- p-values
- 11 exact, 6 threshold
this venue most likely sets them as superscript numerals, which convert to bare digits
per 1000 words
Hedging against amplifying
Per 1000 words. Softening ("may", "suggests") against strengthening ("clearly", "demonstrates"). No referee saw these.
Provenance
PeerReviewAgents 0.5.1
Review record
v2/r1
initial
Reviewed file
v2 · 2209 KiB
sha256 f4abfcf42553c618cc971f4e…
Debate rounds
2
Run cost
$1.93
Manuscript read as
markdown
Converted by rustypaper 0.2.0. Quotations match the manuscript text.
Desk screen
triage gate
In Silico commit
c6cd6372
Configuration
sha256 8c26c06db029
Journal profile
insilico
sha256 c416856b2c4a
Which model wrote which report
| Stage | Model |
|---|---|
| Editorial audits (×2) | claude-haiku-4-5 |
| Advocate / skeptic | claude-sonnet-5 |
| Specialist reviewers (×5) + desk screen | claude-haiku-4-5 |
| Editor, debate synthesizer | claude-opus-5 |
| debate synthesizer (override) | claude-sonnet-5 |
| journal recommender (override) | claude-haiku-4-5 |
The widest fan-out runs on the cheapest model. Only the agents that decide the verdict run on the most capable one.
What the referees looked up (14 searches)
| Agent | Search | Answered by | Hits |
|---|---|---|---|
| reviewer contribution context | find related workLLM large language model metadata curation screening classification | N/A | 10 |
| find related workRNA-seq metadata extraction standardization public databases | N/A | 10 | |
| search preprintsLLM metadata curation biological data extraction | N/A | 5 | |
| search preprintslanguage model GEO SRA dataset screening filtering | N/A | 4 | |
| find related workIkeda biological terms BioSample LLM extraction metadata | N/A | 2 | |
| find related workPubTator semantic search biomedical literature AI | N/A | 5 | |
| search preprintsRiquelme ontology annotation biological samples LLM | N/A | 0 | |
| find related workopen-weight models local inference metadata classification | N/A | 5 | |
| find related workCistromeMeta ChIP-seq metadata extraction GEO | N/A | 5 | |
| search preprintsGEOmetadb pysRADB metadata retrieval query | N/A | 0 | |
| find related workprompt engineering classification precision recall trade-off | N/A | 5 | |
| find related workWei PubTator 3.0 AI-powered literature biomedical | N/A | 5 | |
| find related workRiquelme annotation biological samples ontologies LLM | N/A | 5 | |
| search preprintsGaio microbiome sample origins LLM semantic classification | N/A | 1 |
Run against arXiv, Semantic Scholar, PubMed and bioRxiv while the review was being written. A search returning zero hits is kept: it is the evidence behind a referee saying it found no prior art.
What each agent cost
| Agent | USD |
|---|---|
| editor | $0.7074 |
| skeptic | $0.4399 |
| advocate | $0.3401 |
| reviewer contribution context | $0.1201 |
| debate synthesizer | $0.1051 |
| audit citation integrity | $0.0647 |
| desk screen | $0.0534 |
| audit methods completeness | $0.0259 |
| reviewer data analysis | $0.0180 |
| reviewer ethics | $0.0158 |
| reviewer reporting reproducibility | $0.0149 |
| reviewer scientific validity | $0.0138 |
| journal recommender | $0.0134 |
Cite this review
Permanent: this review only
This URL is a permanent link to this specific review, and will not change.
In Silico (2026). Review of "A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases". In Silico. https://pgarrett-scripps.github.io/insilico/reviews/2026/a-pilot-evaluation-of-open-weight-large-language-10-64898-2026-02-16-706241/v2/r1/
@misc{insilico-a-pilot-evaluation-of-open-weight-large-language-10-64898-2026-02-16-706241-v2-r1,
title = {Review of {A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases}},
author = {{In Silico}},
year = {2026},
howpublished = {In Silico, an AI-refereed overlay journal},
url = {https://pgarrett-scripps.github.io/insilico/reviews/2026/a-pilot-evaluation-of-open-weight-large-language-10-64898-2026-02-16-706241/v2/r1/},
note = {Machine-generated peer review of doi:10.64898/2026.02.16.706241 v2. Produced by PeerReviewAgents 0.5.1. Produced by PeerReviewAgents, doi:10.5281/zenodo.21781895.}
}Please cite the preprint itself as well. This reviews that work, it does not replace it. The review is machine-generated and advisory. If you are citing it as evidence about the paper, say so explicitly.