Review v2/r1 · round 1 · manuscript v2

A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases

Shintani, M., Andrade, D., Bono, H.

Major revisionadvisory · no human graded this paper
DOI
10.64898/2026.02.16.706241
Reviewed
2026-09-03

Panel readout

5 specialists · scored 1–5

84/ 100

Legacy scaled score

Range 4.0–5.0, a spread of 1.0.

  1. ethics5.0Confidence 4 of 5
  2. scientific validity4.0Confidence 4 of 5
  3. reporting reproducibility4.0Confidence 4 of 5
  4. data analysis4.0Confidence 4 of 5
  5. contribution context4.0Confidence 5 of 5

Score is the referee's assessment of the work.Confidence is how sure that referee was of its own reading, recorded separately and never combined. The editor's verdict is its own judgment of the reports, not a threshold applied to this mean. The legacy aggregate is the historical panel mean multiplied by 20. It is not comparable to the Editor-in-Chief's publication readiness score.

Abstract

as posted by the authors

Although the Gene Expression Omnibus and other public repositories are expanding rapidly, curation across these databases has not kept pace. Data reuse is often hindered by unstandardized metadata comprising unstructured text. To address this, we developed a workflow that combines retrieval via application programming interfaces with semantic filtering using large language models (LLMs) to support metadata screening as an initial step in broader curation workflows. As a focused pilot evaluation, we benchmarked multiple LLMs using metadata from 150 candidate Arabidopsis RNA-seq projects to classify projects containing exogenous ABA-treated samples and matched untreated controls. Simple keyword searches yielded many false positives (F1=0.59); classification using LLMs significantly improved performance. Several open-weight models achieved near-perfect classification performance in this defined task (F1>0.98), comparable to that of closed models. We also found that, for some high-performing models, self-reported confidence scores may help identify high-confidence cases that can be prioritized for automated processing. These results suggest that open-weight LLMs can support scalable metadata screening in local environments as an initial step in broader curation workflows, providing a foundation for accelerating public dataset reuse.

The review

Specialist reports

Editorial audits

Factual checklists, not opinions. They skip the debate and go straight to the editor.

The text the panel read

counted, not judged

Counted at ingest, no model involved. These describe theconverted text the referees read, not your PDF.

Size

Words
11,678
Main text
11,335

excluding references

Sentences
478
Display equations
0

Sentences

Median sentence
23 words
Longest tenth
40 words
Over 40 words
9%
Passive
~0.3515/sentence

regex approximation

Evidence on the page

Citations
not countable

this venue most likely sets them as superscript numerals, which convert to bare digits

Numbers
191.47

per 1000 words

p-values
11 exact, 6 threshold

Hedging against amplifying

Per 1000 words. Softening ("may", "suggests") against strengthening ("clearly", "demonstrates"). No referee saw these.

Hedging7.96
Amplifying2.57

Provenance

PeerReviewAgents 0.5.1

Review record

v2/r1

initial

Reviewed file

v2 · 2209 KiB

sha256 f4abfcf42553c618cc971f4e…

Debate rounds

2

Run cost

$1.93

Manuscript read as

markdown

Converted by rustypaper 0.2.0. Quotations match the manuscript text.

Desk screen

triage gate

In Silico commit

c6cd6372

Configuration

sha256 8c26c06db029

Journal profile

insilico

sha256 c416856b2c4a

Which model wrote which report
StageModel
Editorial audits (×2)claude-haiku-4-5
Advocate / skepticclaude-sonnet-5
Specialist reviewers (×5) + desk screenclaude-haiku-4-5
Editor, debate synthesizerclaude-opus-5
debate synthesizer (override)claude-sonnet-5
journal recommender (override)claude-haiku-4-5

The widest fan-out runs on the cheapest model. Only the agents that decide the verdict run on the most capable one.

What the referees looked up (14 searches)
AgentSearchAnswered byHits
reviewer contribution contextfind related workLLM large language model metadata curation screening classificationN/A10
find related workRNA-seq metadata extraction standardization public databasesN/A10
search preprintsLLM metadata curation biological data extractionN/A5
search preprintslanguage model GEO SRA dataset screening filteringN/A4
find related workIkeda biological terms BioSample LLM extraction metadataN/A2
find related workPubTator semantic search biomedical literature AIN/A5
search preprintsRiquelme ontology annotation biological samples LLMN/A0
find related workopen-weight models local inference metadata classificationN/A5
find related workCistromeMeta ChIP-seq metadata extraction GEON/A5
search preprintsGEOmetadb pysRADB metadata retrieval queryN/A0
find related workprompt engineering classification precision recall trade-offN/A5
find related workWei PubTator 3.0 AI-powered literature biomedicalN/A5
find related workRiquelme annotation biological samples ontologies LLMN/A5
search preprintsGaio microbiome sample origins LLM semantic classificationN/A1

Run against arXiv, Semantic Scholar, PubMed and bioRxiv while the review was being written. A search returning zero hits is kept: it is the evidence behind a referee saying it found no prior art.

What each agent cost
AgentUSD
editor$0.7074
skeptic$0.4399
advocate$0.3401
reviewer contribution context$0.1201
debate synthesizer$0.1051
audit citation integrity$0.0647
desk screen$0.0534
audit methods completeness$0.0259
reviewer data analysis$0.0180
reviewer ethics$0.0158
reviewer reporting reproducibility$0.0149
reviewer scientific validity$0.0138
journal recommender$0.0134

Cite this review

Permanent: this review only

This URL is a permanent link to this specific review, and will not change.

Plain text
In Silico (2026). Review of "A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases". In Silico. https://pgarrett-scripps.github.io/insilico/reviews/2026/a-pilot-evaluation-of-open-weight-large-language-10-64898-2026-02-16-706241/v2/r1/
BibTeX
@misc{insilico-a-pilot-evaluation-of-open-weight-large-language-10-64898-2026-02-16-706241-v2-r1,
  title        = {Review of {A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases}},
  author       = {{In Silico}},
  year         = {2026},
  howpublished = {In Silico, an AI-refereed overlay journal},
  url          = {https://pgarrett-scripps.github.io/insilico/reviews/2026/a-pilot-evaluation-of-open-weight-large-language-10-64898-2026-02-16-706241/v2/r1/},
  note         = {Machine-generated peer review of doi:10.64898/2026.02.16.706241 v2. Produced by PeerReviewAgents 0.5.1. Produced by PeerReviewAgents, doi:10.5281/zenodo.21781895.}
}

Please cite the preprint itself as well. This reviews that work, it does not replace it. The review is machine-generated and advisory. If you are citing it as evidence about the paper, say so explicitly.