A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases

Manuscript statistics

Major revisionpanel verdict ยท 2026-09-03

Manuscript Statistics โ€” 1 4 1 5 1 6 1 7 1 8 1 9 20 21 22 23 24 25 26 27 28

Measured deterministically at ingest, with no model involved. Every figure describes the converted text the panel read, not the PDF.

How the file converted

  • Format: markdown via rustypaper 0.2.0
  • Section map: read from the document model
  • Conversion health: clean
  • Fused tokens: 0.0 per 1000 words
  • Hyphenated line breaks: 0.0 per 1000 words
  • Lost sentence spaces: 0.0 per 1000 words
  • Markdown headings emitted by the converter: 74
  • Blank-line-separated blocks: 299
  • Text matching no known section: 11%

Size

  • Words: 11,678
  • Main text (excluding references): 11,335
  • Reference list: 343 words
  • Sentences: 478
  • Display equations: 0
  • Table rows: 0

Prose

  • Sentence length: mean 25.44, median 23.0, 90th percentile 40.0 words
  • Sentences over 40 words: 9%
  • Lexical diversity (MATTR): 0.4867
  • Passive constructions: 0.3515 per sentence (regex approximation)

Claims and evidence

  • In-text citations: too few detected to count reliably โ€” this venue most likely sets them as superscript numerals, which convert to bare digits
  • Bibliography: 5 entries typed by the converter
  • Numbers: 191.47 per 1000 words
  • Hedging language: 7.96 per 1000 words
  • Amplifying language: 2.57 per 1000 words
  • p-values: 11 exact, 6 reported only as a threshold

By section

Measured over each section separately. The bibliography is left out: hedging and sentence length over a reference list describe a dozen journals' house styles rather than this manuscript.

SectionWordsSentencesMean sentenceCitations/1kHedges/1kBoosters/1k
_preamble1,3005424.530.011.542.31
results8,76134125.690.08.93.08
supplementary91191.00.00.00.0
acknowledgements1,1191958.890.00.00.0

โ† All documents in this review