A Pilot Evaluation of Open-Weight Large Language Models for Screening RNA-seq Metadata in Public Databases
Manuscript statistics
Major revisionpanel verdict ยท 2026-09-03
Manuscript Statistics โ 1 4 1 5 1 6 1 7 1 8 1 9 20 21 22 23 24 25 26 27 28
Measured deterministically at ingest, with no model involved. Every figure describes the converted text the panel read, not the PDF.
How the file converted
- Format: markdown via rustypaper 0.2.0
- Section map: read from the document model
- Conversion health: clean
- Fused tokens: 0.0 per 1000 words
- Hyphenated line breaks: 0.0 per 1000 words
- Lost sentence spaces: 0.0 per 1000 words
- Markdown headings emitted by the converter: 74
- Blank-line-separated blocks: 299
- Text matching no known section: 11%
Size
- Words: 11,678
- Main text (excluding references): 11,335
- Reference list: 343 words
- Sentences: 478
- Display equations: 0
- Table rows: 0
Prose
- Sentence length: mean 25.44, median 23.0, 90th percentile 40.0 words
- Sentences over 40 words: 9%
- Lexical diversity (MATTR): 0.4867
- Passive constructions: 0.3515 per sentence (regex approximation)
Claims and evidence
- In-text citations: too few detected to count reliably โ this venue most likely sets them as superscript numerals, which convert to bare digits
- Bibliography: 5 entries typed by the converter
- Numbers: 191.47 per 1000 words
- Hedging language: 7.96 per 1000 words
- Amplifying language: 2.57 per 1000 words
- p-values: 11 exact, 6 reported only as a threshold
By section
Measured over each section separately. The bibliography is left out: hedging and sentence length over a reference list describe a dozen journals' house styles rather than this manuscript.
| Section | Words | Sentences | Mean sentence | Citations/1k | Hedges/1k | Boosters/1k |
|---|---|---|---|---|---|---|
| _preamble | 1,300 | 54 | 24.53 | 0.0 | 11.54 | 2.31 |
| results | 8,761 | 341 | 25.69 | 0.0 | 8.9 | 3.08 |
| supplementary | 91 | 1 | 91.0 | 0.0 | 0.0 | 0.0 |
| acknowledgements | 1,119 | 19 | 58.89 | 0.0 | 0.0 | 0.0 |