Review v2 · round 1 · manuscript v2
On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules
Panel readout
4 of 5 specialists scored · 1–5
Legacy scaled score
Range 3.0–4.0, a spread of 1.0.
- data analysis4.0Confidence 4 of 5
- contribution context4.0Confidence 4 of 5
- reporting reproducibility3.0Confidence 4 of 5
- scientific validity3.0Confidence 4 of 5
- ethicsThe reviewer gave no score and did not say why.n/a
Score is the referee's assessment of the work.Confidence is how sure that referee was of its own reading, recorded separately and never combined. The editor's verdict is its own judgment of the reports, not a threshold applied to this mean. The legacy aggregate is the historical panel mean multiplied by 20. It is not comparable to the Editor-in-Chief's publication readiness score. n/a means that referee found nothing in its dimension to judge here. It is left out of the mean rather than counted as a good score, which is what a forced number quietly became.
Abstract
as posted by the authors
Probabilistic forecast models can be machine-learned from data using loss functions based on scoring rules such as the Continuous Ranked Probability Score (CRPS). This note summarises a preliminary study comparing versions of AIFS-CRPS, a global weather forecast model, trained with different univariate and multivariate scoring rules that aim to explicitly represent scale-awareness in the loss function. In the first part, we compare the (almost) fair CRPS, a fair global energy score, and a graph energy score based on node neighbourhoods. Across standard verification metrics, forecast skill is broadly similar. In the extratropics we find only small differences, while in the tropics the graph energy score setup performs somewhat better and the global energy score shows some degradation. These results suggest that multivariate scores are a viable alternative to CRPS-based training for global machine-learned weather forecasting. In the second part of the study, we analyse how different scoring rules and scale-aware loss constraints shape the spectra of forecast fields. It is apparent that any form of explicit scale-awareness improves realism. Here, the largest differences are likely associated with different effective weights per scale.
The review
- SummaryThe panel's assessment in brief.
- Decision letterThe editor's verdict and what it requires.
- Desk screenWhether the submission cleared the bar for full review.
- Advocate / skeptic debateThe case for and against, in full.
- Debate synthesisThe condensed account of the debate the editor read.
- Venue suggestionsWhere this might be submitted.
- Manuscript statisticsDeterministic counts over the text the panel read.
Specialist reports
Editorial audits
Factual checklists, not opinions. They skip the debate and go straight to the editor.
The text the panel read
counted, not judged
Counted at ingest, no model involved. These describe theconverted text the referees read, not your PDF.
Size
- Words
- 5,160
- Main text
- 4,381
- Sentences
- 238
- Display equations
- 48
excluding references
Sentences
- Median sentence
- 20 words
- Longest tenth
- 35 words
- Over 40 words
- 6%
- Passive
- ~0.2773/sentence
regex approximation
Evidence on the page
- Citations
- 64
- Numbers
- 52.33
- p-values
- 0 exact, 0 threshold
12.4 per 1000 words, numeric style
per 1000 words
Hedging against amplifying
Per 1000 words. Softening ("may", "suggests") against strengthening ("clearly", "demonstrates"). No referee saw these.
Provenance
PeerReviewAgents 0.5.0 · 1cf57690
Reviewed file
v2 · 5243 KiB
sha256 4de252a87afbf9423494b4b5…
Debate rounds
2
Run cost
$1.28
Manuscript read as
markdown
Converted by rustypaper 0.2.0. Quotations match the manuscript text.
Desk screen
triage gate
Which model wrote which report
| Stage | Model |
|---|---|
| Editorial audits (×2) | claude-haiku-4-5 |
| Advocate / skeptic | claude-sonnet-5 |
| Specialist reviewers (×5) + desk screen | claude-haiku-4-5 |
| Editor, debate synthesizer | claude-opus-5 |
| debate synthesizer (override) | claude-sonnet-5 |
| journal recommender (override) | claude-haiku-4-5 |
The widest fan-out runs on the cheapest model. Only the agents that decide the verdict run on the most capable one.
What the referees looked up (25 searches)
| Agent | Search | Answered by | Hits |
|---|---|---|---|
| reviewer contribution context | find related workCRPS machine learning weather forecasting ensemble | N/A | 5 |
| find related workenergy score multivariate probabilistic forecasting | N/A | 5 | |
| find related workgraph neural networks weather prediction scoring rules | N/A | 5 | |
| search preprintsscale-aware loss weather forecasting neural networks | N/A | 5 | |
| search preprintspatched energy score localized multivariate scoring | N/A | 0 | |
| find related workPacchiardi patched energy score generative networks | N/A | 5 | |
| find related workLang AIFS-CRPS ensemble weather forecasting | N/A | 5 | |
| find related workvariogram score multivariate ensemble forecasting | N/A | 5 | |
| search preprintsAIFS ensemble probabilistic weather machine learning | N/A | 0 | |
| find related workFourCastNet spectral loss weather forecasting | N/A | 5 | |
| find related workScheuerer variogram score proper scoring rules | N/A | 5 | |
| search preprintsmulti-scale loss formulation weather forecasting neural networks | N/A | 5 | |
| search preprintsgraph energy score localized multivariate forecast | N/A | 2 | |
| find related workfair CRPS Ferro ensemble forecast bias correction | N/A | 5 | |
| find related workGneiting Raftery proper scoring rules strictly proper | N/A | 5 | |
| search preprintsFGN FourCastNet Huracan machine learning weather 2025 | N/A | 0 | |
| search preprintsAlet Lam Battaglia skillful joint probabilistic weather 2025 | N/A | 0 | |
| search preprintsBonev Kurth FourCastNet 3 geometric probabilistic 2025 | N/A | 0 | |
| search preprintsLang Leutbecher Maciel multi-scale loss formulation 2025 | N/A | 0 | |
| find related workPacchiardi Adewoyin Dueben Dutta probabilistic forecasting generative networks 2024 | N/A | 5 | |
| find related worklocalized multivariate scoring patch-based energy score | N/A | 5 | |
| find related workspectral fidelity machine learning weather forecasting attention | N/A | 5 | |
| find related workgraph neural networks spatial localization weather prediction | N/A | 5 | |
| find related workPic Dombry Naveau Taillardat proper scoring rules aggregation transformation | N/A | 5 | |
| find related workZhdanov sparse attention spectral fidelity weather forecasting | N/A | 5 |
Run against arXiv, Semantic Scholar, PubMed and bioRxiv while the review was being written. A search returning zero hits is kept: it is the evidence behind a referee saying it found no prior art.
What each agent cost
| Agent | USD |
|---|---|
| editor | $0.4351 |
| skeptic | $0.2406 |
| advocate | $0.2151 |
| reviewer contribution context | $0.1549 |
| debate synthesizer | $0.0760 |
| audit citation integrity | $0.0484 |
| desk screen | $0.0303 |
| audit methods completeness | $0.0190 |
| journal recommender | $0.0144 |
| reviewer reporting reproducibility | $0.0132 |
| reviewer scientific validity | $0.0127 |
| reviewer data analysis | $0.0126 |
| reviewer ethics | $0.0042 |
Cite this review
Permanent: this review only
This URL is a permanent link to this specific review, and will not change.
In Silico (2026). Review of "On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules". In Silico. https://pgarrett-scripps.github.io/insilico/reviews/2026/on-the-sensitivity-of-machine-learned-2607-19161/v2/
@misc{insilico-on-the-sensitivity-of-machine-learned-2607-19161-v2,
title = {Review of {On the sensitivity of machine-learned probabilistic weather forecast models to scale-aware scoring rules}},
author = {{In Silico}},
year = {2026},
howpublished = {In Silico, an AI-refereed overlay journal},
url = {https://pgarrett-scripps.github.io/insilico/reviews/2026/on-the-sensitivity-of-machine-learned-2607-19161/v2/},
note = {Machine-generated peer review of doi:10.48550/arXiv.2607.19161 v2. Produced by PeerReviewAgents 0.5.0. Produced by PeerReviewAgents, doi:10.5281/zenodo.21781895.}
}Please cite the preprint itself as well. This reviews that work, it does not replace it. The review is machine-generated and advisory. If you are citing it as evidence about the paper, say so explicitly.