dnoise: Fast Native Data Reduction for Bruker timsTOF

Decision letter

Minor revisionpanel verdict · 2026-09-02

Decision Letter

VERDICT: minor

Summary of Evaluation

This manuscript describes dnoise, a Rust tool that removes points from Bruker timsTOF frames and writes the result back as a native-compatible .d directory, and benchmarks the default MS1-only configuration on the public Generation Beta three-species standard across ddaPASEF and diaPASEF at 5- and 15-minute gradients.

The panel was unanimous in its assessment: five specialist reports, all scoring 4/5 (ethics 5/5), and the advocate–skeptic debate produced no issue that either side considered fatal — the skeptic explicitly disclaimed arguing for rejection. My own reading agrees. The central claim is modest, clearly stated, and supported by the evidence actually presented: default MS1-only denoising reduced the frame binary by 35–53%, ddaPASEF search results were byte-for-byte unchanged (as they must be, since MS/MS is untouched — and the authors say so rather than presenting it as a result), diaPASEF identification counts moved by ≤2.2%, and label-free accuracy, precision, and feature-level intensity agreement were retained. The controls are the right ones: a per-stage ablation, a removal-matched intensity-threshold comparator, a byte-identical round-trip check, and a decoy-to-target loss analysis that gives a mechanistic account of why the optional MS/MS mode costs identifications. Data, code, version tags, and a Zenodo archive are all present. The limitations section is genuinely candid, including the disclosure that the DDA/DIA reduction difference is confounded by on-instrument denoising and that the 15-minute ddaPASEF arm is not fully out-of-sample.

The issues the panel raised are real but, on inspection, all resolvable from data the authors already hold or from the text itself. Three deserve comment.

First, the parameter-selection circularity (flagged as a HARD reproducibility issue by one referee and escalated by the skeptic in the debate). The concern is legitimate — these defaults propagate into every reported condition — but the sweep was performed and is referenced as Table S2; what is missing is its full disclosure and an explicit decision rule. That is a reporting gap, not a missing experiment. It is further mitigated by the fact that three of four benchmark arms (5-minute ddaPASEF and both diaPASEF acquisitions) played no role in selection and show the same behavior.

Second, the "native-compatible" claim, which the debate identified as a collective blind spot: it is validated by dnoise's own reader plus the two pipelines used for the benchmark. I agree the framing in the Introduction (which lists MaxQuant, AlphaTims, OpenTIMS, i2MassChroQ, rustims) implies more than was tested. I am requiring this be closed, but either by one additional reader test or by scoping the claim — so it does not force new work.

Third, the diaPASEF MS1-area ratios moving toward expectation after removing supposedly uninformative points. This is an unexplained observation, not a contradiction of any accuracy claim, and I am asking for candid discussion rather than resolution.

Several other items — quantified-protein-set overlap, the composition of the lost diaPASEF precursors, the halo filter's quantitative contribution — are tabulations or reanalyses of search outputs the authors already have, and their plausible outcomes would refine rather than overturn the conclusions. Given the venue's standard, where reports are published alongside the preprint and the paper's value rests on a reader being able to judge exactly how far the claim extends, the correct outcome is a minor revision with a thorough list.

The contribution is incremental and the authors do not claim otherwise. That is not a criticism here. A well-scoped, honestly reported, immediately usable tool with a checkable benchmark is exactly the kind of work this venue should list.

Required Revisions

  1. Publish the complete parameter sweep and the decision rule. Provide Table S2 in full — every tested (min_feature_length, max_internal_gap) pair with its quantified coverage, precision, and intensity-fidelity outcomes — and state explicitly the criterion by which gap 2 / length 5 was chosen over the alternatives. "Prioritize stricter local continuity and greater point removal" is not a checkable rule. Also state, from the sweep data you already have, how much the headline reduction and quantification outcomes would have differed under the nearest neighbouring settings, so readers can judge sensitivity to this choice.

  2. Report the bootstrap confidence intervals in the main text and reconcile "unchanged" with "moved toward expectation." Tables S10/S11 are cited but the intervals are not surfaced where the accuracy claim is made. Bring at least the diaPASEF condition pairs with visible shifts in Figure 3 into the main text with their 95% CIs, and define operationally what "LFQ accuracy was preserved" means — statistically indistinguishable from the original, or within a stated tolerance. Section 3.2 currently says accuracy was "unchanged" in one sentence and that regulated-species ratios "moved toward their expected values" a few lines later; these need to be stated consistently.

  3. Scope the headline claims to what was tested. The Abstract, Introduction, and Conclusions present the 35–53% range and the preservation result without instrument, load, or sample qualification, while Section 3.7 correctly restricts them to one timsTOF Ultra 2, one laboratory, one 50 ng three-species digest, two gradients, two acquisition modes. Add that qualification where the claims are first made. In particular, note at or near the first side-by-side presentation of the ddaPASEF and diaPASEF reduction figures that on-instrument denoising was enabled only for the ddaPASEF survey scans, so the two numbers are not a mode-level comparison. The body already says this; the Abstract does not.

  4. Close the native-compatibility claim. Either (a) demonstrate that at least one independent native-.d reader that you did not use in the benchmark (e.g. AlphaTims, OpenTIMS, or MaxQuant) opens a denoised directory without error, or (b) restate the claim explicitly as "compatible with the timsrust round-trip and with the Sage and DIA-NN workflows used here," and remove any implication of ecosystem-wide drop-in compatibility. Option (b) is acceptable; what is not acceptable is a title and framing that promise "native" interoperability validated only against the tools used to generate the paper's own results.

  5. Report the quantified-protein and -peptide set overlap under default MS1-only denoising. Table S5 shows small count changes between the original and MS1-only arms. State the overlap of the quantified sets, how many features enter and exit, and whether the movement is attributable to crossing the two-peptide/two-replicate or LFQ q-value thresholds. You already give this explanation for the MS1+MS/MS arm in Section 3.3; extend the same accounting to the default mode so that "preserved" is unambiguous.

  6. Characterize the diaPASEF precursors lost to MS1-only denoising. For the 0.2–2.2% of precursors and 0.2–1.6% of protein groups that change, report their species composition and abundance rank relative to the retained set. If the loss is concentrated in one species or at the low-abundance end, that materially conditions the accuracy claim and should be stated; if it is unstructured, saying so strengthens the claim. This is a tabulation of existing DIA-NN reports.

  7. Discuss, candidly, why direct precursor MS1-area ratios improve after denoising (Table S12). As written, this is in tension with the framing that the removed points carry no analytical signal. Offer the candidate explanations (e.g. removal of weak co-eluting background biasing area integration, or the asymmetry created by on-instrument denoising being off for diaPASEF), and state plainly that the mechanism is not established by these data. Do not present a favorable shift as confirmation.

  8. Clarify the halo filter's default status and quantify its analytical effect. Section 2.1 states the filter "can be disabled," Figure 1 shows it as the third of three default stages, and at least one referee read it as off by default — the text is ambiguous. State the default unambiguously. Then, since Table S4 shows only its point-count contribution ("a small final trim"), add its effect on quantified coverage, accuracy, and precision, and give the basis for halo_peak_fraction = 0.15. If the effect is small, this is inexpensive to show.

  9. State and discuss the calibration of the matched intensity-threshold control. Make explicit in Section 3.4 that the cutoff was calibrated per acquisition rather than globally, and acknowledge the limitation the panel identified: the comparison establishes that mobility-coherent filtering beats a global per-point cutoff at matched removal, but does not isolate the benefit of mobility awareness from the benefit of not using a single global cutoff. Either soften the section heading and claim accordingly, or add a mobility-local threshold arm if you have the data.

  10. Complete the software and search reporting. Specify the rayon version (and ideally reference the pinned Cargo.lock in the archived release); give the numerical m/z and mobility gate-padding values used in the benchmark configuration explicitly in Table S1, since the graphical abstract notes the benchmark pads edges while the figure does not; and state the FDR thresholds applied at PSM, peptide, and protein level in the Methods rather than only implying them via the "1% LFQ q-value" in Results.

  11. Fix three citation items. (a) Reference 4 (Houthuijs et al. 2026) — confirm publication status and DOI. (b) Reference 19 — the DOI 10.64898/2026.01.29.702266 is unusually formed and appears to duplicate reference 18; verify or correct. (c) The statement that the PNNL PreProcessor "does not support Bruker .d" is load-bearing for your novelty claim — give the basis (a section reference in Bilbao et al., or the tool's documentation), and distinguish "does not currently support" from "cannot support."

  12. Give absolute numbers alongside relative ones. Figure 2's right panel reports counts relative to the original arm; add the absolute PSM/precursor, peptide, and protein-group counts to the figure or its caption so a reader can tell whether a 1.6% change is ten proteins or a hundred. Similarly, add a median or interquartile summary to the runtime range in Section 3.6 rather than only the 7.4–39.0 s and 10.2–68.7 s extremes.

Minor Suggestions

  • Add practical guidance on when the optional MS1+MS/MS mode is justified (e.g. long-term archival of runs whose identifications are already locked in, versus active discovery), and on how to relax the msms_* parameters for sparser fragment data. The current recommendation is correctly conservative but leaves the user without a decision framework.
  • Say something about behavior on low-input or sparse acquisitions even if untested — e.g. which direction min_feature_length should move, and whether any diagnostic output helps a user detect over-filtering. You correctly flag single-cell data as unvalidated; a sentence of practical advice would be more useful than a caveat alone.
  • Document edge-case behavior briefly: absent or malformed method tables (you state gates "silently do nothing" — a warning would be safer than silence), frames reduced to zero surviving points, and non-standard acquisition schemes.
  • Consider contextualizing the runtime against the alternative a user might otherwise run (mzML conversion, or a feature-finding pass), which would make the "fast enough for routine post-acquisition use" claim concrete rather than comparative-in-the-abstract.
  • The optional centroiders (Section 3.5) are reported crisply but sit outside the main benchmark; a one-line statement of their interaction with the default stages and of who should use which would help.
  • Consider tempering "fast enough for routine post-acquisition use" to the narrower and fully supported form: fast enough to run before transfer without delaying it.

We thank the authors for a clearly written, honestly bounded manuscript with genuinely open data and code. The revision requested is substantial in item count but requires no new acquisitions and, with the exception of the optional reader test in item 4, no new computation beyond tabulating results you already hold.

← All documents in this review