mcp-proto-okn: Natural-language access to open scientific knowledge graphs through the Model Context Protocol

Decision letter

Minor revisionpanel verdict · 2026-09-01

Decision Letter

VERDICT: minor

Summary of Evaluation

This manuscript describes mcp-proto-okn, a FastMCP-based Model Context Protocol server that exposes the Proto-OKN family of knowledge graphs (hosted on the OKN Fabric) to LLM assistants, with tools for graph discovery/routing, schema inspection, SPARQL execution, ontology expansion via UberGraph, multi-graph querying, and transcript generation. Two case studies illustrate use: a multi-step spaceflight transcriptomics workflow (NASA GeneLab OSD-244, spoke-genelab → spoke-okn) and an ontology-expansion query against the NIAID Data Ecosystem.

All five specialist reports converged on the same assessment (four scores of 4/5; ethics N/A with no issues found). The panel agrees that the tool addresses a real access barrier, that the implementation is described coherently, that the code and verbatim chat transcripts are publicly available, and that the limitations section is candid. No reviewer identified evidence contradicting the reported behaviour of the system, and no reviewer proposed rejection.

Two substantive issues dominated the debate, and both are, on my reading, claim-scoping and reporting problems rather than design failures:

  1. Evidence-to-claim mismatch on "enablement." The abstract and conclusions assert that the server "enables" natural-language discovery and querying and "lowers the barrier for integrative discovery." The only evidence is two self-selected, successful case studies. Nothing in the manuscript establishes success rate, first-attempt query validity, or failure modes across the 30+ graphs, and the manuscript does not state whether the transcripts are representative or whether failed attempts were discarded. I agree with the skeptic that two success stories cannot support a general reliability claim. I do not agree that this requires a new benchmark before publication: the manuscript is explicit that the cases are "illustrative," and the honest fix — scaling the claim language to feasibility/proof-of-concept, and stating plainly what was and was not attempted — is achievable in text alone. A quantitative evaluation would substantially strengthen a future paper and I encourage it, but I am not making it a condition of acceptance for an applications-note-style contribution at this venue.

  2. The ontology-expansion comparison is confounded as framed. The 447 → 10,000+ contrast is, by ontology construction, largely guaranteed: records annotated only with descendant terms are necessarily invisible to a root-only filter. The number that carries information is the count of unique datasets recovered that the root-only query missed, and the precision cost of admitting 1,592 descendant URIs (whether a user asking about "cardiovascular disease" wants Brugada or Holt-Oram datasets is not obvious). The current wording — "preserving recall across the full ontological subtree" — implies a validated recall/precision property that was not measured. This is fixable either by reporting the overlap and unique-dataset counts (a single re-run of an already-implemented query against the same graph, not a new experiment) or by narrowing the claim to "increases recall relative to root-term-only filtering."

The compliance audits raise a large number of methods-traceability gaps. I have triaged these against what the manuscript actually claims. Items that genuinely block verification of the case studies — a persistent accession/citation for OSD-244, whether the differential-expression results were retrieved from spoke-genelab or recomputed, the provenance and definition of the r = 0.80 statistic, the ortholog-mapping source, the LLM model/version used, and version pinning for the tool and transcripts — are folded into required revisions. Items that belong to the upstream NASA study rather than to this reanalysis (sequencing platform, library kit, read depth, reference build, IACUC protocol, randomization/blinding) are not this manuscript's obligation to reproduce; they are discharged by citing the source study and repository record, and I have framed the requirement that way. The reviewers' and auditors' framing of these as HARD gaps in this paper overreaches: this is a query-interface paper reusing a public repository record, not a spaceflight biology paper. Likewise, the future-dated references flagged by the citation audit are consistent with a manuscript prepared in 2026 and are a verification request, not a defect.

Two further points from the reports deserve author attention and are reflected below. First, the division of labour between the MCP server and the LLM client is genuinely unclear in the text (who performs ortholog mapping, result merging, figure generation, batch reassembly?) — this matters because the paper's contribution claim rests on it. Second, the positioning against TogoMCP asserts differences (routing, hierarchical reasoning, provenance) without demonstrating them; the contribution reviewer is right that scope-versus-capability should be stated candidly rather than argued.

None of this requires new experiments or a reanalysis whose outcome could overturn a conclusion. The verdict is minor revision, with a substantial list.

Required Revisions

  1. Scale the enablement claims to the evidence. Revise the abstract, the introduction's contribution paragraph, and the conclusions so that "enables," "lowers the barrier," and similar phrasing is explicitly framed as demonstrated feasibility on two worked examples rather than validated reliability. Add one or two sentences stating (a) whether the two transcripts are representative of typical sessions or selected successes, (b) whether the reported workflows succeeded on first attempt or required reformulation/retries, and (c) that success rate, routing accuracy and failure modes across the full graph set have not yet been quantified.

  2. Fix the ontology-expansion claim. Either report the count of unique datasets in the expanded result, the count in the root-only result, and their overlap (i.e. how many datasets the naive query actually missed); or narrow the claim to "increases recall relative to root-term-only filtering." In either case, remove or qualify "preserving recall," which asserts an unmeasured property. Also state explicitly whether ">10,000" counts unique dataset records or dataset–disease pairs, and clarify the 284-condition tabulation accordingly. Acknowledge the precision/intent trade-off of automatic subtree expansion in one or two sentences.

  3. Specify the division of labour between server and client. For each step of Case Study 1, state which operations were performed by mcp-proto-okn tools (with the tool name) and which by the LLM assistant's own reasoning or external services: specifically the mouse→human ortholog mapping, the identification of the 322 concordant genes, result merging across the 80 NDE batches, and the generation of assay-design diagrams and concordance plots. A short table or numbered tool-call sequence in Methods or supplementary text would satisfy this.

  4. Document the provenance of the Case Study 1 statistics. State whether the differential-expression results were retrieved as precomputed values from spoke-genelab or recomputed within the session. If precomputed, cite the pipeline and thresholds as recorded by the source repository. Define the r = 0.80 statistic precisely: what quantities were correlated (e.g. log fold changes), across which n (state n = 322 if so), and whether any significance test was applied. If no test was applied, say so and describe the correlation as descriptive; do not describe it as evidence of a "sustained response" without stating the basis. Also state the significance threshold and multiple-testing correction behind "thousands of significant genes," and name the source of the ortholog mapping (e.g. Ensembl, HomoloGene) with the release used.

  5. Provide a data and materials availability statement with persistent identifiers. Include: the NASA OSDR/GeneLab accession for OSD-244 with a resolvable URL and, if available, a citation to the study's primary publication (to which wet-lab and animal-care methods are properly delegated); the query date for the NDE and UberGraph queries; and a deposited copy of the query result tables underlying the Case Study 2 tabulation. State explicitly that the Proto-OKN graphs are live services and that results are time-stamped rather than frozen.

  6. Pin versions. In Methods, report the Python version, the FastMCP version, the mcp-proto-okn release tag or commit hash used for the case studies, and a pointer to the dependency specification in the repository. Report the LLM model and version used in each case study (e.g. the specific Claude or GPT model), and note that LLM outputs are non-deterministic so transcripts represent single sessions.

  7. Archive the transcripts and repository snapshot at a persistent, non-GitHub-dependent location. Deposit the two verbatim transcripts and a code snapshot with a DOI (e.g. Zenodo, Software Heritage) and cite it. The transcripts are the paper's central evidentiary artifact; a main-branch blob URL is mutable and may not resolve for future readers.

  8. Sharpen the positioning against TogoMCP and Emonet et al. Replace the asserted distinction with a candid statement of what is shared (schema-guided LLM SPARQL generation, MCP tool orchestration, ontology-based term expansion — all established) and what is specific here (application to a non-curated, cross-domain graph collection with no shared schema convention; graph routing across 30+ graphs; provenance annotation of multi-graph results). If no head-to-head comparison was performed, say so and say why. Also state whether any query validation step (schema checking, dry-run, syntax validation) is performed before execution, and if not, how hallucinated SPARQL is detected — the Emonet et al. framework's validation step invites this comparison directly.

  9. Document the ontology-expansion mechanics. State the batch size used (80 batches of what size), the default and configurable expansion bounds, the trigger logic (always expand when a recognised ontology identifier is present, or conditionally), whether descendants are cached or fetched per query, and whether any batches failed or returned partial results in Case Study 2.

  10. Substantiate or soften the "coordinated multi-graph querying" and get_join_strategy claims. Neither case study exercises multi_graph_query or get_join_strategy; Case Study 1 queries two graphs sequentially. Either add a brief worked example (a transcript excerpt suffices) or state clearly in the text that these tools are implemented but not exercised in the presented use cases, and define what "coordinated" means operationally (sequential, parallel, or join across results).

  11. Verify the reference list. Confirm the bioRxiv identifier and status for Kinjo et al. (the given identifier, 2026.03.19.713030, could not be resolved by the citation audit) and confirm the years for Madrigal et al. 2026, Tsueng et al. 2026 and Gebre et al. 2025 as advance/final publications. Check that the Gao et al. citation supports the specific characterisation "BioBricks toxicology graphs," or reword to match what that reference actually describes.

Minor Suggestions

  • A held-out evaluation — say 30–50 natural-language questions spanning the graph set, with reported routing accuracy, first-attempt SPARQL validity, and failure taxonomy — would convert the paper's central claim from demonstrated to measured. I have not made this a condition of acceptance, but it is the single highest-value addition available and would be the natural core of a follow-up paper.
  • Report indicative wall-clock times for representative queries, including the 80-batch expansion, and any endpoint timeouts or downtime encountered during development. Practical latency matters to users deciding whether to adopt the tool.
  • Give one concrete example of the "query-analysis metadata and warnings about common issues" returned by the query tool; the feature is claimed but never illustrated.
  • Describe how the assistant handles ambiguous routing (e.g. "genes" matching both spoke-okn and spoke-genelab): clarification prompt, both graphs, or heuristic choice.
  • Consider a short paragraph on how the system behaves when two graphs use the same predicate with divergent semantics, since this is the failure mode most specific to the uncurated setting the paper claims as its distinguishing context.
  • State the repository licence in the manuscript's availability section.
  • In Case Study 1, report how many mouse genes mapped to human orthologs, how many were ambiguous or unmapped, and whether ambiguity was resolved or propagated.

I look forward to the revised version. The core contribution is useful and the openness of the code and transcripts is a genuine strength; the revisions above are about making the claims match what was shown and making the two demonstrations independently checkable.

← All documents in this review