An operational and human-reference benchmark of the xPeer engine — combining operational behaviour, a public same-manuscript human-reference resource, explicit cohort accounting, concern-level diagnostics, recommendation correspondence, and reproducibility controls. The evaluated engine is xPeer; its web front end is xPeerd.com.
The benchmark draws on a large-scale corpus spanning operational simulations, public manuscript records, strict paired analyses, and granular concern-unit extraction.
Retained from 500 original records; 70.4% retention rate.
HTTP-success records persisted from Re3-Sci2.0 F1000Research source.
Two human reports and two usable xPeer reviewer reports on the same v1 manuscript.
Across 1,023 of 1,084 reports — 94.4% report-unit coverage.
Longer reports, more extracted concerns, more explicit manuscript targets, broader category representation, and more explicit revision actions across all disciplinary groups.
More explicit rationale cues, stronger lexical manuscript attestation, slightly higher taxonomy-based scientific relevance, and lower mean redundancy across the strict paired cohort.
The benchmark is structured around two complementary components: an operational arm characterising disciplinary breadth and task-conditioned workload, and a human-reference arm enabling direct source-level comparison.
352 / 500 stable-task simulation reports used to characterise disciplinary breadth, task-conditioned workload, simulated decisions, and page anchoring across five subject groups.
271 paired manuscripts — each with two human reports and two usable xPeer reviewer reports on the same version-1 manuscript. xPeer outputs were persisted before source joining.
The operational component retained 70.4% of the original 500 records. Health Sciences and Physical Sciences were the largest subject groups, each exceeding the declared 0.20 classification-confidence threshold.
Every retained report exceeded the prespecified 0.20 classification-confidence threshold. Multidisciplinary records yielded zero stable-task reports in the operational corpus.
Revision formed the majority of simulated outcomes in every disciplinary group. Rejection reached ~45% in Health Sciences and ~42% in Life Sciences. Acceptance was rare across all fields.
Overall mean issue-level page-anchor fraction: 0.29. Report length showed a weak positive association with anchor fraction (Spearman ρ = 0.13, p = 0.014).
The resource was built from Re3-Sci2.0 F1000Research records restricted to manuscript version 1 with at least two linked human reports. Complete-pair availability was 33.8% of exact-two-human records.
Each stage applied a strict inclusion criterion. The final strict cohort of 271 used zero imputation, reconstruction, or reassignment of reviewer fields.
Total incomplete exact-two-human records: 531. Raw transport responses were absent for all incomplete states.
The strict cohort contains 271 manuscripts, 542 human reports, and 542 xPeer reports. Concern extraction produced 15,563 units across 1,023 of 1,084 reports — 94.4% coverage.
xPeer reports were 2.5× longer (median 1,889 vs. 763 words; rank-biserial 0.836) and contained 3.2× more concern units (median 41 vs. 13; rank-biserial 0.904). Mean concerns per 1,000 words were modestly higher for xPeer (21.597 vs. 18.981).
TRACE-R comprises Targeting, explicit Reasoning, Attested alignment, Category coverage, Executability, and Scientific Relevance — each reported as an observable text property on a 0–1 scale.
xPeer led on targeting (+0.135), category coverage (+0.222), and executability (+0.063). Humans led on explicit reasoning (+0.063), attested alignment (+0.023), and scientific relevance (+0.046).
xPeer showed higher prevalence in statistics, study design, methods and reproducibility, data and results, interpretation and claims, literature context, ethics and reporting, and presentation and clarity. All differences remained significant after FDR correction.
Concern units were assigned one-to-one and accepted above a prespecified lexical-similarity threshold of 0.35. Median matched fraction was zero in both source-normalised views. Only 56 manuscripts yielded at least one accepted pair.
Fraction of human concern units matched by xPeer above threshold.
Fraction of xPeer concern units matched to human units above threshold.
Of 271 strict paired manuscripts in the cohort.
Human recommendation metadata were available for all 542 human reports. Normalised recommendation language was extracted from 380 of 542 xPeer reports (70.1% report-level coverage). At manuscript level, 240 cases had usable source consensus values.
The supported inference is source non-equivalence. Editorial decision use remains under accountable human authority.
The benchmark separates design-level abstractions from directly observed implementation variables. The five-phase pipeline structures how manuscript inputs are transformed into reviewer outputs.
Each phase operates on explicit structured representations of the manuscript and produces intermediates that feed the subsequent stage, enabling traceable concern-level diagnostics.
Checks covered cohort counts, report balance, non-empty report text, concern-unit coverage, metric bounds, paired-row counts, sensitivity thresholds, category reconciliation, recommendation auditing, and output existence.
All prespecified quality checks cleared without exception.
Used for confidence interval estimation across all paired metrics.
Used for null-distribution estimation in significance testing.
Use xPeer as a pre-submission stress test for methodological detail, reporting omissions, unsupported interpretation, statistical issues, and presentation barriers.
Use xPeer as a standardised methodological and reporting sweep before or alongside human review, while retaining accountable human publication authority.
Use the released resource as a common test bed for aligned future evaluation with expert adjudication, cost, latency, and governance measures.
The strict paired cohort contains 271 of 1,108 released records and 271 of 802 exact-two-human records, creating complete-case selection risk.
531 incomplete exact-two-human records include blank, recommendation-only and partial-reviewer states; raw transport responses are absent.
Manuscripts and human reports derive from F1000Research-linked data; transfer to anonymous pre-publication review and other venue types requires external validation.
Human reports are manuscript-linked references and do not constitute ground truth for scientific correctness.
Concern extraction uses deterministic lexical and structural rules that may interact with source style. No blinded expert-annotation subset was available for precision/recall verification.
61 reports yielded zero extracted concern units: 22 human and 39 xPeer. Attested alignment quantifies lexical attestation, not factual correctness or domain-grounded reasoning.
Cross-source concern correspondence is threshold-sensitive (0.25–0.50 sensitivity range) and requires expert adjudication for utility claims. System recommendation labels were observable in only 70.1% of xPeer reports.
KNOWDYN produces the system and the author declares a controlling interest. Public data, code, exclusions, hashes, independent replication, and external adjudication form the conflict-management framework.
All datasets, code, and reproducibility materials are publicly archived. The version-pinned archive includes a frozen repository commit, SHA-256 hash, and fixed random seed to enable full independent replication.
Copyright. © 2026 Khalid M. Saqr and KNOWDYN LTD, as applicable. All rights reserved except where a separate licence is expressly stated.
Disclaimer. This publication is provided for research and informational purposes only. It does not constitute legal, editorial, scientific or professional advice; users remain responsible for independent evaluation and decisions.
Intellectual property. xPeer and xPeerd (.com, .online), including their software, methods, prompts, workflows, interfaces and confidential technical information, are proprietary to KNOWDYN LTD.

Copyright © KNOWDYN. All rights reserved.