Open Benchmark Study · August 2026

xPeer Benchmark: Simulating Scholarly Peer Review

An operational and human-reference benchmark of the xPeer engine — combining operational behaviour, a public same-manuscript human-reference resource, explicit cohort accounting, concern-level diagnostics, recommendation correspondence, and reproducibility controls. The evaluated engine is xPeer; its web front end is xPeerd.com.

Scale at a Glance

xPeerd Comparison with Human Reviewers

The benchmark draws on a large-scale corpus spanning operational simulations, public manuscript records, strict paired analyses, and granular concern-unit extraction.

352

Valid operational simulation reports

Retained from 500 original records; 70.4% retention rate.

1,108

Public manuscript-level records

HTTP-success records persisted from Re3-Sci2.0 F1000Research source.

271

Strict paired manuscripts

Two human reports and two usable xPeer reviewer reports on the same v1 manuscript.

15,563

Extracted concern units

Across 1,023 of 1,084 reports — 94.4% report-unit coverage.

Executive Summary

What the Benchmark Establishes

Observable xPeer Profile

Longer reports, more extracted concerns, more explicit manuscript targets, broader category representation, and more explicit revision actions across all disciplinary groups.

Observable Human Profile

More explicit rationale cues, stronger lexical manuscript attestation, slightly higher taxonomy-based scientific relevance, and lower mean redundancy across the strict paired cohort.

Study Architecture

Two Evidence Layers, One Benchmark

The benchmark is structured around two complementary components: an operational arm characterising disciplinary breadth and task-conditioned workload, and a human-reference arm enabling direct source-level comparison.

Operational Component

352 / 500 stable-task simulation reports used to characterise disciplinary breadth, task-conditioned workload, simulated decisions, and page anchoring across five subject groups.

Human-Reference Component

271 paired manuscripts — each with two human reports and two usable xPeer reviewer reports on the same version-1 manuscript. xPeer outputs were persisted before source joining.

Operational Benchmark

Disciplinary Coverage: 352 Valid Simulations

The operational component retained 70.4% of the original 500 records. Health Sciences and Physical Sciences were the largest subject groups, each exceeding the declared 0.20 classification-confidence threshold.

Every retained report exceeded the prespecified 0.20 classification-confidence threshold. Multidisciplinary records yielded zero stable-task reports in the operational corpus.

Operational Benchmark

Simulated Decisions and Page Anchoring

Simulated Decisions by Outcome

Revision formed the majority of simulated outcomes in every disciplinary group. Rejection reached ~45% in Health Sciences and ~42% in Life Sciences. Acceptance was rare across all fields.

Mean Page-Anchor Fraction by Field

Overall mean issue-level page-anchor fraction: 0.29. Report length showed a weak positive association with anchor fraction (Spearman ρ = 0.13, p = 0.014).

Human-Reference Resource

Cohort Funnel: From Source Corpus to Strict Paired Analysis

The resource was built from Re3-Sci2.0 F1000Research records restricted to manuscript version 1 with at least two linked human reports. Complete-pair availability was 33.8% of exact-two-human records.

Each stage applied a strict inclusion criterion. The final strict cohort of 271 used zero imputation, reconstruction, or reassignment of reviewer fields.

Incomplete xPeer Reviewer-Field States

Total incomplete exact-two-human records: 531. Raw transport responses were absent for all incomplete states.

Paired Human-Reference Analysis

Report Scale: xPeer Reports Were Substantially Longer

The strict cohort contains 271 manuscripts, 542 human reports, and 542 xPeer reports. Concern extraction produced 15,563 units across 1,023 of 1,084 reports — 94.4% coverage.

xPeer reports were 2.5× longer (median 1,889 vs. 763 words; rank-biserial 0.836) and contained 3.2× more concern units (median 41 vs. 13; rank-biserial 0.904). Mean concerns per 1,000 words were modestly higher for xPeer (21.597 vs. 18.981).

Paired Human-Reference Analysis

TRACE-R Source Profile: Six Observable Dimensions

TRACE-R comprises Targeting, explicit Reasoning, Attested alignment, Category coverage, Executability, and Scientific Relevance — each reported as an observable text property on a 0–1 scale.

xPeer led on targeting (+0.135), category coverage (+0.222), and executability (+0.063). Humans led on explicit reasoning (+0.063), attested alignment (+0.023), and scientific relevance (+0.046).

Scientific-Category Prevalence

Broader Detector-Recognised Coverage Across the Taxonomy

xPeer showed higher prevalence in statistics, study design, methods and reproducibility, data and results, interpretation and claims, literature context, ethics and reporting, and presentation and clarity. All differences remained significant after FDR correction.

Source Correspondence

Low Lexical Matching and Low Recommendation Agreement

Concern units were assigned one-to-one and accepted above a prespecified lexical-similarity threshold of 0.35. Median matched fraction was zero in both source-normalised views. Only 56 manuscripts yielded at least one accepted pair.

0.026

Mean human recovery

Fraction of human concern units matched by xPeer above threshold.

0.009

Mean xPeer alignment

Fraction of xPeer concern units matched to human units above threshold.

56

Manuscripts with ≥1 accepted pair

Of 271 strict paired manuscripts in the cohort.

Recommendation Correspondence

Human recommendation metadata were available for all 542 human reports. Normalised recommendation language was extracted from 380 of 542 xPeer reports (70.1% report-level coverage). At manuscript level, 240 cases had usable source consensus values.

Agreement Metrics

Confusion Matrix (n = 240)

The supported inference is source non-equivalence. Editorial decision use remains under accountable human authority.

System Process

Design-Level Process Mapped to Measured Outputs

The benchmark separates design-level abstractions from directly observed implementation variables. The five-phase pipeline structures how manuscript inputs are transformed into reviewer outputs.

Each phase operates on explicit structured representations of the manuscript and produces intermediates that feed the subsequent stage, enabling traceable concern-level diagnostics.

Quality Control & Reproducibility

22 of 22 Prespecified Computational Checks Passed

Checks covered cohort counts, report balance, non-empty report text, concern-unit coverage, metric bounds, paired-row counts, sensitivity thresholds, category reconciliation, recommendation auditing, and output existence.

22/22

Computational checks passed

All prespecified quality checks cleared without exception.

2,000

Bootstrap replicates

Used for confidence interval estimation across all paired metrics.

1,999

Permutation replicates

Used for null-distribution estimation in significance testing.

Reproducibility Parameters

Strategic Implications

Use xPeer as an Additional Scrutiny Layer, Not an Autonomous Editorial Authority

Researchers & Authors

Use xPeer as a pre-submission stress test for methodological detail, reporting omissions, unsupported interpretation, statistical issues, and presentation barriers.

Editors & Publishers

Use xPeer as a standardised methodological and reporting sweep before or alongside human review, while retaining accountable human publication authority.

Benchmark Designers

Use the released resource as a common test bed for aligned future evaluation with expert adjudication, cost, latency, and governance measures.

Limitations

Twelve Principal Inference Boundaries

1

Cohort selection risk

The strict paired cohort contains 271 of 1,108 released records and 271 of 802 exact-two-human records, creating complete-case selection risk.

2

Incomplete record states

531 incomplete exact-two-human records include blank, recommendation-only and partial-reviewer states; raw transport responses are absent.

3

Venue transferability

Manuscripts and human reports derive from F1000Research-linked data; transfer to anonymous pre-publication review and other venue types requires external validation.

4

Human reports are references, not ground truth

Human reports are manuscript-linked references and do not constitute ground truth for scientific correctness.

1

Extraction rule interactions

Concern extraction uses deterministic lexical and structural rules that may interact with source style. No blinded expert-annotation subset was available for precision/recall verification.

2

Zero-unit reports

61 reports yielded zero extracted concern units: 22 human and 39 xPeer. Attested alignment quantifies lexical attestation, not factual correctness or domain-grounded reasoning.

3

Correspondence threshold sensitivity

Cross-source concern correspondence is threshold-sensitive (0.25–0.50 sensitivity range) and requires expert adjudication for utility claims. System recommendation labels were observable in only 70.1% of xPeer reports.

4

Conflict of interest

KNOWDYN produces the system and the author declares a controlling interest. Public data, code, exclusions, hashes, independent replication, and external adjudication form the conflict-management framework.

Data & Code Availability

Public Benchmark Resources

All datasets, code, and reproducibility materials are publicly archived. The version-pinned archive includes a frozen repository commit, SHA-256 hash, and fixed random seed to enable full independent replication.

Evaluation Repository


Copyright. © 2026 Khalid M. Saqr and KNOWDYN LTD, as applicable. All rights reserved except where a separate licence is expressly stated.

Disclaimer. This publication is provided for research and informational purposes only. It does not constitute legal, editorial, scientific or professional advice; users remain responsible for independent evaluation and decisions.

Intellectual property. xPeer and xPeerd (.com, .online), including their software, methods, prompts, workflows, interfaces and confidential technical information, are proprietary to KNOWDYN LTD.

Copyright © KNOWDYN. All rights reserved.