Home/Research/Benchmarks & Experiments
Benchmarks & Experiments

We measure the work,
not the wording.

Most AI benchmarks score a model answering trivia. Operations don't run on trivia. The OneStaff Operations Benchmark scores completed work — was the invoice reconciled, the lead qualified, the ticket resolved — against expert human baselines, on live tenants.

0.9%
task accuracy, expert-graded
0%
of workflows completed with no human touch
0s
median time-to-action per task
0.4%
escalation error rate
OneStaff Operations Benchmark v3

Autonomous vs. the baseline.

Task completion quality across four departments, scored 0–100 by domain experts blind to which system produced the work. Higher is better.

Human ops baseline OneStaff.ai
Sales — lead qualification & follow-up+9 pts
Finance — reconciliation & invoicing+12 pts
Support — ticket resolution+7 pts
Inventory — reorder accuracy+15 pts
0%
Task accuracy
0%
Fully autonomous
Methodology

How the benchmark is built.

01

Tasks are sampled from live work

We draw a stratified sample of real, completed tasks from consenting production tenants each quarter — weighted by department and industry so no single workload dominates the score.

  • ~40,000 tasks per benchmark cycle
  • Stratified across 9 departments and 20+ industries
  • PII-redacted before evaluation
Sample frame
40,120 completed tasks · Q2 2026
Stratification
By department, industry, complexity
Redaction
PII removed · tenant-anonymised
02

Experts grade it blind

Domain specialists score each completed task against a rubric — correctness, completeness, policy-adherence and tone — without knowing whether a human or OneStaff.ai produced it. We report inter-rater agreement on every metric.

  • Double-blind, rubric-scored 0–100
  • Inter-rater agreement (Cohen's κ) ≥ 0.81
  • Disagreements adjudicated by a third rater
Rater panel
28 domain specialists
Rubric
Correct · complete · compliant · tone
Agreement
κ = 0.83 this cycle
03

Results ship with the error bars

We publish cohort size, confidence intervals and the delta versus the previous cycle. Regressions are called out, not buried — a benchmark you can't fail is a marketing slide, not a measurement.

  • 95% confidence intervals on every headline number
  • Quarter-over-quarter deltas, including regressions
  • Full methodology released to enterprise customers under NDA
Accuracy
99.9% ±0.2 (n=40,120)
vs. v2
+1.4 pts autonomy · +0.3 accuracy
Watch item
Legal redlining flat QoQ

Benchmark it against
your own work.

Bring a workflow, bring your success criteria. We'll score OneStaff.ai on it live and show our working.