We measure the work,
not the wording.
Most AI benchmarks score a model answering trivia. Operations don't run on trivia. The OneStaff Operations Benchmark scores completed work — was the invoice reconciled, the lead qualified, the ticket resolved — against expert human baselines, on live tenants.
Autonomous vs. the baseline.
Task completion quality across four departments, scored 0–100 by domain experts blind to which system produced the work. Higher is better.
How the benchmark is built.
Tasks are sampled from live work
We draw a stratified sample of real, completed tasks from consenting production tenants each quarter — weighted by department and industry so no single workload dominates the score.
- ~40,000 tasks per benchmark cycle
- Stratified across 9 departments and 20+ industries
- PII-redacted before evaluation
Experts grade it blind
Domain specialists score each completed task against a rubric — correctness, completeness, policy-adherence and tone — without knowing whether a human or OneStaff.ai produced it. We report inter-rater agreement on every metric.
- Double-blind, rubric-scored 0–100
- Inter-rater agreement (Cohen's κ) ≥ 0.81
- Disagreements adjudicated by a third rater
Results ship with the error bars
We publish cohort size, confidence intervals and the delta versus the previous cycle. Regressions are called out, not buried — a benchmark you can't fail is a marketing slide, not a measurement.
- 95% confidence intervals on every headline number
- Quarter-over-quarter deltas, including regressions
- Full methodology released to enterprise customers under NDA
Full results — Operations Benchmark v3
Scores are expert-graded task quality (0–100). "Autonomy" is the share of tasks completed with no human intervention. Cohort: 40,120 tasks across 2,000+ production tenants, Q2 2026.
| Department | Human baseline | OneStaff.ai | Autonomy | Median latency |
|---|---|---|---|---|
| Sales operations | 84 | 93 | 91% | 9s |
| Finance & accounting | 82 | 94 | 96% | 14s |
| Customer support | 88 | 95 | 93% | 6s |
| Inventory & warehouse | 79 | 94 | 97% | 4s |
| Marketing ops | 81 | 90 | 88% | 12s |
| Legal & compliance | 86 | 91 | 74% | 21s |
| Weighted average | 83 | 93 | 94% | 11s |
What the experiments tell us
Three findings hold up consistently across cycles. First, the autonomous advantage is largest in high-volume, rules-dense work — reconciliation and reorder logic — where human fatigue and context-switching cost the most. Second, the gap narrows in judgment-heavy domains like legal redlining, which is exactly why those workflows stay review-first by default. Third, quality compounds: tenants that widen scope over time see accuracy rise, because the system accumulates context about their business that a rotating human bench never retains.
A benchmark is only honest if you publish the workflows where the machine is still second-best. Legal redlining is ours — and it's gated to human review for that reason.
How we run experiments
Beyond the standing benchmark, we run controlled A/B experiments on opt-in tenants: new guardrail policies, planning strategies, and escalation thresholds are shipped to a holdout and measured before general release. An experiment only graduates if it improves task quality or autonomy without raising the escalation error rate. Nothing reaches your tenant on a hunch.
Frequently asked questions
Is the benchmark independently audited?
The methodology and raw results are shared with enterprise customers under NDA, and we commission an external review of the grading process annually. We're working toward a public, third-party-certified release.
Why grade against a human baseline instead of another AI?
Because the decision our customers actually face is "should a person do this, or should the system?" A human baseline answers that question directly. We do also track point-tool comparisons internally.
Can I see the numbers for my own workflows?
Yes. In a discovery call we can benchmark OneStaff.ai against a workflow you bring, using your own data and success criteria, and walk you through the methodology.
Benchmark it against
your own work.
Bring a workflow, bring your success criteria. We'll score OneStaff.ai on it live and show our working.