Copilot Analytics Lab · Inside the Metrics · Part 2

Measuring the Additional Work AI Agents Make Possible

New Assisted Hours (NAH) estimates the human-equivalent work an AI agent performs alongside a person — additional work done in parallel, not time subtracted from a task. This brief walks through the layered architecture, the six research stages behind it, and the calculations that make it auditable.

When an AI agent runs a multi-step workflow — searching, reading, drafting, calling tools, revising, verifying — it performs work a person would otherwise have to do themselves. New Assisted Hours quantifies that work by asking a precise question: if a person had produced the same output, step by step, how long would it have taken them? That human-equivalent time, aggregated across the workflow, is what New Assisted Hours counts.

Crucially, this is an additive view of time, not a subtractive one. The agent works in parallel with the person, so its output is best understood as human-equivalent hours added to what the person could do alone, rather than minutes shaved off a task they were already doing. This is why assisted hours can grow even when a task's elapsed time barely moves when completed by an agent — a pattern we observed directly in controlled trials, where AI assistance lowered subjective workload with no change in completion time. Time saved — the Agent Assisted Hours metric covered in Part 1 — and New Assisted Hours are complementary time-based views: one measures how much faster a shared task finishes, the other measures how much additional human-equivalent work the agent absorbed.

Summary

  • What it measures. The human-equivalent time a typical worker would need to reproduce the agent's steps — a proxy for the human-augmented hours the agent added, computed per prompt and aggregated upward.
  • Architecture. A five-layer pipeline. Layer 1 reads the agent's steps from telemetry; Layer 2 assigns each step a calibrated human-equivalent time; Layer 3 aggregates with an effort discount. Layers 4–5 (complexity/user adjustment, net capacity) are future work.
  • Evidence base. Six linked research stages: two randomized controlled trials, a whole-task PERT estimation study, a microtask decomposition and time estimation stidy, an LLM benchmark with calibration, and an effort-discounting aggregation.
  • Calibration. Raw LLM time estimates are strongly biased; prediction-bin calibration reduces mean absolute error from ~11.7 min to ~1.99 min under leave-one-task-out cross-validation.
  • Aggregation. A concave, order-aware discount — weight k^α − (k−1)^α, default α = 0.7 — prevents repeated steps from being overcounted, compressing a repetition-heavy worked example by ~58%.

This brief is the accessible overview. For the complete methodology, statistics, and derivations, see the detailed research brief.

Architecture

A layered pipeline

New Assisted Hours is organized as a pipeline whose layers answer different measurement questions and feed forward. The agent's raw activity enters at Layer 1; a defensible assisted-hours figure leaves at Layer 3. The two layers beyond that are named but out of scope for this brief — they are where complexity adjustment and net-capacity accounting will eventually live.

1
Observed agent steps
What did the agent actually do? The ordered sequence of searches, reads, tool calls, drafts, and verifications, read from telemetry.
In this brief
2
Human-equivalent time per step
If a person executed each step, how long would it take? A calibrated lookup, tcal,i, built by research Stages 1–5.
In this brief
3
Sequence (effort) discounting
Aggregate the per-step times without overcounting repetition. Research Stage 6. Output = New Assisted Hours for the session.
In this brief
4
Complexity & user adjustment
Adjusting for task complexity and individual user context.
Out of scope
5
Time saved & net capacity
Translating human-equivalent work into net organizational capacity.
Future work
Figure 1 — The five-layer New Assisted Hours architecture. Layers 1–3 are covered here; the six research stages populate Layers 2 and 3.

The architecture is informed by research conducted across these six stages: Layer 2 is produced by Stages 1–5, which together yield the calibrated per-step human-equivalent time tcal,i stored in a lookup table, and Layer 3 is Stage 6, the effort-discounting rule that aggregates those per-step times into the session-level metric.

Methodology

The six research stages

The metric is constructed from original research designed to build the human baselines an auditable productivity measure requires. Stages 1–5 produce the calibrated per-step estimate tcal,i; Stage 6 aggregates it after applying a discounting weight for repeated steps. The pipeline below shows the stages at a glance, and each is described in turn.

Stages 1–5 · Producing the calibrated step-time inputs (tcal,i)
STAGE1
Observed human time — RCT
Within-subject randomized controlled trial across three knowledge-work tasks. 197 participants
STAGE2
Revised document-review RCT
Re-run on denser, less-searchable documents to force genuine reading. 133 participants
STAGE3
Whole-task PERT time estimation
Three-point human estimation of full-task time across seven task families. 302 respondents / task
STAGE4
Microtask decomposition and time estimation
Workflows mapped to tool-level, effort-bearing units. 14 tool-primitive microtasks
STAGE5
LLM benchmark creation & calibration
Four LLMs benchmarked, then prediction-bin calibrated to observed times. MAE 11.7 → 1.99 min
↓ produces tcal,i ↓
Stage 6 · Effective-effort discounting (Layer 3)
STAGE6
Effective-effort discounting
Concave, order-aware aggregation with repetition-sensitive discounting applied across all steps. discount paramter α = 0.7 · ~58% compression
Figure 2 — The six-stage estimation pipeline. Stages 1–5 build the calibrated per-step time; Stage 6 aggregates it into New Assisted Hours.

Stages 1 & 2 — Observed human time (two RCTs)

The baselines start in the lab. A first within-subject randomized controlled trial measured how long people actually take on three realistic tasks — infographic creation, executive-brief review, and comparison research — each performed both with and without AI assistance, alongside NASA-TLX workload, click-based strain, and abandonment. AI's effect on completion time was task-specific: significant on the infographic and comparison tasks, flat on the executive brief. Notably, subjective effort fell on all three tasks even where time did not — the first direct evidence that capacity and elapsed time are distinct.

Because the first document task could be shortcut with keyword search and copy-paste, a revised trial hardened it with longer, denser, less-searchable documents to force genuine reading and synthesis. Across 133 participants, AI produced a statistically significant reduction in completion time, and the robust medians became stronger ground-truth labels for calibration.

Stage 1 · Study snapshot
Within-subject RCT (Prolific) · 1,111 sign-ups → 197 analyzed · human time = no-AI median
TaskHuman, no-AIWith CopilotImpact
Certification comparison12.5 min5.7 min−32% ★
Infographic creation10.0 min7.1 min−21% ★
Executive brief8.4 min8.8 min+5% n.s.
★ statistically significant (p < .05); n.s. = not significant. Subjective effort (NASA-TLX) fell on all three tasks.
Stage 2 · Study snapshot
Revised document-review RCT (Prolific) · N = 133 · denser, less-searchable documents
TaskHuman, no-AIWith CopilotImpact
Business Analysis Snapshot7.4 min6.0 min−19.5% ★
Operations Snapshot7.3 min7.4 min+1.6%
Overall7.4 min6.8 min≈ −8% ★
Human time = control (no-AI) median. On the Operations task the median was flat; the significant gain appeared in the mean (−28 s, −5.9%).

Stage 3 — Whole-task PERT time estimation

Controlled trials can only cover a handful of tasks, so coverage is extended with whole-task human estimation using an uncertainty-aware three-point PERT method — optimistic, most-likely, and pessimistic times combined as (O + 4M + P) / 6, with the median PERT as the primary estimator. Independent knowledge workers estimated full manual completion time for standardized whole-task vignettes across seven benchmark-style families (from academic editing to open-ended research), with 302 respondents per task. The resulting medians span roughly 92.5 to 700 minutes and are strongly right-skewed (skewness 6.2–17.2) — the long-tailed shape knowledge work characteristically takes.

Stage 4 — Microtask decomposition and time estimation

Agent telemetry is expressed as tool calls, not whole tasks, so the whole-task and RCT baselines must be bridged to the level at which agents actually operate. Stage 4 decomposes workflows into 14 tool-primitive microtasks — the small, effort-bearing units (search, open/read, extract, summarize, verify, refine, and so on) that compose an agent session. These become the unit at which per-step human-equivalent time is estimated and calibrated.

Stage 5 — LLM benchmark and calibration

To scale beyond direct human elicitation, four LLMs were benchmarked against the 14 microtasks, evaluated by mean absolute error (MAE), RMSE, Spearman rank correlation, pairwise concordance, and leave-one-task-out cross-validation. Out of the box the models are not usable as time estimators: raw estimates overestimated 79–86% of the tool-level microtasks, with very weak rank correlation to observed times (ρ ≈ 0.12). A prediction-bin calibration — scaling estimates by predicted-duration bin — corrected the bias and, under leave-one-task-out cross-validation against ground-truth RCT timings, reduced MAE from about 11.7 minutes to about 1.99 minutes, outperforming model-family and complexity-based corrections. The bias direction is level-dependent: models overestimate microtasks but underestimate whole tasks, so calibration against human anchors is not optional.

Predicted-duration binCalibration multiplier
≤ 7 min× 1.25
7–15 min× 0.73
15–30 min× 0.37
≥ 30 min× 0.24
Stage 5 · estimation prompt (excerpt)
You are an expert evaluator assessing task completion time complexity. Estimate how long it would take a skilled professional with relevant domain expertise (but no task-specific context) to complete this task with 50% success rate.
Assessment Criteria · Time Horizon Categories · Output Format — with per-component reasoning and a comparison anchor.

Stage 6 — Effective-effort discounting

Adding up per-step human-equivalent time overlooks how people actually work: humans exert less effort on a step the more they have already done it. The first search establishes context; later searches build on it. The first draft creates the backbone; later revisions make smaller changes. Stage 6 applies a concave, order-aware discount so repeated occurrences of a step receive progressively less marginal credit. Letting tcal,i be the calibrated human-equivalent time for the i-th step, the naive baseline and the discounted rule are:

effective effort = Σi tcal,i × w(occi) where the marginal weight of the k-th occurrence of a step is w(k) = kα − (k−1)α

At α = 1 the weights sum to a plain count and the rule reduces to naive addition. At α < 1 the function is concave, so the first occurrence keeps full credit while each repeat contributes less than the last. The default is α = 0.7 — a governed parameter set mid-range of the effort-discounting literature and stress-tested with sensitivity analysis. Internal agent reasoning steps are excluded from scoring throughout.

Results

What the numbers show

The RCT snapshots above are the observed human baselines. The remaining results establish that the estimation scales and that the discounting behaves as intended.

Whole-task human time spans an order of magnitude

The PERT medians confirm how wide and skewed real task times are — from short clarifications to multi-hour research workflows.

92.5 min
abstract clarification (shortest median)
700 min
prompting-methods test (longest median)
Figure 3 — Whole-task human-time medians across seven task families (Stage 3), strongly right-skewed (skewness 6.2–17.2). Trimmed-10 means range from about 119 to 1,103 minutes.

Calibration turns a biased proxy into a usable signal

The single most important result for scaling: prediction-bin calibration cuts per-microtask error by roughly an order of magnitude against ground-truth RCT timings.

Raw LLM estimate
MAE ≈ 11.7 min
After calibration
≈ 1.99 min
Figure 4 — Mean absolute error against observed human times, under leave-one-task-out cross-validation. Raw estimates overestimate 79–86% of microtasks (ρ ≈ 0.12); calibration outperforms model-family and complexity-based corrections.

Discounting compresses repetition, not the backbone

The discounting layer is grounded in telemetry from 99 Deep Research sessions, which are strongly repetition-heavy. On a worked example, discounting removes credit from late-stage repeats while leaving early backbone steps at full value.

Naive sum
≈ 60.5 min
Effort-discounted
≈ 25.6 units
Figure 5 — A 34-call worked example (2 internal reasoning calls excluded): order-aware discounting compresses ~60.5 calibrated minutes to ~25.6 effective-effort units at α = 0.7 (~58% compression), concentrated on late-stage retries.
31
median scored tool calls per session (90th percentile 46; max 107)
27
of those, median calls that repeat an already-used tool
10.3 min
median agent runtime across sessions with a recorded runtime
99
Deep Research sessions in the discounting telemetry sample

Figure 6 — Session-level telemetry behind the discounting layer: multi-step, repetition-heavy workflows are the norm, which is exactly the setting the concave discount is designed for.

Validation

Why the signal holds together

Two checks support using the metric. First, external agreement: the calibrated estimates closely match independent human estimates of how long the same tasks would take — the property you want before trusting a New Assisted Hours number. Second, internal validity: the calibration is validated under leave-one-task-out cross-validation against ground-truth trial timings, and it outperforms model-family and complexity-based corrections. Two design choices run through the whole program and explain why the results cohere: task difficulty is anchored in observed human time rather than tokens or model internals, and aggregation is built to be robust to outliers so a few unusually large or small sessions do not distort the picture. The consistent lesson is that raw model estimates cannot substitute for human-grounded baselines — calibration against human anchors is what makes the metric trustworthy.

Effort discounting in practice

The discounting layer was exercised on telemetry from 99 Deep Research sessions, including agent runs on tasks similar to those used in the RCTs plus a fully worked-out example. Because the rule is order-aware, the compression a session receives scales with how repetition-heavy it is. A few worked out examples from tasks similar to the RCT tasks make the behaviour concrete (values are agent-session calibrated totals, not the 15-minute-capped RCT times):

Agent session Naive additiveEffective effort (α = 0.7)Compression
Infographic task — similar to the RCT infographic task≈ 21 min≈ 15 min≈ 30%
Document-review task — similar to the RCT Biz-snap document task≈ 102 min≈ 37 min≈ 64%
Fully worked-out example (34-call Deep Research session)≈ 60.5 min≈ 25.6 min≈ 58%

The compression depends on the task as seen above. A light, low-repetition session loses little — there is barely any iteration to discount — while a longer, retry-heavy session is compressed substantially. In every case the discount concentrates on late-stage repeats, so the early backbone steps of a workflow retain their full human-equivalent credit and only the redundant later passes are down-weighted.

Scope

What the metric does and does not claim

New Assisted Hours measures the human-equivalent effort represented by the agent's actions — the depth and volume of work performed. It deliberately does not, on its own, assess output quality, whether the user ultimately used the result, or direct economic outcomes such as revenue or ROI; those are necessary complements addressed separately. It is also a modular layer: the discount strength is set by governance rather than learned per user, and complexity/user adjustment (Layer 4) and net-capacity accounting (Layer 5) remain future work.

  • Human-equivalent time per session — the calibrated, pre-discount estimate summed over steps.
  • Effective-effort units — the discounted total that becomes the New Assisted Hours figure.
  • Repetition ratio — repeated tool calls as a share of scored calls; the primary driver of discounting.
  • Discount factor α — default 0.7; a governed parameter, not a fitted constant.
  • Calibration error (MAE) — tracked against any available ground-truth timings.
Appendix A

Study snapshots

Full results for the two randomized controlled trials that anchor Stages 1 and 2. Human time is the control (no-AI) median; impact is the change from that baseline.

Stage 1 · Study snapshot
Within-subject RCT (Prolific) · 1,111 sign-ups → 197 analyzed · human time = no-AI median
TaskHuman, no-AIWith CopilotImpact
Certification comparison12.5 min5.7 min−32% ★
Infographic creation10.0 min7.1 min−21% ★
Executive brief8.4 min8.8 min+5% n.s.
★ statistically significant (p < .05); n.s. = not significant. Subjective effort (NASA-TLX) fell on all three tasks.
Stage 2 · Study snapshot
Revised document-review RCT (Prolific) · N = 133 · denser, less-searchable documents
TaskHuman, no-AIWith CopilotImpact
Business Analysis Snapshot7.4 min6.0 min−19.5% ★
Operations Snapshot7.3 min7.4 min+1.6%
Overall7.4 min6.8 min≈ −8% ★
Human time = control (no-AI) median. On the Operations task the median was flat; the significant gain appeared in the mean (−28 s, −5.9%).
Appendix B

The Stage 5 estimation prompt

The exact prompt used in Stage 5 to elicit per-microtask human-time estimates from the benchmarked LLMs. Model outputs from this prompt are the raw estimates that the prediction-bin calibration then corrects against observed human times. Bracketed items are placeholders filled at runtime.

Stage 5 · LLM time-estimation prompt

You are an expert evaluator assessing task completion time complexity. Estimate how long it would take a skilled professional with relevant domain expertise (but no task-specific context) to complete this task with 50% success rate.

Task to Evaluate
  • [INSERT TASK DESCRIPTION HERE]
Assessment Criteria
  • Cognitive Steps Required: Count distinct reasoning/action steps
  • Information Gathering: How much research/exploration needed?
  • Error Recovery: How forgiving of mistakes?
  • Verification Burden: How hard to confirm correctness?
  • Context Requirements: How much domain knowledge assumed?
Time Horizon Categories
  • Atomic (< 1 minute): Single-step actions, immediate decisions
  • Short (1–10 minutes): Simple multi-step procedures, basic problem-solving
  • Medium (10–60 minutes): Complex procedures, requires planning
  • Long (1–8 hours): Substantial projects, multiple components
  • Very Long (> 8 hours): Major undertakings, extensive coordination
Output Format
  • Estimated Time: [X minutes/hours]
  • Confidence: [Low/Medium/High]
  • Primary Time Drivers: [List 2-3 main factors]
  • Comparison Anchor: [Similar to: "writing a bug fix" or "creating a presentation"]
Provide reasoning for each assessment component before giving final estimate.
References
  1. Stevens, S. S. (1957). On the psychophysical law. Psychological Review, 64(3), 153–181.
  2. Białaszek, W., Marcowski, P., & Ostaszewski, P. (2017). Physical and cognitive effort discounting across different reward magnitudes. PLoS ONE, 12(7), e0182353.
  3. Shenhav, A., Botvinick, M. M., & Cohen, J. D. (2013). The expected value of control. Neuron, 79(2), 217–240.
  4. Kurzban, R., Duckworth, A., Kable, J. W., & Myers, J. (2013). An opportunity cost model of subjective effort and task performance. Behavioral and Brain Sciences, 36(6), 661–679.
  5. Cuevas Rivera, D., Strobel, A., Goschke, T., & Kiebel, S. J. (2020). Modeling dynamic allocation of effort in a sequential task using discounting models. Frontiers in Neuroscience, 14, 242.
  6. Donner, Y., & Hardy, J. L. (2015). Piecewise power laws in individual learning curves. Psychonomic Bulletin & Review, 22(5), 1308–1319.
  7. Hart, S. G., & Staveland, L. E. (1988). Development of NASA-TLX. In Human Mental Workload (pp. 139–183). North-Holland.
  8. Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192.
  9. Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work (NBER Working Paper No. 31161).
  10. Kwa, T., West, B., Becker, J., et al. (2025). Measuring AI ability to complete long tasks. arXiv:2503.14499.