The best answer to how to measure AI productivity is not hours saved. Measure whether AI increases the value and quality of completed outcomes, reduces total cycle time and rework, and leaves the organisation more capable of making the next decision. Time saved belongs on the scorecard. It should not be allowed to chair the meeting.
Time saved is the productivity metric everyone loves because it can be counted before anyone asks what happened to the work.
That does not make it useless. It makes it incomplete. An hour removed from one task may become an hour of review elsewhere. A faster draft may create a better decision—or a small crisis in Legal.
The suspiciously tidy number
"Hours saved" thrives in AI business cases because it is legible, flattering and available early. Subtract the new task time from the old and multiply by headcount. The result fits neatly into a cell, which is more than can be said for most organisational change.
Value arrives later and may belong to a team rather than one user. It depends on whether the task was worthwhile, the outcome improved and the saved time was used productively.
METR's May 2026 survey illustrates the difference. Its 349 technical-worker respondents reported a median threefold change in speed, but a median 1.4-to-twofold change in the value of their work. METR's own formulation is the useful one: "speed changes would typically overstate value changes."
The caveat belongs beside the number. This convenience-sample survey relied on self-report; response rates were low, selection bias was plausible and METR urged scepticism about the magnitude. It exposes a category error rather than a universal multiplier: speed and value are related, but not identical.
Faster at the task, slower in the system
Task efficiency measures what happens inside one activity. System throughput measures how quickly valuable work crosses the whole workflow.
The distinction appears in DORA's generative-AI research, updated in April 2026. Developers with greater adoption reported higher individual productivity. Yet a 25 per cent increase was associated with 1.5 per cent lower delivery throughput and 7.2 per cent lower delivery stability. DORA points to larger batches and slower review as plausible mechanisms.
These are associations from software development, not proof of causation and not a template for every sector. The systems lesson travels further: output can accelerate into a review queue, governance gate or dependency that has not moved.
A bottleneck with more documents approaching it is still a bottleneck. It is merely better informed about the traffic.
Productivity includes the work AI creates
AI can remove drafting, searching and formatting. It can also create prompt iteration, source checking, correction and the peculiarly modern task of finding which confident sentence invented its evidence.
If AI removes ten minutes of drafting and creates 25 minutes of forensic reading, recording only the first number is not measurement. Recording only the second is not fair either. Both belong to the same workflow.
METR's February 2026 methodological update shows why measurement is changing. Concurrent agent work made self-reported task time unreliable. Developers also changed which tasks they submitted because they did not want to complete them without AI. The tool altered both the clock and the choice of work.
Count the labour around the output: review minutes, corrections, source checks, escalations, duplicate work and explanation. Time saved should be net of the work needed to make the result usable.
Work can improve while capability declines
A good result today and a capable team tomorrow are different assets.
An MIT Media Lab preprint studied LLM-assisted essay writing. Across the first three sessions, 54 participants used an LLM, search engine or no tool; 18 completed a fourth crossover session. The LLM group showed the weakest measured brain connectivity, reported the least ownership of their essays and struggled more to quote their own work.
This is early evidence from a small educational task, not proof of general workplace deskilling. It is a reason to ask whether people can still explain, challenge and improve the work.
Knowledge retention, confidence calibration and exception handling are productive assets. If a system improves this week's answer while weakening next month's judgement, the liability belongs on the same balance sheet.
How to measure AI productivity: five measures, not one
No universal AI productivity metric can capture every workflow. A credible scorecard can, however, ask five consistent questions.
| Measure | The question | Example evidence |
|---|---|---|
| Outcome value | Did the work improve a result that matters? | Revenue, completion, risk reduction, customer outcome, decision quality |
| Quality | Is the finished work more accurate and useful? | Error rate, acceptance rate, defects, source coverage |
| Total cycle time | Did the whole workflow move faster? | Signal-to-decision, decision-to-action and end-to-end completion time |
| Verification and rework | What new checking or correction did AI create? | Review minutes, reversals, duplicate work, escalation and correction rates |
| Human capability | Is the team better able to judge the next case? | Knowledge retention, confidence calibration, exception handling, learning transfer |
For each use case, choose one primary outcome and two safeguards. A support tool might optimise resolution quality while guarding rework and escalation. A research tool might optimise cycle time while guarding source coverage and retained understanding.
The point is not to produce a more ornate dashboard. It is to prevent a convenient proxy from becoming the objective.

Run an AI productivity trial that can survive contact with finance
A credible trial requires a baseline, a comparison and some resistance to decorative arithmetic.
- Define the outcome first. Name the result the workflow should improve before introducing the tool.
- Establish a baseline. Use comparable work and record its quality, cycle time, rework and exceptions.
- Measure end to end. Include preparation, AI use, review, approval, integration and completion.
- Separate perception from observation. Ask users about uplift, but distinguish estimates from recorded change.
- Capture quality and exceptions. Record defects, corrections, rejected outputs, escalations and unusual cases.
- Review after 30, 60 and 90 days. People change how they work and which tasks they attempt as they learn the system.
Teams can use matched cases, phased roll-outs or periods before and after deployment, provided they document material differences. Finance should see the baseline, assumptions and uncertainty—not merely an ROI percentage with suspiciously excellent posture.
A Danish study published as NBER Working Paper 33777, revised in October 2025, found rapid chatbot adoption, reported productivity benefits and new AI-related tasks. Yet its administrative data showed no measurable effects larger than two per cent on recorded earnings or hours in the first two years. Work changed before the headline measures moved.
It is a working paper from one national setting, not the last word. Its lesson is methodological: behaviour, task mix and organisational structure may change before lagging indicators. A 90-day trial should measure workflow change without pretending to settle the long-term economics.
The Attimo view: attention is part of the result
Productivity also depends on what the organisation can attend to—and safely forget.
Attimo's Attention Cycle follows four moments:
- 01Notice: does the right signal surface when it matters?
- 02Clarity: does the system reduce reconstruction and uncertainty?
- 03Action: does the right person move the work with the right context?
- 04Memory: does the learning remain available and inspectable?
Attention is not an unlimited human resource. As our article on cognitive friction in AI workspaces argued, automation can reduce manual effort while increasing the mental labour of checking, switching and reconstructing context.
AI should increase human capability, not reduce human authorship. The best system does not simply make a task disappear faster. It helps people notice what matters, decide with better context, complete the work and retain what was learnt.
The answer to how to measure AI productivity is to keep time saved, but place it beside value, quality, total cycle time, verification and human capability. Then ask what the timesheet cannot: what can the organisation now notice, decide and complete that it could not before?
Speed is a benefit. It is not the balance sheet.
Explore how Attimo builds human-first AI around the moments that matter.
Frequently asked questions
How should companies measure AI productivity?
Measure one primary business outcome and at least two safeguards for each AI use case. Assess outcome value, finished-work quality, end-to-end cycle time, verification and rework, and retained human capability. Include the whole workflow and distinguish self-reported uplift from observed change against a baseline.
Why is time saved a weak AI productivity metric?
Time saved measures a task-level input, not necessarily a business result. It can miss low-value work, displaced review, correction and integration costs, or a downstream bottleneck. It remains useful when measured alongside outcome value, quality, total cycle time and the work required to verify the result.
What is the difference between task speed and business value?
Task speed is how quickly an activity is completed. Business value is the improvement in an outcome the organisation cares about, such as revenue, completion, risk, customer experience or decision quality. AI can accelerate a task without improving the result—or enable valuable work that was previously impractical.
How can organisations measure the hidden cost of AI verification?
Track the minutes spent checking sources, correcting outputs, repeating prompts, integrating results and explaining decisions downstream. Record rejected outputs, reversals, duplicate work and escalations. Compare these costs with the labour removed from the original task to calculate net time and workflow value.
What should an AI productivity scorecard include?
Include the target outcome, baseline, observed change, quality indicators, end-to-end cycle time, verification and rework costs, exception rates and measures of retained human judgement. Label survey estimates separately from operational data, and review the scorecard as users' behaviour and task choices change.
Sources
- METR2026 AI Usage Survey, 11 May 2026.METR
- METRUplift Measurement: A February 2026 Methodological Update, 24 February 2026.METR
- DORAState of AI-Assisted Software Development, updated April 2026.DORA
- MIT Media LabYour Brain on ChatGPT, preprint.MIT Media Lab
- NBER Working Paper 33777Revised October 2025.NBER