"We Use AI a Lot" — Compared to What? Three Questions That Measure Whether AI Helped
"We use AI a lot" is a claim about effort volume: it says how much was used, not what came out. The same sentence can describe pushing against capacity and describing an unplanned quarter, so on its own it carries no information. This piece separates the usage metric from the impact metric, shows how framing AI as a coach rather than a production machine changes what you measure, and sets out three questions — Comparison, Decision, Durability — with evidence from METR, McKinsey, PNAS and the QJE, each with its limits stated. It closes with the serious objections and a short FAQ.
For the past year I have heard the same sentence in almost every meeting: "We use AI a lot." It is said as an announcement of success, and the report on the table backs it up: licences held, active users, prompts per month, content produced. Everyone nods. In me that sentence raises the same question every time.
A lot compared to what?
When I ask, the answer usually doesn't come, because the sentence carries no denominator. "A lot" is not a measure; it is an impression. This piece neither defends AI nor attacks it; it asks whether that sentence is a measure or an impression. If the answer is impression, I owe you something to put in its place — that is the debt this piece has to pay.
A lot compared to what?
By a usage metric I mean any measure that counts the resource spent — whatever is on that report. All are real numbers, all are easy to collect. What they share: none asks what changed at the end of the work.
The numbers say the denominator is missing too. In Bick, Blandin and Deming's NBER adoption study (32966, revised February 2025) — by adoption I mean the share of people who have started using the tool, not how intensively they use it — 23% of workers had used generative AI at least once for work in the previous week. Against that, only 1% to 5% of all working hours were AI-assisted, and the time savings people reported came to 1.4% of total hours. The savings are self-reported, quality was not measured, and the figures are a snapshot of late 2024. The number does not size the gain; it sizes the distance between adoption and the gain we can see.
The corporate version is sharper. In McKinsey's 2025 global survey (1,993 respondents, 105 nations, fielded June-July 2025), 88% of respondents say their organisation uses AI regularly in at least one business function, and only 39% report any level of EBIT impact — and most of those say it accounts for less than 5% of EBIT. This is a self-reported executive survey, and McKinsey sells consulting on the same subject. I have to apply my own rule to this number as well: 39% cannot tell "AI produced no effect" apart from "the organisation could not see the effect it produced". That is not an objection to my thesis; it is the thesis — the organisation is talking without a denominator about its own AI claim too.
And when you cannot measure, two opposite explanations stay standing side by side: we are pushing against our capacity, or the work we entered without a plan is working us hard. If a number cannot tell apart the two options of the decision you want to make with it, it is not evidence for that decision. It stands as a measure; it does not stand as a reason.
Three measurement layers, and effort's shelf life
I sort everything we measure into three layers. An input metric counts the resource spent: prompts sent, hours in the tool. An output metric counts the product: a forty-slide module, a ten-page report. An outcome metric counts the changed state of the work: a decision made better, a mistake that didn't repeat. The distinction from the opening sits here: a usage metric is the name of the input layer, and what I call an impact metric is the outcome layer. The output layer sits between them and keeps standing in for the outcome in reports.
I have watched one example of this for ten years. In corporate training, the completion rate is an output metric; nobody claims it measures learning, yet it never leaves a report — it is easy to collect and it looks green. Today the prompt count is at the start of the same road. I have seen this film once; it ended with nobody defending the number and nobody removing it either. I have written separately about why the completion rate doesn't survive the CFO's desk.
The problem is less the number itself than how we obtain it. In METR's 2025 randomised controlled trial, sixteen experienced developers completed 246 real work items in their own mature repositories; afterwards they said they had worked 20% faster, and the measured time had risen by 19%. This does not say AI slows people down — sixteen people, and a mature codebase is the hardest scenario AI faces. METR then changed the design of that experiment in February 2026 and wrote down its current view: selection effects made the central estimate unreliable, but developers are probably getting more speed-up in early 2026 than in early 2025. The finding did not close in my favour; it stayed open. A serious team tells you the shelf life of its own headline.
An old principle follows: when a measure becomes a target, it stops being a good measure — the principle is Charles Goodhart's, the formulation Marilyn Strathern's. Even so, there is one phase in which effort volume does carry information: the exploration phase, and its validity ends the moment the limit is found. An effort metric is not wrong; it is a metric with a shelf life — and when the date is never declared, it turns into a permanent excuse.
One last distinction, because "efficiency" left undefined justifies everything. Efficiency is the resource spent per unit of output; effectiveness is producing the right output. AI often raises efficiency and lowers effectiveness: the wrong thing gets produced faster.
A production machine, or a coach?
AI is not a production machine. It is closer to a coach. That is a slogan; behind it sits a design decision, and the consequence of the decision is a metric: a machine counts volume, a coach counts the improvement in a person's work. You can object that we count inputs simply because inputs are easy to count. The objection is correct and it is not enough: ease explains what we count, not what we treat as success. Putting the prompt count on a panel is convenience; making it a quarterly target is a framing decision.
I separate the two along two axes. Direction of information: with a tool I speak and it acts; with a coach it asks, and it tests my framing. Location of responsibility: a mistake gets charged to the tool, while in coaching the athlete still runs the race. The most valuable thing AI can do is not answer my question but tell me I am asking the wrong one.
One experiment shows the framing flipping the sign of the result. In the field experiment Bastani and colleagues published in PNAS in 2025, roughly 1,000 students at a high school in Türkiye were split into three arms; the arm using the standard ChatGPT interface scored 48% better on practice problems and 17% worse than the control group on the exam taken after AI access was removed. In the "GPT Tutor" arm the harm was largely erased — the harm removed, not learning added: GPT Tutor did not raise learning. One school, high-school students, the GPT-4 of 2023 — it says nothing directly about adult learning inside a company. What transfers is not the number but the mechanism: the gap that appears when the support is taken away.
Where the gain comes from has to be asked too. In the experiment Noy and Zhang published in Science in 2023, 453 professionals cut their time on writing tasks by 40% and raised quality by 18%. The authors' own reading is that ChatGPT substituted for the worker's effort rather than complementing their skill. What was measured here is performance at the moment of the task; whether anyone had become a better writer was not measured. Soderstrom and Bjork's 2015 review draws exactly this distinction: what we can measure is performance, and performance is an unreliable index of learning. The review says nothing about AI and rests on laboratory work; but it describes the gap in the Bastani experiment from a completely different literature. The 40% and the 18% show that a tool was at hand that day, not that a person had improved.
The metaphor has a limit that should be written down: AI is not a coach in any real sense — it has no stake in your development and no memory of your career. The metaphor is not a claim but a design constraint that decides what we measure.
A tool executes an instruction; a coach asks a question. The first produces a file; the second produces your next decision.
Did AI actually help? Three questions
What we need to measure is not how much AI we used — how much of the actual work it moved. Left there, the sentence stays a wish. So I ask three questions: Comparison, Decision, Durability.
First a principle. The answer to unmeasurability is not a bigger metric but a narrower claim. "AI raised revenue by X%" cannot be proved; "this job used to take a specialist four hours, it now takes forty minutes of production plus twenty minutes of checking" can be proved and can be falsified.
Comparison
Decision
Durability
Only Comparison is independent of the framing — it is the minimum threshold any framing has to clear. Decision and Durability are the two tests the coach framing adds; under the production-machine framing both are meaningless, because there whether the person changed is not something to be measured. An organisation that passes only the first has become cheaper, not better.
The weakest point of these three tests is my own evidence standard: all three rest on self-report, and METR's experiment shows exactly that danger. So I build them as records rather than surveys: take the old duration from the calendar, write down the decision that changed the moment the session ends, and have a third person read the month-one and month-three output against the same standard. A report is not a test; a record is.
Two experiments together show why one test alone isn't enough. In the pre-registered 2023 field experiment Dell'Acqua and colleagues ran with 758 BCG consultants, on tasks chosen inside AI's capability frontier quality came out more than 40% higher than the control group's; on a task chosen outside that frontier the probability of reaching the correct solution fell by 19 points. That 40% is quality, not productivity; and it is one company, one profession, with the "outside" task deliberately built to trip AI up. In Peng and colleagues' 2023 Copilot experiment, a task written from scratch shows a 55.8% speed-up, with a confidence interval running from 21% to 89% — a preprint that has not been peer reviewed, with authors at Microsoft and GitHub; one task, and quality was not measured. Same technology, different class of task, opposite result: the frontier is jagged.
Keep the record to one line — it takes under a minute: name of the job · how long it used to take · how long it takes now · what it replaced · who approved the quality. If the measurement costs more than the work, shrink the scope rather than the measurement.
My own work is not exempt: a training module we used to prepare in three days now comes out in forty minutes — our own record, not an independent measurement. It passes Comparison; if behaviour doesn't change, Decision and Durability are still open. Passing one test is not passing three.
Three questions decide whether AI helped: what did it replace, which decision did it change, and did it reduce the need the second time round.
The serious objections
"Usage metrics are necessary too; you cannot measure impact without measuring adoption." True, and I am not throwing them out. The distinction: usage is a diagnostic metric, not an outcome metric. When impact comes out low, "it was never used" and "it was used and it didn't help" call for different treatments. The rule: usage may sit in the diagnostic column, never in the outcome column.
"Impact in knowledge work cannot be measured anyway; AI is not the only cause of the output." A serious objection, and the answer is the principle above: AI's share of a company P&L cannot be separated out, a comparison at task level can. If you need a stronger claim, run a 90-day pilot with a control team.
"The coach metaphor is romantic; I want to finish my work, not be questioned." You are right — this is exactly where the boundary sits. In routine work with a fixed specification the tool framing is correct and coaching is waste. The evidence points the same way, and it points against me: in the study Brynjolfsson, Li and Raymond published in the QJE in 2025, across 5,172 customer support agents the number of issues resolved per hour rose 15% on average; the gain is concentrated among the least experienced (around 30%). Among the highest-skilled there is no meaningful change in productivity: small gains in speed, and small but statistically significant declines in quality — resolution rates and customer satisfaction. That finding does not support my thesis; it bounds it. One firm, one function; not a randomised experiment but a quasi-experimental estimate drawn from a staggered rollout. When the specification itself is the problem, a machine that skips the question multiplies the wrong work. My claim isn't "don't use AI as a tool"; it is know which of the two you are in, and don't leave that decision to the tool's framing.
How we build this into Blink AI
When I look for the three tests inside our own product, the picture is this: we partly measure two of them and we do not measure the third at all. Comparison has no counterpart in the product — we do not know how long a piece of training would take without AI. The closest we stand to Decision is the depth of the interaction; which session changed a decision, we cannot measure. For Durability all we hold is the direction of pre- and post-test — a signal that claims no causation and is not even a reporting line: it is a run-time signal that shapes the coach's answer in the moment, and the coach is forbidden to write the number out. We did not delete the usage counts either: tokens, cost and call counts sit on one panel, engagement depth and the skip rate on another — separate, because they do not measure the same thing.
The coach's first rule is a prohibition: the first line of the Socratic engine's system prompt reads "never give a direct answer". After a wrong answer the coach can neither celebrate nor write out the correct answer; it steers towards the right concept with a question, and normalising a wrong answer as "perfectly natural" is a separate violation type. We also measure the coach breaking that rule: five violation types are scanned on every turn of the flow and written to a persistent table. What we measure is not how much the model talks but how closely it follows the rule — and we track the false-positive rate of those patterns against our own content corpus, though that is an internal check, not a published finding.
Our unit is not "how many tokens" but "how deep": median message length, question rate, substantive-reply rate — all deterministic, with no LLM call. The coach's "what should I go back to?" suggestion is not invented either; the list is built only from content objects the learner actually reached. The streak advances on qualified days rather than logins, badges carry no ranking or comparison, and inside the micro-feedback flow the programme closes for that calendar day once a module is finished.
What we don't measure decides just as much. No proctoring, no screen time, no attention or gaze tracking. No calendar integration — we write no blocks into anyone's diary. No spaced-repetition scheduler. And no automatic competency level for any topic.
To be honest, the hours metric has not left the product: a team list sorted by minutes still sits in the manager's panel. We didn't remove it; we put the depth signal beside it — rather than deleting a number that says nothing on its own, putting next to it a number that says what it means is more informative.
Key takeaways
- "We use AI a lot" is a statement of effort volume; because it describes both pushing against capacity and working without a plan, it carries no information on its own.
- A usage metric sits in the diagnostic column, not the outcome column. Zero usage guarantees zero impact; high usage guarantees nothing.
- The framing decides the metric: treat AI as a production machine and you count volume, treat it as a coach and you count the improvement in a person's work.
- Apply the three questions to a job, not to an organisation: what did it replace, which decision did it change, did it reduce the need the second time round. Let the answers come from a record, not a report.
- The goal is not more production. Not more hours. More focused ones.
Frequently asked questions
What if the team starts gaming these three tests?
They will — my own instrument is not exempt from Goodhart. The moment the Decision test becomes a target, people learn to write "my decision changed". Three safeguards help: apply the test to the job rather than the person; keep the record as free text and never convert it to a score; and read all three together, because the calendar checks Comparison and unaided output checks Durability.
Do the three tests belong in a performance review?
They should not. These tests measure where the tool helps, not how well a person works — the same employee can produce a large gain on a task inside the frontier and a loss on one outside it, and what makes the difference is the class of task. The moment the record is tied to an appraisal it stops being honest. Report the results at the level of the task class, not the person.
Usage is up but output hasn't improved — what now?
Look at the class of task first: is the work inside or outside AI's capability frontier (Dell'Acqua and colleagues, 2023). Then look at the shelf life: are you in the exploration phase, and is there a declared end to it? Third, apply the Decision test. If there is production but judgement isn't moving, change where the tool is installed rather than dropping it — move it from a job on the input side to a job where a decision is made.
What happened to the competency signal in your earlier posts?
I am correcting the framing here. Blink AI does not generate an automatic competency level for every topic; the schema and the endpoint exist, the automatic generation does not. Two things get measured: engagement depth (median message length, question rate, substantive-reply rate) and the direction of pre- and post-test — the second with no claim of causation. Completion still sits in the record; it just doesn't sit in the outcome column.
Does Blink AI monitor how employees use AI? (GDPR/KVKK)
A manager sees only their own team and the rows carry names; in the engagement report what is measured is the form of the conversation rather than its content — length, question rate and substantive-reply rate are computed in the database, and the raw text is never loaded into the application for this report. Beyond that, a manager can see a learner's learning-object answers in programmes they created themselves; the only role that can read the chat text verbatim is the platform administrator, and only inside the session of the relevant programme. Cohort summaries apply a k-anonymity threshold (5 by default); individual rows do not. The camera is active in exactly two places and only with explicit consent — the occupational-safety attendance selfie (with optional location capture) and the learner-initiated video reflection; both record the consent version, the retention period and the deletion. There is no camera access anywhere else. No proctoring, no screen time, no attention tracking.
When the measure changes, the target changes with it. The goal is to spend hours on fewer, better-chosen pieces of work. Not more hours. More focused ones.
Related reading
- 9 min read
The Speed Trap: AI Made Content Easy, Not Learning Ownership
AI can turn an idea into a working product in hours — but producing a course and owning the learning are two different jobs. Shipping two AI apps back to back surfaced a speed trap that sits at the heart of learning design too. This piece reframes the "user manual" trap, generic soulless content, and the three filters to run before shipping — empathy, legal, first impression — for L&D teams.
Read post - 12 min read
L&D on the CFO's Desk: 3 Visible Numbers, 7 Hidden Costs
"94% completion rate" — that number doesn't carry the L&D budget into the next quarter. It measures attendance, not behaviour change. This piece walks through the three financial metrics a CFO reads directly off the P&L and the seven hidden costs that never appear on it — with research from Deloitte, LinkedIn and McKinsey — followed by a 90-day decision matrix and a short FAQ.
Read post