VSP Demo?

Back to Blog

Your Copilot ROI Story Is Missing the Baseline — And That's Why Your CFO Doesn't Believe It

By Asif Rehmani
Updated September 14, 2026
The Copilot ROI Baseline
VisualSP
Blog
Your Copilot ROI Story Is Missing the Baseline — And That's Why Your CFO Doesn't Believe It

Nine months into a Copilot rollout, the deck lands in the CFO's inbox. Weekly active usage above 60%. Self-reported time savings of 3.5 hours per user per week. Extrapolated across the licensed population, the number reads north of six million dollars in productivity gains. The CFO reads it twice, sets it down, and asks the one question the deck was never built to answer: compared to what?

That's where most Copilot ROI stories quietly die. Not because the impact isn't real. Because the person telling the story never captured what the "before" actually looked like. And once you're nine months in, you can't go back and measure it.

This is the pillar of the executive Copilot adoption plan that gets the least attention and does the most damage when it's skipped. Governance gets a policy. Training gets a LMS module. Change management gets a comms plan. Baseline capture gets a bullet point that everyone agrees is important and no one actually owns. Then, twelve months later, the CFO wants a number, and the only thing on the table is a self-reported survey the finance team was never going to accept.

The problem with self-reported time savings

Every Copilot ROI deck I see leans on the same source: users telling you how much time they saved. Sometimes it's an in-flow prompt. Sometimes it's a quarterly survey. Sometimes it's a well-run focus group. The number that comes out is directionally interesting and financially useless.

Self-reported savings tell you three things at once, and you can't separate them. It tells you what people believe Copilot did for them, what they want the program to look like (people who championed Copilot report bigger savings — surprising nobody), and what they can remember about a workflow they've now been doing differently for six months. None of that survives an audit. None of that survives a room where someone with a spreadsheet is looking for reasons to cut spend.

The Microsoft-published benchmarks — 14 to 26 minutes saved per user per day, 29% speed-up on certain tasks, 150–400% first-year ROI when governance and adoption are executed properly — are real and useful as an outside-in reference. They are not a substitute for measuring your own environment. When you present a general-industry benchmark as your result, you are telling your CFO you did not measure your own result. That's the moment the credibility drops.

What "baseline" actually means for Copilot

A defensible baseline isn't a single number. It's a small, specific set of pre-Copilot measurements captured against the exact workflows Copilot is meant to change, on the exact populations who will get licensed, in the two to four weeks before rollout begins. The whole discipline is: pick your battles, measure them cold, then measure them again.

Three examples that hold up in a CFO meeting.

Cycle time on a named workflow. Pick a workflow with a clear start and end — first-draft of an RFP response, month-end close narrative, weekly sales report, HR case triage. Measure how long it takes today, in hours, using timestamps you can pull from the system of record. Not a survey. A timestamp. Do this for a defined cohort of ten to fifty people who will be in the pilot. Then measure the same workflow at 30, 60, and 90 days post-Copilot with the same cohort.

Volume per person on a repeatable output. Emails resolved, tickets closed, proposals shipped, meeting summaries produced, drafts submitted for legal review. Whatever your team already counts. Get the pre-Copilot four-week average per person for the pilot cohort and a matched control cohort. Then track the delta.

Quality signal on the same output. Rework rate on the RFP. Manager kickback rate on the draft. Revision count on the deck. Something the business already tracks that says "this got done well." Because if your speed metric goes up 30% and your quality metric drops, that isn't a win. That's a rework problem you now have to explain.

Three metrics. Two cohorts, pilot and control. Two-week capture window before you turn Copilot on. That is the whole exercise. It doesn't require a new tool. It requires someone to own it.

The control cohort is the part everyone skips

The single move that turns your ROI story from "we think Copilot helped" to "we can prove Copilot helped" is a matched control cohort — a group in the same role, doing the same work, who don't get a Copilot license for the first 60 to 90 days. Same measurements, same window, no license. When your pilot cohort's cycle time drops 32% and your control cohort's drops 4% over the same period, you have attribution. Without a control, you have a graph that trended in the direction you were rooting for.

Yes, this is politically hard. Somebody in that role is going to feel left out. That's a change management problem, not a measurement problem, and it's solvable — the control cohort gets Copilot on day 91, and they get told exactly why upfront. What you get in return is a defensible number, which is the only kind of number that keeps the program funded in year two.

What to do this week

If your rollout hasn't started yet, name the three workflows you will measure and the person who owns capturing them. Not "IT" and not "the adoption team." A named human. Write it in the plan. If your rollout is already underway and you skipped this, don't fake it retroactively. Instead, pick the next wave — the next department, the next persona, the next 200 licenses you're about to activate — and run the baseline discipline on them. You won't be able to claim ROI on the users already live, but you'll have a defensible number for everyone from wave two forward, and that is the number the CFO will use to decide whether to renew.

The executive Copilot adoption plan has eight pillars. Vision. Governance. License strategy. Persona enablement. Change management. Training. Measurement. Lifting low-adoption pockets. Six of them decide whether Copilot gets used. The measurement pillar decides whether it gets funded again. Own the baseline before the rollout starts, and the ROI story writes itself. Skip it, and you'll spend year two defending a number nobody in finance believes.

VisualSP accelerates digital adoption, digital transformation & user training.

Get a Demo
Table of Contents

Stop Pissing Off Your Software Users! There's a Better Way...

VisualSP makes in-app guidance simple.