I've written about Copilot piece by piece for a year and a half now — its arrival, its DAX skills, its move to mobile, the day it got switched on by default. This is the post that steps back and draws the whole map, because enough time has passed to say, with some confidence rather than speculation, what's genuinely load-bearing, what's a pleasant convenience, and what's still theatre. If you only read one thing I write about Copilot, I'd want it to be this, because the individual features matter far less than the single principle that sorts them.
Here's that principle first, because it does almost all the work: the usefulness of any Copilot feature is decided by one question — is its output checkable by the person receiving it? Everything else follows from that. Let me use it to sort the whole landscape into three honest tiers.
Tier one: genuinely load-bearing
These are the uses I'd actually miss if they vanished — where Copilot earns its place because a human stays positioned to verify what it produces.
- Drafting DAX and other code. Ask for a measure, get a competent first draft, read it, test it, ship it once you understand it. The output is code you can inspect, so its errors are visible and correctable. This is the single most valuable thing Copilot does, precisely because it's the most checkable.
- Explaining existing work. Point it at an inherited report or a wall of someone else's DAX and get a plain-language explanation. It turns "I'm afraid to touch this" into "I understand this," and you can validate its explanation against the artefact in front of you. Quietly transformative for anyone maintaining what they didn't build.
- Summarising and articulating. Turning a chart into a first-draft paragraph of "here's what changed" — with a human editing before it goes anywhere — takes real, repetitive articulation work off the plate. Useful, and safe, because you're right there to catch it.
The common thread: in every one, the person using Copilot can see whether it's right before anything depends on the answer. That's the whole game.
Tier two: useful convenience, handle with care
These are real conveniences with a catch — genuinely helpful, but riding closer to the line where checkability gets harder.
- Natural-language querying by capable users. An analyst asking a question in plain English of a model they understand can move fast and sanity-check the result against what they know. Useful. The same feature in the hands of someone who can't judge whether the answer is plausible slides straight into tier three.
- Pushed summaries to mobile and inboxes. Convenient for busy decision-makers, and I've argued it's the riskiest convenience, because it delivers interpretation to exactly the people least positioned to verify it. Worth having for well-governed models you'd stake your reputation on; worth real caution everywhere else.
Tier two isn't hype — these things work. They just require you to hold the guardrail the convenience quietly removes: keep the path back to the source close, and be honest about who's receiving the output and whether they can judge it.
Tier three: still theatre
And then the tier that gets the applause and delivers the least.
- "Ask your data anything and trust the answer." The headline demo. On a curated demo model it's magic; on a real, messy semantic model it's confident nonsense, because Copilot answers from what your data is named, not what it means, and almost nobody's model is clean enough for those to be the same. This is the pitch that peaked the whole "AI replaces the analyst" narrative, and it's the one that keeps not surviving contact with a real estate.
- The end-to-end "build me a finished report from a sentence" moment. Spectacular on stage, marginal in practice, for the same reason: it collapses the instant the underlying model isn't pristine, which is always.
Tier three isn't useless forever, and I'm not sneering at the ambition. But as of today it's a promise, not a deliverable, and treating it as a deliverable is how organisations wire confident wrong answers straight to decisions.
The Copilot features that survived eighteen months are the checkable ones. The ones still stuck in the demo are the ones that ask you to trust an answer you can't inspect. That's not a coincidence — it's the whole pattern.
What actually determines your mileage
Here's the part that matters more than the tiers, and it's the same conclusion I keep arriving at from every direction: which tier a given feature lands in, for you, depends mostly on your semantic model, not on Copilot. A clean, well-named, governed model pulls features up the tiers — natural-language querying becomes genuinely safe, summaries become trustworthy, because the thing Copilot reads from is sound. A messy model drags everything down — even the good stuff gets riskier, because the foundation it's working from is ambiguous.
So the highest-leverage thing you can do to get value from Copilot isn't a Copilot activity at all. It's the unglamorous work of getting your semantic model clean, named like a human should understand it, and governed. Copilot didn't reduce the value of that work. It repriced it upward, because now your model quality determines not just your report quality but whether your AI helps or misleads.
The honest bottom line
Copilot in Power BI and Fabric is genuinely useful, genuinely over-sold, and the gap between those two is entirely predictable once you apply the checkable test. Adopt tier one enthusiastically — it's real, it's safe, it'll make your capable people faster today. Use tier two with the guardrails intact. Treat tier three as an impressive demo you don't yet bet decisions on. And spend your actual energy where it compounds: on the model underneath, which decides how much of Copilot's promise you get to safely keep. The hype will keep insisting the AI is the point. A year and a half in, I'm more sure than ever that the model was always the point, and Copilot is just the thing that made caring about it urgent.