Copilot in Power BI has reached general availability, which means the feature everyone's been demoing for the better part of a year is now something you can actually put in front of real users with real questions. Ask in plain English, get a narrative summary, have it draft a measure or explain a report. The demos are genuinely impressive. And I want to give you the verdict I'd give a client who asked "should we turn this on", which is more useful than either the hype or the backlash.

Here it is: Copilot in Power BI is a mirror. Point it at a clean, well-named, well-governed semantic model and it's genuinely useful. Point it at the sprawling mess most organisations actually have, and it will reflect that mess back at you — fluently, confidently, and wrongly. Whether Copilot is good for you is mostly a question about your model, not about Copilot.

What it does well, honestly

Let me be concrete about where the value is real, because it is real — just narrower than the keynote suggests.

  • Explaining, not just answering. Ask Copilot what a report shows or what a gnarly measure actually calculates, and it's a fast, capable interpreter. For the analyst inheriting someone else's work, or the stakeholder who doesn't speak DAX, that's a genuine unlock.
  • Drafting the first version. "Write me a measure for year-over-year growth" gets you a competent starting point in seconds. It's not always right, but it's almost always faster to correct than to write from scratch. As a fast junior who does the first pass, it earns its keep.
  • Narrative summaries. Turning a chart into a paragraph of "here's what changed and by how much" is the kind of small, repetitive articulation work that Copilot genuinely takes off your plate.

Notice what these have in common: in every one, a human stays in the loop as the checker. The analyst reads the explanation and judges it. Corrects the drafted measure. Sanity-checks the summary against the visual. That's Copilot at its best — accelerating a person who remains responsible for the output.

Where it quietly fails — and why it's the model's fault

The trouble starts at the use everyone actually wants: the business user typing a question and trusting the number that comes back without an analyst in between. That's the dream being sold. And it works right up until the moment your data doesn't cooperate — which is most of the time, for reasons that have nothing to do with the AI.

Copilot answers questions using your semantic model: the tables, the relationships, the measures, and crucially the names of all of those. So when someone asks "what were sales last quarter," Copilot goes looking in your model for something that means "sales." If your model is clean — one clearly-named Total Sales measure, unambiguous relationships, a single table that obviously represents customers — it finds the right thing and gives a right answer. If your model is what most real models are — three columns that could plausibly be "sales", a measure called Sales_v2_FINAL, two customer tables nobody's reconciled — Copilot picks one, and gives you a fluent, confident answer built on the wrong choice. It won't flag the ambiguity. It can't. It doesn't know there was a decision to make.

That's the danger, stated plainly: Copilot doesn't fail loudly with an error. It fails quietly, with a plausible number. And a plausible wrong number handed to someone who trusted it is worse than no number at all, because they'll act on it.

Copilot doesn't know what your data means. It knows what your data is named. In a well-modelled world those are the same thing — and almost nobody lives in a well-modelled world.

The uncomfortable prerequisite

Which leads to the conclusion the demos skip. The organisations that get real, safe value from Copilot in Power BI are the ones that did the unglamorous work first: named their measures like a human should understand them, collapsed their duplicate tables, agreed what the contested words mean, and marked the trustworthy model as the one to use. Copilot didn't remove the need for a well-governed semantic layer. It raised the stakes on having one, because now a messy model doesn't just confuse the analyst who knows to be careful — it misleads the executive who doesn't.

I made a prediction at the start of this year that the semantic model would become the real battleground of 2024, and Copilot's arrival is exactly why. The moment you let people query in natural language and trust the result, the definitions underneath stop being an internal modelling concern and become the thing standing between your leadership and a confidently wrong decision. Copilot turns your data model into a public-facing product whether you were ready for that or not.

So should you turn it on?

Yes — with a clear-eyed sense of which Copilot you're turning on. Enable it as an accelerator for the people who can check its work: your analysts and report builders, who'll use it to draft, explain, and summarise faster, and who have the judgement to catch it when it's wrong. That's low-risk, high-return, available today.

Be far more careful about promoting it as a self-service oracle for business users querying models you know aren't clean. Not because the feature is bad, but because in that configuration it will confidently launder your data-quality problems into executive decisions, and you won't find out until one of those decisions goes wrong.

The through-line, as ever with this stuff: the AI got impressive, and in doing so it made the boring foundational work matter more, not less. Copilot is a fast, fluent, useful reflection of your semantic model. If you don't love what it reflects, the fix was never going to be a better mirror. It's a tidier room.