I've just completed DP-100, the Azure Data Scientist certification, which caps off a deliberate stretch of crossing from my home territory — data engineering — into the neighbouring country of data science. It started a while back with AI-100, the AI Engineer certification, and has finished, for now, with DP-100. And having made the trip, I want to write about what the border between data engineering and data science actually looks like from someone who lives firmly on one side and went to genuinely understand the other — because the two get lumped together constantly, and they are more different than the "data" in both names suggests.

Why an engineer goes exploring next door

First, the honest motivation, because "collecting certifications" isn't it. As I've led a platform team, the work has increasingly touched data science — the models, the ML, the analytical work that sits on top of the platforms I build. And I kept feeling the same discomfort I felt before I sat the analyst certification: that I was building for a discipline I understood only second-hand, making decisions that affected data scientists' work without really knowing what their work was like from the inside.

So this was the same move as the analyst cert, pointed at a different neighbour: go and properly understand the people my platform serves, this time the data scientists, so I can build for them better and know where my engineering helps them and where it quietly gets in their way. AI-100 came first, focused on the AI-engineering side; DP-100, on the data-science and machine-learning side, is where I really felt the difference from my home discipline.

The border, as I found it

Here's what the crossing actually taught me about how the two disciplines differ, beyond the surface fact that both involve data:

  • Engineering seeks the right answer; science seeks a good enough one. My whole training as a data engineer is oriented around correctness — the pipeline is right or it's broken, the number is accurate or it's a bug. Data science lives in a fundamentally more probabilistic world: a model isn't right or wrong, it's better or worse, more or less accurate, useful within confidence bounds. Crossing that border was a genuine mental adjustment — learning to be comfortable with "good enough and improving" where my instinct demands "correct."
  • The rigour is real, but it's a different rigour. I'll admit I half-carried the engineer's quiet prejudice that data science is "just fitting models." It isn't. The discipline of doing it well — validating properly, avoiding the traps of overfitting and leakage, being honest about what a model can and can't claim — is real and demanding rigour, just aimed at different failure modes than engineering rigour. I came away with a lot more respect for the craft of it.
  • The relationship to the data is different. As an engineer, I make data correct and available. As a scientist, you interrogate it for patterns and predictions, and you relate to its messiness differently — noise and uncertainty aren't bugs to eliminate but properties to model. Same data, genuinely different posture toward it.

Understanding those differences from the inside has already changed how I build. I can see, now, where a platform decision that felt neutral to me actually helps or hinders the probabilistic, exploratory, iterative way a data scientist works — and I can build for that rather than assuming they work the way I do.

Am I a data scientist now? No — and that's the point

Let me be clear-eyed about what this trip did and didn't do, because it's the same honesty I brought to the other certs. Passing DP-100 does not make me a data scientist, any more than passing the analyst cert made me an analyst. It made me a data engineer who genuinely understands data science — which is a different and, for my actual job, more useful thing.

Because I'm not trying to become a data scientist. I'm trying to be a better platform leader for data scientists, and for that, the goal was never to do their job but to understand it well enough to build for it, communicate with them without a translation gap, and make good decisions about the analytical work my platforms support. Depth in my own discipline plus genuine literacy in the neighbouring one is exactly the combination that lets me serve the border between them, which is where a lot of the friction and value in data work actually sits.

Crossing into an adjacent discipline isn't about switching careers. It's about erasing the translation gap between you and the people you build for — so you stop making decisions about their work from the outside and start making them with an understanding from the inside.

The pattern, again

I notice this is the third time I've made essentially the same move — going and learning a role adjacent to mine (analyst, now data scientist) not to become it but to serve it better — and I don't think that's coincidence. It's becoming a deliberate practice: as I lead more, and build platforms and teams for an increasingly varied set of people, the way I stay good at it is by periodically going and genuinely inhabiting the perspective of the people downstream of me, so I never drift into building for an imagined version of them.

The certifications are just the structured vehicle for that. What I'm actually accumulating isn't badges; it's a widening literacy across the disciplines my work touches, so that the borders between them — engineering and analysis and science — become places I can translate across rather than walls I build blindly against. From AI-100 to DP-100 was one more crossing in that project. I came back to my own side, as I always do, but I came back understanding the neighbours a great deal better — and that, not the certificate, was the whole point of the trip.