Most people who work seriously with data arrived from code, or maths, or statistics — the quantitative doors. I came in through a different one. My training is in communication: how messages are made, how they travel, how they land or fail to land in another person's head. And now, having just finished a long piece of research turning thousands of people's words into data, I find myself standing at the threshold of a field I entered sideways, carrying a conviction I want to write down properly, because it's becoming the thing I believe most about this whole subject.

Here it is, plainly: data is, overwhelmingly, people communicating. And if you read it as mere numbers, you miss most of what it's actually saying.

This isn't a mystical claim. It's a practical one, and it changes how you handle data at every step. Let me make the case.

A dataset is not found. It's said.

Start with where data comes from, because that's where the misreading begins. We talk about data as though it were found — discovered, neutral, lying around in the world like rocks waiting to be picked up. That framing is comforting and it is wrong. Almost all the data you'll ever work with was made, by people, through acts that are fundamentally communicative.

Every row in a customer database is a person, at some point, filling in a form — deciding what to type, what to leave blank, what to round off, what to fudge. Every entry in a system of record is someone making a choice about how to record something, under time pressure, according to rules they may or may not have understood. Every gap in a table is a message about what someone didn't think was worth capturing. Every tweet in a dataset of tweets is a human being saying something, to someone, for a reason, in a context.

I felt this in my hands during my research, coding article after article, tweet after tweet, by my own judgement. You cannot build a dataset by hand — deciding personally what counts as what, thousands of times — and ever again believe that data simply exists. You feel, physically, that every data point was said by someone, or decided by someone. The dataset is a transcript of human choices, not a photograph of neutral reality.

Once you've felt that, "read the data as numbers" starts to sound like "read the letter as ink." Technically the letter is ink. But if that's all you see, you've missed the entire point of it.

The hard part was never the counting

Here's what follows, and it's the part I think the purely quantitative doors into this field can obscure. If data is communication, then the hard part of working with it is not the counting. It's the meaning.

Consider the most ordinary disaster in any organisation: two departments show up with different numbers for the same thing — different revenue, different customer counts, different definitions of "active." Everyone reaches for the technical explanation: a data quality problem, a pipeline bug, a mistake to be fixed. But usually it isn't. Usually both numbers are correctly computed. The two departments simply mean different things by the same word, and their data faithfully reflects two different meanings. That's not a counting error. It's a communication failure — a shared word with unshared meaning — wearing the costume of a technical one.

No amount of better tooling fixes that, because the tool isn't where the problem lives. You can build a flawless system on top of a definition nobody agreed on, and all you've done is industrialise a misunderstanding — compute the wrong thing faster and more reliably. The semantic layer, the layer where you decide what the data actually means, is where the real work is. And it's not a technical layer. It's a communication layer, and often a political one, because agreeing what a word means is a negotiation between people who currently, quietly, disagree.

This is why I think a communication background is not a quirky detour into data work but a genuine qualification for it. The field has plenty of people who can build the machinery. It has far fewer who instinctively ask the question the machinery can't answer: what does this actually mean, to whom, and do we agree?

Reading data as a message, in practice

So what does it actually change, to read data as a message rather than a number? Concretely, it changes the questions you ask at every stage.

  • When you receive a dataset, you don't ask only "is it clean?" You ask "who made this, through what act, and what were they really recording?" Because the answer tells you what the data can and can't honestly say. A "customer status" field means something very different if it was set by a salesperson optimising their commission than if it was set by an automated process — and only the communicative question surfaces that.
  • When two numbers disagree, your first instinct isn't "find the bug." It's "do these two things even mean the same thing?" More often than not, the disagreement is semantic, and chasing it as a technical fault is a way to waste a week.
  • When you build something on top of data, you treat the definitions as the foundation, not an afterthought — because you know the whole structure inherits whatever ambiguity you left unresolved at the bottom.
  • When you present a finding, you remember there's a human on the other end who will receive it through their own context and assumptions — and that a true number, badly communicated, changes no minds and no decisions. The analysis isn't finished when the number is right. It's finished when the meaning has actually crossed into another person's understanding.

That last point is where my two worlds fuse completely. A finding that stays in the analyst's head, or arrives as a number the decision-maker doesn't truly grasp, has done nothing. The final, essential step of data work is communication — getting the meaning to land in someone who'll act on it. Treat that as an afterthought and you've built a beautiful bridge to nowhere.

The belief I'm carrying forward

I'm at a beginning here. I can feel the field I'm moving into — the pipelines, the tools, the technical depth I don't yet have and mean to go and get. I have no illusion that reading data "as a message" excuses me from learning the machinery; it doesn't, and I intend to learn it properly. But I also don't intend to let the machinery talk me out of the thing I already know to be true, the way it seems to for some who only ever came in through the quantitative door.

Because here's my bet on where this all goes. As tools get more powerful — as it gets easier and easier to compute over data, to generate numbers and even language from it at scale — the technical barrier to working with data will keep falling. What won't fall, what will if anything become more valuable, is the human judgement about what the data means: whether the definitions are sound, whether the number answers the question actually asked, whether the meaning survives the trip to the person who has to decide. The more the counting gets automated, the more the meaning becomes the whole game.

Data is people communicating. Read it as numbers and you'll be precise about things that don't matter. Read it as a message and you'll start asking the only questions that do.

So that's the conviction I'm building on, entering this field from the side door. Not that the numbers don't matter — they do, and I'm going to go and get properly good at them. But that underneath every number is a message someone sent, a decision someone made, a meaning someone intended — and that the people who can read that, who never forget there's a human at both ends of the data, will always be the ones who matter most. I read data like a message because that's what it is. Everything I do from here is going to be built on that.