Somewhere along the way, "lakehouse or warehouse?" stopped being an architecture question and turned into a tribal loyalty test. There's a camp that treats the data warehouse as a relic and the lakehouse as the obvious future, and a camp that regards the lakehouse as an overcomplicated way to reinvent things the warehouse settled decades ago. Both camps argue with the fervour of people defending a football club, and both are, in the way of most tribal arguments, mostly missing the point. Because after building on both, and watching organisations succeed and fail with each, I've come to think the lakehouse-versus-warehouse question is almost never the one that decides the outcome. The architecture that fits your organisation is settled by things nobody's shouting about online. So stop asking which tribe to join, and start asking the questions that actually matter.
First, a quick, fair statement of what the two things are, without the war paint. A data warehouse is the mature, structured approach: data modelled into clean tables, strong on SQL, governance, and reliable BI, built on decades of hard-won practice. A lakehouse tries to combine a data lake's flexibility and cheap open storage with warehouse-like structure and performance on top — one place that handles both the messy, varied, large-scale data and the clean analytical layer. Modern platforms increasingly blur the two anyway: Microsoft Fabric, currently in preview, deliberately offers both a warehouse and a lakehouse over the same storage and rather pointedly invites you to stop choosing. Which is a hint that the binary was never as sharp as the arguments make it sound.
The questions that actually decide it
Here's what I ask instead, and these are the ones that predict success:
- What does your data actually look like? If your world is overwhelmingly structured, relational, SQL-shaped data feeding BI, a warehouse (or the warehouse side of a unified platform) is a comfortable, proven fit and you needn't overcomplicate it. If you're genuinely wrangling large volumes of varied, semi-structured, or unstructured data — logs, events, documents, things destined for data science — the lakehouse's flexibility earns its keep. Match the architecture to the data you have, not the data in a vendor's demo.
- What are your people good at? A lakehouse in the hands of a team that lives in SQL and knows warehousing can become an unfamiliar, underused mess. A warehouse imposed on a team of Spark-and-Python data scientists will chafe. The best architecture for your organisation is partly the one your people can actually operate well, and that's a real constraint, not a compromise.
- What's your data-science ambition, honestly? If serious ML is central to where you're heading, the lakehouse's affinity for that work is a genuine advantage. If you mostly need trustworthy dashboards and clean reporting, a warehouse does that superbly and the lakehouse's extra flexibility may be complexity you'll pay for and never use.
- What's the real cost and operational picture? Cheap open storage is a genuine lakehouse draw, but "cheap storage" and "cheap to operate" are different sentences. Factor in the skills, the tooling, the governance maturity each demands in your context, over years, not the sticker price of the storage.
- How mature is your governance, right now? A warehouse's structure imposes a certain discipline almost by default — schemas, types, a modelled shape. A lakehouse hands you flexibility, and flexibility without governance becomes a swamp faster than anything else in data. If your governance practice is still young, the lakehouse's freedom is a liability before it's an asset; the warehouse's guardrails may be exactly the scaffolding you're not yet ready to do without. Choose the amount of freedom your organisation can currently handle safely, not the amount that sounds most modern.
The pattern behind the answers
Notice that not one of those questions is "which is technically superior," because that question has no context-free answer and arguing it is how people waste months. Every question that does decide the outcome is about fit — fit to your data, your people, your ambitions, your constraints. This connects to something I keep coming back to, most directly when I wrote about medallion architecture without the buzzwords: the winning architecture is rarely the most fashionable one; it's the one that matches the organisation it has to serve. Fashion is a terrible architect. Fit is a good one.
Where this is all going anyway
Here's the part that should take some heat out of the whole debate: the industry is quietly resolving it by refusing to choose. The convergence is real — platforms are increasingly offering warehouse and lakehouse experiences over one shared, open storage layer, so you can use the right tool for each workload without maintaining two separate worlds and copying data between them. Fabric is the most visible bet on exactly this. In a few years I suspect "lakehouse or warehouse" will sound as dated as arguing about which single database to standardise the whole company on — a question that assumed a scarcity of options that no longer exists.
The counter-view, so I'm not accused of dodging: convergence doesn't mean the distinctions vanish, and "use both" can be an excuse to avoid the discipline of deciding what each workload actually needs. A unified platform still asks you to choose the right engine for each job; it just spares you from maintaining two estates to do it. So the thinking doesn't go away — it moves from "pick a tribe" to "match each workload to the right tool within one coherent platform," which is a much healthier question. Either way, the move that never pays off is choosing your architecture by which side of an internet argument you find more persuasive. Ask what fits. The data, the people, the ambition, the cost — those decide it, quietly, every time, while the tribes are still shouting.