Data Lake
Data & TrackingAlso: Data Lakehouse
Quick definition
A data lake is a storage system that holds raw data in its original format, structured or not, until someone needs to use it. Unlike a data warehouse, it doesn't force data into tables and columns upfront. It's built for volume and flexibility rather than instant readability.
How it varies across Australia
Data lake adoption in Australian mid-market businesses lags well behind enterprise. Most smaller organisations run on a single data warehouse or a handful of connected tools, and only reach for a data lake once they're combining event data, CRM exports and third-party feeds at real volume.
See data and tracking maturity across Australian industries →What it actually means
Picture a warehouse where every item arrives pre-labelled, shelved and catalogued before anyone touches it. That's a data warehouse. A data lake is the loading dock. Everything gets dropped there in whatever shape it arrived: raw event logs, CRM exports, PDF invoices, clickstream data, social media feeds. Nothing gets forced into a schema until someone actually needs to use it.
This matters because marketing teams generate messy, high-volume data that doesn't always fit neatly into rows and columns. Attribution data, user-level clickstream events, unstructured customer feedback. A data lake lets you keep all of it without deciding upfront how it'll be queried.
The catch is that a data lake with no governance turns into what practitioners call a data swamp. Nobody knows what's in it, nothing is documented, and it becomes cheaper to ignore than to use. The value of a data lake isn't the storage. It's the discipline applied on top of it, usually through a data layer or pipeline that eventually feeds a data warehouse or reporting tool.
Most marketing teams never need a data lake directly. It's typically owned by a data or engineering function and marketing consumes a cleaned, modelled version downstream.
A data lake without governance isn't an asset. It's a swamp with a nicer name.
How it shows up
A data lake shows up as the underlying infrastructure behind more sophisticated reporting setups. It's rarely visible to a marketer directly. Instead it shows up as the answer to 'where does that come from' when someone asks how a cross-channel attribution model or a lifetime value calculation was built. If your organisation has one, it usually lives in a cloud platform like Amazon Web Services (AWS), Google Cloud or Azure, and a data engineering team manages what flows in and out.
The Australian context
Data residency matters more in Australia than in many markets, particularly for businesses in finance, health or government-adjacent sectors. The Privacy Act and sector-specific rules mean where a data lake physically stores information can be a compliance question, not just a technical one. Local businesses often choose Australian regions within AWS, Google Cloud or Azure specifically to keep customer data onshore.
Where people get this wrong
Data Lake vs Data Warehouse
| Data Lake | Data Warehouse | |
|---|---|---|
| Data format | Raw, unstructured or mixed | Structured, modelled into tables |
| Ready to query on arrival? | No, needs processing first | Yes, built for fast queries |
| Cost profile | Cheap storage, expensive to use well | More expensive storage, cheaper to use |
| Typical owner | Data engineering | Analytics or BI team |
| Marketing team interaction | Rare, usually indirect | Common, via dashboards and reports |
Related terms
Common questions
Does a small business need a data lake?
Almost certainly not. Data lakes solve problems at high data volume and variety, typically found in larger organisations. Most small and mid-sized businesses get more value from a well-structured data warehouse and clean tracking than from raw storage infrastructure.
What's the difference between a data lake and a data warehouse?
A data lake stores raw data in its original format and structures it later, if at all. A data warehouse stores data that's already been cleaned and organised into tables built for fast querying. Many organisations use a lake to feed a warehouse rather than choosing one over the other.
What is a data swamp?
A data swamp is a data lake that's lost governance. Nobody documents what's stored, ownership is unclear, and the data becomes too messy or untrustworthy to use. It's the most common failure mode for data lake projects and usually stems from storing data without a plan to use it.
Where do data lakes typically live?
Most run on cloud platforms like Amazon Web Services, Google Cloud or Microsoft Azure. Australian businesses in regulated sectors often choose Australian data centre regions specifically to keep customer data onshore for compliance reasons.
Debrief
Get the next one
No spam. No fluff. Just the next article, straight to your inbox.
Keep exploring
About New Rebellion
New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.
How we think →