Data Lake

Data & Tracking

Also: Data Lakehouse

What it isRaw storage for all your data
Versus warehouseUnstructured, not modelled
Watch forBecoming a data swamp
FeedsAttribution, segmentation, reporting

Quick definition

A data lake is a storage system that holds raw data in its original format, structured or not, until someone needs to use it. Unlike a data warehouse, it doesn't force data into tables and columns upfront. It's built for volume and flexibility rather than instant readability.

How it varies across Australia

Data lake adoption in Australian mid-market businesses lags well behind enterprise. Most smaller organisations run on a single data warehouse or a handful of connected tools, and only reach for a data lake once they're combining event data, CRM exports and third-party feeds at real volume.

See data and tracking maturity across Australian industries

What it actually means

Picture a warehouse where every item arrives pre-labelled, shelved and catalogued before anyone touches it. That's a data warehouse. A data lake is the loading dock. Everything gets dropped there in whatever shape it arrived: raw event logs, CRM exports, PDF invoices, clickstream data, social media feeds. Nothing gets forced into a schema until someone actually needs to use it.

This matters because marketing teams generate messy, high-volume data that doesn't always fit neatly into rows and columns. Attribution data, user-level clickstream events, unstructured customer feedback. A data lake lets you keep all of it without deciding upfront how it'll be queried.

The catch is that a data lake with no governance turns into what practitioners call a data swamp. Nobody knows what's in it, nothing is documented, and it becomes cheaper to ignore than to use. The value of a data lake isn't the storage. It's the discipline applied on top of it, usually through a data layer or pipeline that eventually feeds a data warehouse or reporting tool.

Most marketing teams never need a data lake directly. It's typically owned by a data or engineering function and marketing consumes a cleaned, modelled version downstream.

A data lake without governance isn't an asset. It's a swamp with a nicer name.

How it shows up

A data lake shows up as the underlying infrastructure behind more sophisticated reporting setups. It's rarely visible to a marketer directly. Instead it shows up as the answer to 'where does that come from' when someone asks how a cross-channel attribution model or a lifetime value calculation was built. If your organisation has one, it usually lives in a cloud platform like Amazon Web Services (AWS), Google Cloud or Azure, and a data engineering team manages what flows in and out.

The Australian context

Data residency matters more in Australia than in many markets, particularly for businesses in finance, health or government-adjacent sectors. The Privacy Act and sector-specific rules mean where a data lake physically stores information can be a compliance question, not just a technical one. Local businesses often choose Australian regions within AWS, Google Cloud or Azure specifically to keep customer data onshore.

Where people get this wrong

Building a data lake before there's a clear use case.Storage without a consumption plan just accumulates cost and risk. Most teams should start with the report or model they need and work backwards, not build infrastructure speculatively.
Assuming a data lake replaces a data warehouse.They solve different problems. A lake stores raw, flexible data. A warehouse structures it for fast, reliable querying. Most mature setups use both, with the lake feeding the warehouse.
Letting a data lake go ungoverned.Without documentation, ownership and quality checks, a data lake becomes unusable clutter. The data swamp problem is the single biggest reason data lake projects get abandoned.

Data Lake vs Data Warehouse

Data LakeData Warehouse
Data formatRaw, unstructured or mixedStructured, modelled into tables
Ready to query on arrival?No, needs processing firstYes, built for fast queries
Cost profileCheap storage, expensive to use wellMore expensive storage, cheaper to use
Typical ownerData engineeringAnalytics or BI team
Marketing team interactionRare, usually indirectCommon, via dashboards and reports

Related terms

Common questions

Does a small business need a data lake?

Almost certainly not. Data lakes solve problems at high data volume and variety, typically found in larger organisations. Most small and mid-sized businesses get more value from a well-structured data warehouse and clean tracking than from raw storage infrastructure.

What's the difference between a data lake and a data warehouse?

A data lake stores raw data in its original format and structures it later, if at all. A data warehouse stores data that's already been cleaned and organised into tables built for fast querying. Many organisations use a lake to feed a warehouse rather than choosing one over the other.

What is a data swamp?

A data swamp is a data lake that's lost governance. Nobody documents what's stored, ownership is unclear, and the data becomes too messy or untrustworthy to use. It's the most common failure mode for data lake projects and usually stems from storing data without a plan to use it.

Where do data lakes typically live?

Most run on cloud platforms like Amazon Web Services, Google Cloud or Microsoft Azure. Australian businesses in regulated sectors often choose Australian data centre regions specifically to keep customer data onshore for compliance reasons.

Debrief

Get the next one

No spam. No fluff. Just the next article, straight to your inbox.

Keep exploring

About New Rebellion

New Rebellion is a marketing intelligence consultancy. We build tools, score Australian businesses on how their marketing actually performs, and publish Debrief every day. This dictionary is part of how we work in the open.

How we think →