Introduction
Not every data team operates at the same level of maturity. Some teams are basically in a rocket 🚀: distributed compute, automated pipelines, streaming, scalable storage. Others are still riding a bike 🚲 through the countryside with Excel, a few SQL queries, and a PowerPoint deck.
Both are moving. Both are doing Data. But they are clearly not playing the same game.
That gap is what this article is about. Data teams can operate at very different levels of tooling maturity, from mostly manual workflows to fully automated, event-driven platforms. The point is not that everyone should immediately build the most advanced stack. It is to understand where your team stands, what its current tooling allows, and what the next meaningful step looks like.
So where does your data team actually stand? That is what these four maturity levels are meant to show, from basic manual workflows to scalable, automated Data platforms.
🔴 Level 0 – Unstructured Data Management
Data? Oh yeah, we’ve got spreadsheets for that.
At this stage, data processes are completely ad hoc. There is no structured data management, and most tasks are handled manually. Teams do not have a dedicated data infrastructure, and reporting usually relies on Excel files, CSV exports and PowerPoint slides.
Key Characteristics:
- Ingestion: No dedicated ingestion tools. Data is manually copied, downloaded or pasted into files.
- Compute: Processing takes place on local machines, mostly through desktop tools.
- Storage: Scattered files on shared drives, personal folders or local machines.
- Analytics: Mostly Excel, pivot tables and simple dashboards built from files.
- Machine Learning / AI: Little to no ML activity. AI may be used through standalone tools, but it is not connected to the company’s Data stack.
- Automation: None or almost none. Reports and datasets are updated manually.
- Processing capacity: Small, ad hoc datasets.
Challenges:
- No scalability: Workflows quickly become difficult to maintain as data volumes grow.
- Error-prone: Manual handling creates inconsistencies, duplicated files and broken formulas.
🟡 Level 1 – Structured Data Interaction
We know how to query using SQL and use some data tools, but our stack is still fragmented.
At this level, your team starts querying structured databases but still lacks centralization. BI and analytics depend on local SQL queries, Excel workbooks and isolated data sources. While some automation exists, data processes remain largely manual and fragmented.
Key Characteristics:
- Ingestion: Manual SQL scripts or low-code tools like Alteryx or Dataiku.
- Compute: Processing mainly takes place on local machines or individual environments (Excel, Python, Jupyter Notebooks, etc.).
- Storage: No dedicated analytical storage. Data is accessed directly from operational systems, files or other source databases.
- Analytics: BI tools and notebooks are used directly on source data.
- Machine Learning / AI: Data Science and ML experiments may exist, usually in notebooks or isolated environments, with little production integration.
- Automation: Basic automation exists but workflows are not integrated into reliable pipelines.
- Processing capacity: Mostly ad hoc or scheduled batch processing.
Challenges:
- Lack of standardization: Your team may create its own queries without company-wide coordination, leading to duplicated logic and redundant work.
- Scalability issues: Queries on live production databases can cause performance bottlenecks. Never run a SELECT * on a production database! (see article Analytical vs Transactional)
🟢 Level 2 – Centralized Data Platform
We have a centralized data platform and reliable pipelines, but we still mostly work in batch mode.
At this level, data is centralized into a dedicated analytical platform instead of being queried directly from operational systems. Pipelines ingest and transform data on a schedule, giving analytics teams a shared and more reliable source of truth.
The stack is now industrialized and automated, but still mostly batch-oriented.
Key Characteristics:
- Ingestion: Automated ingestion, transformation and scheduling, using platform-native capabilities or tools (e.g. Airbyte, Fivetran, dbt, Airflow).
- Compute: Processing runs on dedicated infrastructure, on-prem or in the cloud.
- Storage: Centralized analytical storage such as a Data Warehouse, Data Lake or Lakehouse.
- Analytics: BI workloads use shared datasets and centrally managed logic rather than querying raw operational sources.
- Machine Learning / AI: ML workloads can use the centralized Data platform for training and batch inference.
- Automation: Data pipelines run automatically, usually on predefined schedules.
- Processing capacity: Primarily batch processing for structured and semi-structured data.
Challenges:
- Batch latency: Data freshness depends on pipeline schedules, so new events are not immediately available for analysis.
- Governance gaps: As more pipelines get added, ownership and documentation often lag behind, leading to duplicated transformation logic and eroding trust in the “single source of truth.”
🔵 Level 3 – Advanced Data Architectures
We process both batch and real-time data. Our infrastructure is scalable and AI-ready.
Here, the data team moves beyond mostly batch-oriented analytics and starts handling real-time events, streaming workloads and more complex data processing. Compute becomes more elastic and distributed, while pipelines can react to events instead of only running on a schedule.
Machine learning can also be integrated into production workflows, but the real shift is architectural: the platform can now handle data at much higher volume and velocity, across more varied workloads.
Key Characteristics:
- Ingestion: Batch and event-driven pipelines, using tools like Kafka, Flink or platform-native services.
- Compute: Scalable, on-demand and distributed compute (Spark clusters, streaming engines, serverless compute, etc.).
- Storage: Centralized analytical storage able to handle all types of data: structured, semi-structured and unstructured data.
- Analytics: BI and advanced analytics can consume both historical and continuously arriving data.
- Machine Learning / AI: ML models, including LLM-based and agentic systems, are integrated into production workflows with automated inference and retraining.
- Automation: Pipelines can react to events and trigger downstream processing.
- Processing capacity: Supports both batch and streaming workloads.
Challenges:
- Higher complexity: Streaming, distributed compute and event-driven pipelines require stronger engineering skills.
- Operational overhead: More moving parts also mean more monitoring, observability and failure handling.
Key Takeaways
You know I love a good cheatsheet, so here’s the whole article in one view:

The goal is not to reach Level 3 at all costs. A mature data team is one whose tooling matches its actual needs. If batch pipelines, a centralized platform and standard analytics already meet your needs, Level 2 may already be exactly where you should be.
What matters is knowing where you are, what is limiting you, and whether moving to the next level solves a real problem.
👉 If tooling maturity shows you how the game is prepared, this next article is about the cards and the players. Curious to see the deck? Roles in the Data/AI World: Poker Cards 🎴