Skip to content
Paradox Machines
All posts
By Andre Casimiro6 min readData ArchitectureAI Strategy

How Paradox Machines maps data architecture

Most companies can't answer who owns their data, what it's for, or whether it can be trusted. A four-layer reference architecture is the map we use with customers to help answer those questions, or identify gaps.

Part 1 of a series on our data reference architecture: this post lays out the map; the four that follow take the layers one at a time.

AI is pushing companies to interrogate and leverage their data like never before. The AI models themselves are a commodity anyone can buy; what turns them into an advantage is a company’s own first-party data, and any AI built on it will be only as good as that data. Yet ask three plain questions and most organizations can’t answer any of them confidently:

  1. Who is accountable for the data?
  2. What is it used for?
  3. Can you trust it to make decisions?

These are important questions that when answered properly can shape how an organization works with and benefits from its own data. To help answer these and many other questions, we use a reference architecture to map out everything related to a company’s data: the business needs that require them, the governance that defines who and how the data is managed, the definitions of the data assets needed to support the business and finally the platform that brings it all together. It’s a map for navigating the complexities of your own data. This series is about how we build and use this map in partnering with our customers.

One map, multiple payoffs

A reference architecture pays off in many ways, and for a long period of time:

  • Shared vocabulary: teams can tell whether they’re debating governance or the platform, a capability or a tool, instead of talking past each other.
  • A clear view of your current state: the exercise of mapping what you have against the reference forces a complete inventory: what’s covered, what’s missing and what’s wasteful.
  • Knowing what to build next: after the mapping exercise, the gaps that surface become your roadmap.
  • Priorities rooted in the business: every item traces to a business need in a view you can reason about, so you fund what the business actually needs instead of a capability for its own sake.

AI raises the stakes

All of this can feel like the slow path in an age where an LLM will seemingly answer anything, and the temptation is to hope AI lets you skip the fundamentals. It does the opposite. Point a capable model at messy ownership, wrong data and an unaware platform, and you get confident, fluent answers built on data no one should trust, and an agent that acts without guardrails. AI accelerates how we work, and the faster you go the more control it takes to stay on course.

A reference architecture keeps everything in check, it is a view across the organization. When AI is duplicating efforts or creating schema drift or metric definition collision, it can refer back to the architecture to understand where everything lives and who signs off on what. It provides a guardrail to AI (and humans), enabling teams to move faster.

The four layers

A reference architecture can list dozens of moving parts, but they sort into four layers, and each one is framed as a question you should answer:

  • Business - what data, information, and intelligence does the company actually need, gathered from every department? The needs sort into four: to comply, to operate, to decide, and to automate. This layer sits on top and drives the rest.
  • Governance - how does the data stay trustworthy? Three pillars answer that - people, processes, and technology: who owns the data, who is accountable, and what has to be true for teams to rely on it.
  • Data assets - what data must the company capture, and in what form? Five stages, from the raw operational record through to the intelligence the platform produces, with the reconciled and analytical layers in between.
  • The platform - what does the technology actually have to do? Thirteen capabilities that store, transform, and serve the data, wrapped by the cross-cutting work of running and governing it.

The four layers stack top down, and the influence only runs one way. The business sits at the top and sets what matters: the decisions and outcomes worth investing in. Those needs decide what is worth governing and what data is worth capturing, which in turn shapes the data assets below. The platform sits at the bottom and exists only to store, transform, and serve those assets.

By framing the architecture in this top-down manner, you ensure that technical decisions made at the bottom are tightly coupled to business outcomes at the top.

The architecture is not a checklist

Seeing all four layers laid out can read like a checklist to complete. It is not. The map shows everything a data setup can include, not everything yours should. A complete map does not mean a complete build. A small company selling one product has very different needs from a bank or hospital, and many of the boxes on the map will remain empty; the key is they remain intentionally empty, not because they weren’t considered. The value is not in scoring full marks against the map; it is in choosing the parts that serve what the business actually does and leaving the rest, so effort and resources are diverted to where they are needed.

The architecture is tool agnostic

A critical piece of the reference architecture is that it remains tool agnostic. In any data conversation there is a strong pull to jump straight to products: which warehouse to buy, which BI tool to standardize on. The map holds every capability as a job to be done, kept separate from the specific product that might do it. Keeping the two apart matters, since the same tool can be a great answer to one need and a poor answer to another; whats more, tools evolve faster than the needs above them, so anchoring on a product dates your thinking.

Tools follow structured decision making, not the reverse.

How to read this series

The four posts that follow consider the layers one at a time: business, governance, data assets, and finally the platform. Each stands on its own, but should be read in the same order as the architecture proposes.

This map is how we approach the data platforms and systems we build and deploy at Paradox Machines. I’m excited to dive in deeper in this series, and share the overall framework we use to approach each layer in it.

Next up: the business layer. Stay tuned!

Ready to build your data foundation?

Tell us where your reporting hurts and we'll show you what a reliable foundation looks like.