top of page

The Great Customer Data Boomerang

1 day ago
11 min read
We spent years centralizing customer data—only to keep sending copies of it back out. The next era of audience activation brings the decision closer to the data.
We spent years centralizing customer data—only to keep sending copies of it back out. The next era of audience activation brings the decision closer to the data.

We spent a decade centralizing customer data. Now we need to stop copying it back out again.


This Signal & Noise exclusive is brought to you by AUDIENCES.


We Centralized the Data—then Started Exporting It


For the better part of the last decade, enterprises have been trying to solve one of marketing’s most persistent problems: customer data was everywhere. Purchase history lived in one system. Website behavior lived somewhere else. Loyalty data had its own database. Mobile activity, service interactions, media exposure, product usage and identity data were scattered across dozens—sometimes hundreds—of platforms.


So companies did what seemed logical: they brought it together. To do this, they invested heavily in cloud data platforms, data lakes, warehouses, CDPs and identity infrastructure. Snowflake, Databricks, BigQuery and other platforms increasingly became the place where enterprises could finally assemble a more complete, governed view of the customer.


Then right on cue, marketing started shipping the data back out again. To activate audiences, brands routinely extract customer data from their central environment and send it to advertising platforms, social networks, DSPs, retail media networks, clean rooms, measurement providers and marketing technology vendors. Each destination gets its own slice of the customer. Each has its own integration, refresh schedule, identity methodology, permissions and governance requirements.

We centralized customer data—and then recreated fragmentation downstream.

The result: a new version of the problem we just spent years trying to solve. Sure, we centralized customer data. But then we recreated fragmentation downstream and outside the enterprise. This model may have made sense when marketing platforms were largely self-contained systems. For example, if a DSP needed an audience, you sent it an audience file. If a social platform needed customer data for matching, you uploaded customer records. If a marketing application needed a segment, you synchronized the segment into the application.


But the underlying architecture of the enterprise has changed—probably forever. Increasingly, the richest and most authoritative customer intelligence already lives inside the company’s own data environment. We’re talking about identity, transactions, behavioral signals, consent, customer value, propensity models and suppression rules—all of which can increasingly be computed there.


So why should the default answer still be to move all that data somewhere else? If you ask me, the better question is what does the activation platform actually need to know? Usually, it doesn’t need the actual customer record. Often, what it really needs is a signal. Is this person a customer, and did they purchase recently? If so, are they eligible for this offer, or should they be suppressed? Are they high value, or did their behavior change? Should the brand bid more—or not bid at all?


You'll notice these are very different things from handing another platform a copy of the underlying customer data, and this critical distinction must reshape how first-party data gets activated.

The next generation of audience activation will be less about moving customer records from system to system and more about making governed customer intelligence available wherever a decision needs to be made. In other words, bring the activation closer to the data, instead of continually sending the data to the activation.


We finally brought customer data together—then immediately started sending copies of it back out across the marketing ecosystem.
We finally brought customer data together—then immediately started sending copies of it back out across the marketing ecosystem.

The Audience Should Live Where the Data Lives


For years, marketers treated the audience as something that lived inside the marketing stack.

You collected customer data somewhere, pushed it into a CDP or audience management platform, created a segment, matched those records against whatever identity framework the destination supported, and then shipped the resulting audience somewhere else for activation.


This approach made sense because the marketing technology stack was where most of the intelligence lived. I would argue this is no longer true. Today, many enterprises already have an environment where customer data is being stored, joined, modeled, governed, and analyzed at enormous scale. And increasingly, that environment is a modern cloud data platform like Snowflake, Databricks, BigQuery or Microsoft Fabric.


This reality changes the role of the marketing stack. If the data platform already knows who bought yesterday, who hasn't purchased in six months, who is likely to churn, who is eligible for an offer, who opted out, who has a high predicted lifetime value, and who just crossed a behavioral threshold, why do we need to reproduce all of that logic someplace else before marketing can act on it? Pro tip: we shouldn't.


The audience itself can increasingly be defined where the underlying data already lives. Let's break down why by starting with what an audience really is. No, it isn't just a CSV file containing 2.3 million customer IDs—that's simply one way we have historically transported it.


I would argue an audience is actually the result of applying customer logic—let me explain what I mean. Customer logic is how you discern between customers who purchased Product A but not Product B, those whose lifetime value exceeds a certain threshold, or those who visited the website three times in the last seven days but haven't converted. Customer logic can be a way to identify existing customers who should be excluded from an acquisition campaign, or pinpoint lapsed customers whose recent activity suggests they may be ready to return.

The activation platform usually doesn’t need the customer record. It needs the right signal at the right moment.

Once you begin to think about audiences this way, the architectural question changes dramatically.

Instead of asking, “How do I get all of this customer data into my marketing platform?” we can ask, “How do I allow the marketing platform to act on the output of this logic?” This is a much smaller and far more manageable problem.


It's also a much better one. For one thing, using this pattern the logic stays closer to the enterprise systems that actually understand the customer. The marketer isn't depending on a stale copy of the data exported yesterday—or last week—to determine what should happen today. Governance also gets easier. Consent, suppression, eligibility, and privacy rules don't have to be recreated independently across every destination, so they can increasingly be applied upstream, before a marketing system ever receives something it shouldn't.


And perhaps most importantly, marketers can stop creating competing versions of the customer.

This is one of the dirty little secrets of modern marketing technology. Every platform promises to help create a better customer view, but the result can be five, 10 or 20 slightly different customer views spread across an organization. The CDP has one, and the CRM another. The email platform has another, and the DSP yet another. Your retail media partners also have their own versions—and everyone spends an extraordinary amount of time reconciling why the numbers don't match.


Fortunately, the modern data environment offers a different and better model: keep the richest customer context in one governed place, build the audience logic there, and expose only what downstream systems need in order to take action. No, this doesn't mean marketing platforms disappear. Far from it. DSPs still need to buy media, social platforms still need to reach consumers, and measurement systems still need to measure outcomes. Creative systems still need to determine what message gets shown.


But their roles change and they become execution environments rather than alternate systems of customer truth. This distinction matters because the more customer intelligence we generate—from transactions, behavioral events, machine-learning models and eventually autonomous agents—the less practical it becomes to continuously copy every new piece of information into every platform that might someday need it.


The better architecture is the inverse. Keep the intelligence close to the data, move only what is necessary, and let the audience travel as logic and signals rather than as another sprawling copy of the customer.


The audience doesn’t need to live in every downstream platform. Build the logic where the richest customer data already exists, then send only the signals each activation system needs.
The audience doesn’t need to live in every downstream platform. Build the logic where the richest customer data already exists, then send only the signals each activation system needs.

Stop Thinking About Records, and Start Thinking About Signals


If audiences can be defined where the customer data already lives, the next question becomes fairly obvious: what actually needs to leave? Probably a lot less than you think.


For decades, marketing technology has been built around moving records. We upload email addresses, phone numbers, hashed identifiers, customer IDs, device IDs, and dozens of associated attributes because the receiving platform needs enough information to figure out who the customer is and what to do with them. In many cases, we have treated the customer record itself as the unit of activation.


I think this approach is starting to look woefully out-of-date. Most advertising systems don't actually need to know everything the enterprise knows about a customer. They don't need their complete purchase history, every interaction with the brand, lifetime value calculation, loyalty status, service history, and dozens of behavioral attributes. All they need is enough information to make a decision.


This is where signals come in. In this case, I define a signal as the output of intelligence that exists upstream. The signal might, for example, tell a media platform that this person is an existing customer and they purchased in the past 30 days. Or it could inform the platform which customer belongs to a high-value segment, in which case they should be suppressed from a win-back campaign. Sometimes, it could be as simple as yes or no—eligible or ineligible, high propensity or low propensity, acquire, retain or suppress.

Traditional audience activation has largely been membership-based: determine which bucket someone belongs in and send the bucket somewhere. Signal-based activation can become decision-based: determine what should happen next and make that decision available wherever activation occurs.

The complexity lives behind the signal, not necessarily inside the system consuming it. Think about a customer with hundreds of attributes in an enterprise data environment. Maybe the company knows what products they own, when they last purchased, how frequently they interact with the brand, what channels they prefer, their predicted lifetime value, their likelihood to churn, and what offers they're currently eligible to receive.


Sorry, but a DSP doesn't need all of that to do its job. It may simply need enough to decide, for example, whether it should I bid on a particular impression. Sure, this decision could be informed by an incredibly sophisticated combination of first-party data, predictive models, business rules, and real-time behavior, but it doesn't mean the platform needs to be exposing all of the underlying information used to reach the decision.


This is an important distinction because we've historically confused having access to data with needing possession of the data. News alert: they're not the same thing. You can benefit from customer intelligence without necessarily creating another copy of the customer record and shipping it somewhere. And once you start thinking this way, a lot of activation use cases begin to look different.


Take suppression, for example. Brands routinely send customer lists to media platforms simply to tell those platforms who not to advertise to. That's a strangely data-intensive way to communicate what is ultimately a binary instruction: do not spend money trying to acquire this person. Or consider propensity modeling. An enterprise may use hundreds of variables to determine that a customer has a high likelihood of purchasing a particular product. Does the activation platform need all of those variables?

Probably not.


What it does need is the resulting propensity signal. The same thing applies to eligibility, churn risk, lifetime value, product affinity, and dozens of other marketing decisions. The sophisticated work can happen close to the underlying customer data. What travels downstream is the information required to act.


This matters for privacy and governance, but it also matters for speed, because customer records are surprisingly heavy and expensive to move, synchronize, and manage. They need to be extracted, transformed, matched, synchronized, and refreshed. Every time identity changes, consent shifts, or a transaction takes place, customers move into and out of segments. Yesterday's audience is not necessarily today's audience, and today's audience may be wrong by tomorrow morning. The more data we replicate across platforms, the more synchronization becomes its own problem.


Signals can be much more dynamic. A customer doesn't have to permanently “belong” to an audience, meaning they can qualify for a particular action at a particular moment based on whatever the enterprise knows right now. I realize this may sound like a subtle distinction, but it's a profound one.

Traditional audience activation has largely been membership-based: determine which bucket someone belongs in and send the bucket somewhere. Signal-based activation can become decision-based: determine what should happen next and make that decision available wherever activation occurs.


That gets us closer to the way modern marketing actually needs to work. Customers aren't static. Their value changes, intent changes, and eligibility changes. Their relationships with the brand shift over time, and increasingly, those changes can be detected almost immediately inside the enterprise data environment. The activation layer should be able to respond just as quickly.


This is where the conversation begins to move beyond better audience management and toward something much more interesting. Because once activation systems can consume governed signals instead of relying exclusively on replicated customer records and static audience files, they become capable of making decisions much closer to real time. And once those decisions can be made dynamically, the economics of the entire activation stack begin to change.


Stop moving entire customer records when the activation system only needs the signal. Less data movement means less complexity—and faster, smarter decisions.
Stop moving entire customer records when the activation system only needs the signal. Less data movement means less complexity—and faster, smarter decisions.

The Payoff: Faster, Cheaper, Smarter Activation


There is also a very practical reason to rethink this architecture: moving less customer data can make the entire activation ecosystem simpler and more efficient. Think about it. Every additional copy of customer data creates work. It has to be moved, matched, refreshed, governed, reconciled, and eventually deleted. Anyone who has worked in enterprise IT can tell you that integrations break, identity match rates fluctuate, consent changes, and different platforms refresh on different schedules. Before long, teams are spending as much time maintaining the plumbing as they are improving the marketing.


Keeping more of the intelligence in the enterprise data environment changes the economics. Fewer unnecessary copies mean fewer pipelines to maintain, less reconciliation between systems, faster access to new customer intelligence, and tighter control over how data is used. More importantly, it shortens the distance between what the company knows and what marketing can actually do with it.


This becomes especially important as marketing becomes more automated. An intelligent agent deciding whether to bid, suppress, personalize, increase spend, or change a message or treatment cannot rely on an audience file created three days ago. It needs current context—what happened, what the customer is eligible for, what the brand is trying to accomplish and what action is permitted right now. This doesn’t require shipping the entire customer record into every system an agent might touch. In fact, doing so would make the problem worse.


A better model is to keep the customer intelligence governed at the source and make the right signal available at the moment a decision needs to be made. That is ultimately where I think audience activation is headed. Not toward more copies of customer data. Instead, toward a future with fewer copies, better signals, and a much shorter distance between intelligence and action.


Less data movement means less integration work, lower costs, and faster activation—while keeping the intelligence needed for smarter decisions closer to the source.
Less data movement means less integration work, lower costs, and faster activation—while keeping the intelligence needed for smarter decisions closer to the source.

Stop Sending Your Customer Everywhere


We spent the last decade solving customer data fragmentation by bringing more of that data together. It would be a strange outcome if the next decade were spent recreating the same fragmentation by copying that customer into every platform that wants to use it. This needs to stop.


The better model doesn't eliminate activation platforms or stop data from moving altogether. Instead, it acts much more deliberately—and judiciously—about what actually needs to move. The best approach keeps the richest customer intelligence governed close to the source, applies the logic there, and gives downstream systems only the signals they need to act.


The future of audience activation should not be about getting more customer data into more places. Instead, we need to focus on how to make better decisions with less movement of the underlying data. Maybe the smartest thing we can do with customer data is finally stop sending so much of it everywhere.


------------------------------------------------


Rio is an executive with 20+ years at the intersection of strategy consulting, AdTech, data, and media. He's a trusted advisor on customer experience, digital strategy, and marketing transformation. He's a partner at Credera, Omnicom's consulting arm. He's also a podcast host, writer, and public speaker focused on the future of advertising and AI-driven infrastructure.




Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page