Episode 11: Why dirty customer data is killing growth
Customer advocacy and data expert Irwin Hipsman explores the hidden impact of poor customer data—and why dirty databases are quietly hurting go-to-market strategies, customer engagement, AI initiatives, and revenue growth.
Drawing from years of experience in customer marketing and database strategy, the conversation explores why data quality is no longer just an operational issue. From failed cross-sell opportunities to delayed renewals and unreliable AI outputs, Irwin explains how inaccurate customer data creates measurable business risk across B2B organizations.
The episode also dives into practical ways to improve database health, the importance of data governance, and the idea of "minimum viable data quality"—a more realistic approach to building scalable customer data systems.
During this episode, we'll answer:
- Why is dirty customer data a revenue problem, not just a technical one?
- How does poor data quality affect AI initiatives and go-to-market strategies?
- What makes customer data actually usable for decision-making?
- What is "minimum viable data quality" and why does it matter?
- How can segmentation improve customer engagement and advocacy performance?
Host 0:00
Welcome to the Advocast, where we discuss all things customer advocacy, discover the latest trends, best practices, and tangible advice as we talk to some of the brightest and sharpest professionals in the industry.
Matias Vivone 0:16
Hi everyone, welcome to a new episode of the Advo Hub, where we talk about what actually moves the needle in customer marketing and advocacy without the fluff. Today's episode is a bit of a reality check. We are tackling something that sits underneath every go-to market strategy and every AI ambition, but rarely gets to a spotlight that it's dirty data, or more specifically, why dirty data widely kills revenue and makes even the best AI look like useless. I'm joined by two amazing people here who bring the perfect combo of perspective. First, I have here with me Irwin Hipps, man, product and data authority, longtime advisor to teams trying to operationalize operationalize data quality and AI in this messy real world. So, Irwin, great to have you here.
Irwin Hipsman 1:09
Yeah, great. Looking forward to talking.
Matias Vivone 1:12
And then on the other side, I have Agusto D'odorico here. Augusto has been deep in real world implementations, seeing what breaks, what works, and what teams think is clean data versus what actually is clean data. Agust, welcome to this episode.
Augusto D'odorico 1:27
Thank you very much for having me here. It's a pleasure to me.
Matias Vivone 1:31
Thank you. Quick note, before we jump in, we are giving listeners a small gift today, instead of prompt you can use with your team to diagnose data quality fast, we'll share how to get those at the end, so stay with us until the end of the episode. And also this is an introduction to a webinar that we will host on January with more practical guidance and on a more conversational setup. All right, so let's get into it. A lot of teams say they want AI, they want personalization, they want better cross-sell, better retention, but they're building on a foundation of messy data. So, first question to open the floor is, What does clean data actually mean in practice?
Irwin Hipsman 2:23
So, I'll kick it off to me. Clean data, I think you used the right word, is the foundation. I've done research, and the average customer contact database - we're talking about customers, not prospects - is about a 47% so that's a failing grade in school, and the impact of, of this, think about who it impacts across the organization. Certainly, it impacts customers, because if I'm a customer and I'm a VP, but you think I'm still a manager, and you communicate with me like I'm a manager, or you think I'm in Boston, but I live in Argentina. You communicate with me like I'm in Boston. Customers expect that their vendors will have clean data about them. Then you think about, you know, group like Demand Gen, they want to know if one of your customers have left the company and they've gone to another company who happens to be an ideal customer profile, if no one's keeping on top of that, you know, all you're doing is waiting for that customer six months from now to raise their hand, but why don't you nurture them, incentivize them, you know, etc. And then there's, you know, the title, when they first met you, you know, they filled out an ebook, and they were summer summer intern and now they're a director four years later and you think there's still the summer intern because who wants to update titles and title becomes very important in cross sell so if you're trying to do a cross sell you're probably not going to cross sell to the specialist you want to cross sell to a director or senior director etc so when you think about healthy data, think about who it impacts across the organization, and then you see the complete picture.
Augusto D'odorico 4:10
For me, accuracy and guiding data can be summarizing for different worlds, because we talk about accuracy data, but the question is, what's actually accuracy data? For me, accuracy will mean meaningful data, so information you actually want to hear updated information. You don't only want the data to be there, you also want to make sure the data is up to date. Then you also want complete information. It sounds basic, but for being the most basic thing is also one of the most difficult to find, right. You have this particular field, and you have 10,000 rows, and you can bet, and out of those 10,000 rows, you will find maybe 150 cells that are complete, so that means only. 50 contacts out of 10,000 have complete information that you can use for that, and finally you also want data, and you can actually use right, not only meaningful that you can actually use, and to make something out of it, and those are the words that come to mind when you are where I'm managing databases here, and I'm migrating databases. I pick the data, and I have seen small databases, which are maybe I don't know, couple 100 rows, and they are super complete, that I have seen databases that are extremely big, and I know just as a glance that those databases have not been well maintained, and why? Because you see that database is sloppy, that information has not been standardized, and that whatever has been written has been written by a number of people.
Matias Vivone 5:53
I like that. I like this concept of data you can actually use, and not just having data for the record, or just to show a final number, but actually useful data that you can actually use to take business decisions. So this leads me to a question that also the main impact in my eyes is that it creates hidden revenue loss. Do you have any real examples of hidden revenue loss because of dirty data or not updated data that you have seen in your experience in real life?
Irwin Hipsman 6:31
Yes, I've done research into this question, and it obviously depends. The worse it is, the bigger the revenue impact, but typical you see about a two to 3% impact on ACV with bad data, so it'll impact things like renewals. So, typically, what happens, a classic situation is, you know, it's renewals three months before the renewal for the account, and the main point of contact left three months ago, and you don't know who the main point of contact is, so now everyone has to scramble and hustle and find the person, and there's nobody at the company who's a main point of contact. You finally get a junior person. The renewal maybe happens, but it might get delayed three or four months because you're starting all over again. So you multiply that across 49% of data being unhealthy, the impact on ACV is about anywhere from two to 4% which is not insignificant.
Augusto D'odorico 7:27
Well, I have seen a number of examples already. Typically, they go down to campaigns where clients realize that about 10% of their database is not learning reachable, not because the information is wrong, but because the information has been is updated, and as a real-life example, some months ago I was managing a database for a particular client, and I created a dashboard for them. Okay, and that's what dashboard was working with their opportunities, and just out of proactivity, let's say we extracted a couple of metrics on how much money they were living on the table due to incomplete information for the opportunities. Okay, I did not the average, and that mean on the standard case of what they're living out of out of the books, because it was not registered was $65 million That's how much it can impact not having information up to date.
Matias Vivone 8:26
So, guys, for, for also, I feel that for leadership, actually, there are a lot of myths around data and around AI, probably on their eyes, like head of departments or even directors, they feel that AI comes to fix everything when it comes to data, that's the main myth I have in my mind. And do you have any other myth, or what's the most common myth that you hear from your from from leaderships at companies about data quality and AI?
Irwin Hipsman 9:01
So I have a quote from Ed King. He's the CEO of Open Prize, and he said AI is likely the one technology that will expose your data quality issues the most, which may help raise the awareness and urgency of dealing with your go-to-market data quality issues. Maybe this will finally make the executives care about and invest in data quality, so you know what he's saying is that right now data quality can sort of be swept under the rug, because you assume that in many companies people assume it's the customer success managers for the for the customer contact data is staying on top of that, and the reality is, you know, a customer success manager, you know, not every account gets a customer success manager, or some customer success managers have 100 accounts, and the most they can track are the one or two people who come to their, you know, quarterly business reviews, the 10 other people who are so. Associated with the account, they can't go into LinkedIn on every night and see if there's any change, and we could talk about how often data should be updated, but I think what will happen is as AI comes in, people will start to realize how poor and how unhealthy the database really is, because it will expose the problems,
Augusto D'odorico 10:21
so for me the thing with data quality, and particularly with AI, is that AI is a tool, and as any tool it will work as long as the person will need actually wants it to work as intended. The thing with AI is not the difference that between AI and other different tools is that it's not just something that do what you want to, is that it will then learn from what you did, learn from what you intended it to do to do, and from there it will try to update itself, and what everyone is thinking is okay, it will then improve itself, it will do it better, and the question is, will it, will it actually do it better, because what it will do is it will try to do what you already asked it from your instructions, and it will do it faster and faster and better and bigger, but that doesn't mean that the original instruction was correct, doesn't mean that the original intention was correct. Okay, so with EI you are running two risks here, and I would like to bring the concept that is the slope. I know that everyone here knows that slope is just for a refresher. The AI slope means when you have this content created by EI and the content is purely generated just for quantity over quality, so for databases where you want to have the exact piece of information that's actually meaning and that you can use for any type of purpose, the slope is the most natural things you can work for two reasons. First of all, because you are having your tools and your particularly UEI affecting and trying to bring that information and to process that information to provide results, and those results will probably be trash, or will be contaminated. Let's say they are contaminated using wrong data, and the second part, the most very short one, is that you keep filling your model with trash after trash after trash, or slope after slough after slope, and the model will learn that the slope is actually good. Let's imagine this kid here, and you keep feeding him with McDonald's every day, hamburger after hamburger after hamburger. Have you seen Super Size Me? Well, you can imagine this scenario, right? Every day you eat hamburger, and you never tell him no, this is bad for you, and you tell them this is super good for you, and they will think that hamburgers are good, right? We don't, they are not, but at the end of the day, they will process hamburgers, and in this case, they will process bad information as if they are something meaningful.
Matias Vivone 12:59
I would love to be the kids eating burgers every day, aside, but also I take this that AI comes to expose the bad data and not to fix it, right? Like, you can have the very best tool AI tool, but if you don't have healthy data, probably that will not bring positive results, right.
Augusto D'odorico 13:23
Sorry, just to round up that idea. In the good hands, it will actually expose it, because the person behind the AI knows what they are doing and knows what they are looking for. In the wrong hands, it will actually just aggregate the problem, because you are leaving the actual decision power to the AI, and you are never taking into consideration what you are looking for when you have a database and you have data in there. The first thing you need to ask is why you want this information to hear, and the AI will not answer that question for you.
Irwin Hipsman 13:54
And a big place that shows up is clearly in segmentation. I mean, there are a lot of ways to segment, and you can get yourself a little crazy with too much segmentation, and there are a lot of people who are gray, and which segment do they belong in? But I think of sort of three layers of segmentation: one is segmentation by title, so you've got your C-level people, your VPs, directors, managers, specialists, and everyone can define it differently. You don't need 20 levels, it's probably five or six levels. Then you have your segmentation by the way the people you know, the relationship they have to your product. So, are they the main point of contact, more the business relationship, are they the administrator? Are they an advocate? Are they a frequent user? They log in, let's say, once a week, and you'd have to decide what a frequent user is. Are they an infrequent user? They log in, you know, less than once a month, and because you'll have some situation where there's a C-level person who's a frequent user. Or you know, with other situations where there's a specialist who's the main point of contact, you know, in a small company, so you know you've got to balance those two together, and then the third is the way that they, they engage with you, so are they engaged with the product, we talked a little bit about that, but they engage with you as a company, you know, do they open up your emails, do they register for your events, do they, you know, do they contact support? So now you've got a, you know, a three three dimensional chess board, and I think with those three pieces of information you can do some decent segmentation, and there will be sub segmentation, so for example, if there's a C-level person who's also a product user, you know, you have to think about them a little bit differently than if you know, and they're engaged with you, you have to think them a little bit differently than if they're, you know, a specialist who's engaged with your product, so you know, you could have, you know, if you only have, you know, 100 customers. Segmentation becomes less important, but imagine you've got, you know, 10,000 customers with 100,000 customer contacts. You know, you can get into, you know, 15 to 20 different sort of sub segments, and the key thing is to send them the information they need. If a C-level person is not engaged with your product, why you sending them a product update? You know they have people who want to get the product update, so you know, and then it becomes the, and this is where it's not just a data exercise, it's also people have to understand their customers,
Matias Vivone 16:40
and I feel based on all of these, one of, and coming back to the myths, I feel that everyone feels that this is a woman, one man show, and on the contrary, this is a team work, right, you need to be really sync with all with all the other departments, and this leads me to this topic of data governance, like not in a bureaucratic way, but in a practical way. Who owns the data, or maybe it's multiple departments who maintains it? I feel that probably you will, when you had this conversation multiple times with advocacy teams, but also sales teams, so trying to align them together internally and see, okay, who's in charge of this, and probably each of them feels that the other one is taking care of these, so who takes the lead, right? Do you have any insight around data governance?
Irwin Hipsman 17:28
Yeah, in the research I've done, the number one factor in having a higher customer contact database is a data governance team, or a team that thinks about it's not just purely the database, but team of people who thinks about how do we communicate with our customers. So, if I was, you know, at a company, and I've done this before at the companies I've been at, the very first thing I would do is I would do an inventory of all of the customer communications. What's what are those automated communications that get sent out automatically. Let's say to a new user. All the people need to communicate with customers. What are they sending out? And you'd be amazed when you see that list. There'll probably be 30 things on it. And product sends its thing, and this group sends that thing. And you'll see a few things. One is you need to understand, like, you know, who was sending out, and how does it get sent? Does it get sent through the through an automated the application itself? Does it get sent through, you know, you know, emails from the CSMS? Does it get sent out through an email marketing system? You know, what are the tools being used? Because you might be surprised that certain groups are using their own tool, and everyone else is using another tool, and, and then I think, most important, what are the open rates for these? Because a product open rate, you know, will be pretty high. The newsletter, you know, might be 50% if you're doing some segmentation. If the newsletter might only be 20% which is about the average, and are there any of these communications that we can maybe combine, and I might also take a look at what those communications look out and send it to the brand team and say, brand is everyone in compliance with with brand, or are we all calling, you know, the product the same way? What is the look and feel, because sometimes an email, particularly automated emails, got written three years ago by somebody, and they have the old logo, and they were written by an engineer, and don't sound like they're written by a human, so that overall inventory to me is the is the very first step, and I think once you have that inventory, then you could take it to other people across the organization and say here's what I discovered, and we need to have some data governance on who we communicate to. Then, when it comes to the actual database itself, things like, you know, the, you know, who's going to own it, obviously, but how often do we clean it up? Is it nightly? Mm-hmm. Probably not. Is it monthly? Maybe not. Is it quarterly? Probably. Is it twice a year? Absolutely. Is it once a year? Not too late. If you don't once a year, you have too many problems. And then, and what are our goals? Because, for example, the we talked about earlier, the main point of contact, that's your business relationship. That person should probably anybody with with that role accuracy, accuracy should be 99% Do they still work for the company? Where do they live, their title, etc. If they're a specialist and they never log into the platform, then you know what, maybe you shouldn't even worry about them. So you might have 50,000 customer contacts, but you've decided that 25,000 of them are the ones that we need to monitor and keep track of, because we want to keep track of our main point of contacts, our admins, our advocates, and our frequent users. You know, if we can have those four done, you know, at a 90 plus percent accuracy, then we're doing pretty well, and yes, there'll be some people that we don't do, but you have to prioritize, so that's part of that data governance, is who do we even need to, you know, focus on. I know Agustow has other other thoughts around data governance, so I will let you take over.
Augusto D'odorico 21:17
Thank you. Thank you. Now, actually, a lot of what you guys say lies a lot of what I'm into, what I'm seeing every day when I'm managing migrations and databases, and I'm telling you, data governance is not being discussed enough, is not being talked about enough, as it should, for me, data governance is the core of all that we are seeing here, because again it's at the center of everything, and at the beginning of everything, you cannot have a healthy process, you cannot have a healthy inventory, you cannot have healthy communications, you don't have healthy data governance policies in place, and that means that the people who is actually managing the information needs to know exactly what they are managing, and why you get a standard database. You get the information that comes to you, and you see names that are incomplete, names that have been put with different standards, company names that are noted in one way or another, locations that are incomplete. You have the city, then the city, and the postal code, and you have the state. It's not USA, maybe you don't have the state, you have the city only, and the country. Countries are in different notations. Well, you have many of examples, but the point is all that information is what is coming into the database, right? Is the food that all your process is taking, and we go to the hamburger again, you feed hamburgers, and you are not going to have a healthy body. So, for me, the biggest part with data governance, and the part that should be prioritized for every company, is not only the inventor itself, but also the practices, and make sure everyone understands who is managing the masters, how the information needs to be input, if possible. Try to go into a more detail level, try to close the fields, make sure they are not text-based, and you have options to produce the margin, and also make sure that those revisions you mentioned are happening at least quite often, because a lot of studies have suggested that for contact information every year 30% of that information becomes obsolete. That means on average, a database that has not been reviewed in, let's say, nine years - that's a lot of time, of course - but you can guess information that has not been reviewed nine years. You can take that complete database and just drop it, erase it, because it's not worth
Irwin Hipsman 23:47
it. I had lunch with somebody yesterday. She was a VP at a very large company. She talked about how they once cleaned up their database and they found people who had left the company four years ago, so for four years they thought that person still worked at the company, but that's what happens when you don't pay any attention to
Augusto D'odorico 24:06
it. I have,
Irwin Hipsman 24:07
and they've had three jobs since then.
Matias Vivone 24:10
Guys, probably this might sound, or could sound scary and complex for a lot of people out there, right? So one concept that you brought to our previous call was this minimum viable data quality, right? So you don't need perfect data, but it's enough with a baseline where AI and automation stop hallucinating and start helping. So what do you want to say about minimum
Irwin Hipsman 24:38
follow up on Agust for question. So, one of the things that you know, I've been around marketing a lot, and marketers spend a lot of time worrying about open rates, and they'll say, let's A/B test the message. So, do we say this or that? Let's A/B test. Should we should we have a plain text, or should we have a pretty envision? All, so how do we increase our open rate, like 1% or 2% The easiest way to increase open rates is to have an accurate database, so you know, for example, if 20 20% of people leave their employer every year, if you can take them out of that list, you're going to increase your open rate by 5% without any AB testing or anything alone, so that alone, and it'll be more accurate, also because everyone I would argue that most companies, when it comes to customer contact open rates, are probably 5% lower than they should be, and it's probably 10 to 15% if you did segmentation, so when we look at the minimum viable data, I look at three levels, and we've touched upon the first one. First one is what I call cleanup, that's the data entry errors. Second is so the updated enrich, the way you would do it is you would manually go into LinkedIn and figure everyone out, so you go in there, you've got 10,000 people, you hire somebody, and they spend three months, and they fix everybody, but by the time they're done, you know things already started changing. So you want to have the right employer, the right location, and the right title. And the third is, and this is where you get it to way beyond minimal, is there's for every customer contact, there's probably two to three contacts within that company who are not in the CRM, and so how do you add those people, and you can't just go to LinkedIn and just say, oh, well, who else is in this company, but how do you, how do you double or triple the customer contact database? So those are the three levels that we think about, the third level being the most mature. When you're at that point that you think, how can we expand the customer contact database? And those people need to be communicated differently, because they're not a prospect, they're not a user, there's somewhere in between. You have to think about how to, how do you communicate differently than you would somebody who's a, who's a user of your product.
Augusto D'odorico 27:05
For me to add into this concept, there is one particular item I would like to bring to conversation, which is actually the duplicate entries. Okay, which are part of the actual cleanup, but while there should be the map, maybe they are overdue, because you are talking about how the information should be standardized, should be rich, should complete, but from the point of view of database administrators, when you are looking at the actual entries, you see that I don't know, you have three vendors, and those vendors include information for other customers, so one of the first things I do when I get a new database is to try to catch these particular duplicate entries, and I tell you, there are a lot that means that those databases are not being properly cleaned, so that plays a lot into the data governance concept we spoke before, which is when you are trying to include information, make 100% sure that you are not creating the same entity,
Matias Vivone 28:05
I like that. So, standardization, duplication, some of the main concepts for these basics, and I think also it's important to say that you don't need to start with the whole database, but you can just pick a portion of it, a portion that you feel that it's relevant for the business, and that you will actually use for a go-to-market activity, right? So, Irwin Agust, thank you very much for bringing the real talk today. This was so, so valuable. Wrapping this up, if you listen this far, you, the audience, hear the headline: Dirty data is not a technical issue, it's a revenue issue. So, one clear action for you that are listening to this: register for the showering live webinar that we will have with Irwin and Agust as well. Seats are limited because we will actually work through live prompts with the audience, not just talk at you. So, we will have a small group attending, and as promised, about the gift about the data quality prompts that we mentioned at the beginning. Irwin, do you want to expand a bit on that?
Irwin Hipsman 29:10
So, I've got a series of prompts that you can use, and also some lessons learned where you can add those into Gemini, and it'll spit back out the, you know, new columns, the corrected information. I mean, you could do it yourself. It's not rocket science, but you know, we've spent the hours to figure out all the little subtleties.
Augusto D'odorico 29:31
I remember, if you're missing the correct prompts, you can end up making it worse. So pay attention to your win.
Irwin Hipsman 29:37
Exactly, I'll have to show them to you, Agust O' before we shout to everyone else,
Augusto D'odorico 29:41
thank you. Thank you from the master.
Matias Vivone 29:45
So, to everyone listening, if this episode was useful, share with a leader your organization who is pushing AI or any growth initiative, because I feel that everyone should know about the value of data. So, if you want. Join us live in January. Stay tuned for the registration link, and for the prompts that you will offer after this. Thank you very much to you both, guys, and to everyone listening in. See you next time. Thank you. Thank you much. Have a nice day. Bye, bye, bye, bye.
Transcribed by https://otter.ai