Stephanie Wiechers: You are listening to the Data Edge, where we talk about what actually happens behind the dashboard. I am Stephanie Wiechers CEO of Pearstop, where we work with facilities management, construction and infrastructure companies to make their spend data actually visible and usable. I have a background in data science and AI, and I've spent the last few years deep in procurement data for some of the largest built environment companies in Europe. Today We are talking about why your spend data probably is not telling you what you think it is and what changes when it does. So last Tuesday, last Tuesday I spent about three hours in ⁓ a bunch of Excel files ⁓ staring at spreadsheets and The reason I'm staring at these spreadsheets is not because I don't know what's going on in there. Right? So there's lots of data in there. And I actually have a very, very good idea about what's happening there. I would argue that I know this data better than almost anyone. But I'm doing a quality check and I'm going over the procurement, the spent items line by line. And somewhere around, I guess, row 200. My brain just stopped being able to I wouldn't be able to tell you, is it right what we're seeing here? Is it wrong? I I literally cannot process it anymore. So we get to line, I don't know, by then 312, and we see cleaning agent, five liters, Viking Direct. I know this this sounds very sexy. It's floor cleaner. Is it floor cleaner? So my mind goes, is that Correct? Is this miscategorized? Should it be general purpose cleaner? And so this is obviously a bit of a perfectionism trap if you go do this by hand and and these are things that you're gonna be looking at. But it's also really important because what the system is doing is yeah, it's automatically classifying every single line. So every item that's been bought last month, every item that's been bought last quarter. every item that's been bought last year. And it tells us, okay, from well like that's a Viking direct, from whatever supplier, what have we actually bought from them? What were all of the line items? What were the unit costs? And what does it mean? What are we looking at here? So in the end I'm making these judgment calls and and at some point I also get confused by just looking at the same stuff so over and over again. The reason I do this and I do this regularly is because doing the reviews trains our models. And so every time doing this costs me more energy than anything else I do. And here's the thing: when it costs me that much, and I built this, well, with an awesome team who deserve a big shout-out here, imagine what it costs your procurement manager or your finance lead. Whoever got handed the job of Just have a look at the spend data on top of everything that they are already doing. So, what you walk away with today, we're gonna talk about one thing. Why is spend data so hard to read? And what has to happen before it becomes useful? So, we're not gonna do a technology pitch here. ⁓ we're gonna do an explanation of what's going on under the hood. So, when we want to make this spend data useful for our commercial teams, for our head of finance, and of course for procurement. What are the analyses that we want to build and what do we need to get to those analyses? Let me tell you a story. ⁓ a few months ago I was speaking with the head of procurement at a mid-sized cleaning services company. Their operations were across several European countries. So there's hundreds of invoices coming in every single month, multiple languages from dozens of suppliers. So this head of procurement was not particularly a disorganized person. A very competent guy, ⁓ like structures, systems in place, a real strategy, plans on okay, what are the initiatives that we're gonna run? Of course, what kind of cost savings are we intending to make, but also what are the crucial supplier relationships. Anyway, what's happening? ⁓ anytime the invoice comes in. ⁓ they get logged, they get saved into a folder, and the totals did go into a dashboard. So at the end of the month, the CFO of the company has the total spend by supplier. Boom. Problem solved, right? Total spend on supplier level. So we go a bit into this conversation and basically I asked him, Do you know? How much you're spending on well, we're back to it, floor cleaner across all of your Dutch operations. And he laughed. He's like, Yeah, of course, it would be great to see exactly what's been happening per site, but like, are you gonna clock this data? ⁓ are you gonna log it? So I agreed, like, there is no way that you're gonna have someone do all of this by hand, it just It's an insane amount of work and it's an insane amount of time where we can see that there's a big return in it, but does that is that matching with the amount of work and and the amount of effort in a business that's, you know, already operating on tight margins to go do that. Anyway, so we went looking in this data and We found okay, all of the invoices are there, every single one, they're they're really well organized in these folders. It's basically a data gold mine. and just looking at these PDFs and going over it, ⁓ it's impossible. You you're not really reading data. You're gonna just go insane. So why is unstructured data harder than having no data at all? Right? What if we only had the aggregated totals? This is what we spend per supplier, and we don't even have to look into the invoices. That's gonna you know, that actually is much easier than having to go through all of the unstructured data. It's a bit counterintuitive. ⁓ but basically there's ⁓ a scientist called Thompson, and many many years ago, like the the 19, I think it was 1935, he did a bunch of research which was about the human brain and where we were able to focus. And what he discovered was people are notoriously bad at repetitive jobs. And so looking at data, going over it line by line, is a very repetitive job. ⁓ Let's say we're looking at lines that say TLT 5 liters, 1240. Vendor well what does your brain do with that? You need some context to know who are these guys, what are they producing. But also if you've gotta go do that line by line by line by line, you're gonna get so brain dead. So you're gonna start to take shortcuts, you're gonna stop actually evaluating what you see, which means you still don't get to the actual quality that you might want while still having to spend a lot of time because you know, even if you're fast It does take some time to assign a category to a line and say, like, okay, we've copied all of the fields here, and now we know what we're talking about. Anyway, basically the big fix here is you don't need to give it to someone smarter. there we don't need more people. We just need to give this data structure before anyone looks at it. And so then there's a process that we set up. ⁓ or that everyone actually can set up. It's invoices come in on the PDFs, the data points get extracted, so in a structure table format, like Excel works fine, or using a procurement system, an ERP, and then clean and classify all of those lines so that they're all in the same format, the same structure, and they have just a code or anything that links to say This is what this actual item is. Bingo. The question on how much were we spending on floor cleaner in ⁓ the Netherlands or in Germany or in ⁓ our London sites? It's it's a question that now takes four seconds to answer. And then you can add the next the next one and the next one ⁓ and the next one. And so I'm gonna Do the blue bucket example to show you what this looks like in practice. Right? Okay, so you find out you've been buying the same 50-liter blue buckets from let's say four different suppliers. And the prices of these blue buckets have ranged from 620 to 1180. Same bucket, no one noticed because the product descriptions. They're all slightly different. There's different suppliers, different invoices, and no one is looking at it that way. Because again, people have ⁓ responsibilities and a job, a day job. So you find out that there's one site consistently running 40% over the category average on consumables. And so this usually happens because these sites don't just don't know better. They've been ordering through a local supplier that they have a good relationship with. especially when there's been some MA, this you know, these kind of things stay in the company. And so they'll be buying from suppliers that are not on the preferred list. And what we see is very often it's just because they didn't know that there was a supplier on the preferred list, and no one from Finance or no one from another office actually knew that they were spending outside of the preferred list. So so no one was really able to point that out. and then another big one that we've been seeing and it it's ⁓ well been laughing about it a bit is how much do you spend on Amazon? And there's of course the how much you spend on Amazon, but how much does your company internally spend on Amazon? Because we didn't realize that we we're gonna run out of gloves and we just didn't make the order. Or we've placed, I don't know, seventeen separate orders for Nespresso Cups last month, whilst actually we've been doing that month after month after month. So maybe we can just order maybe once or twice because we know how much we're gonna need. so I mean unless you care more about the margins of ⁓ Jeff Bezos than of your own internal ones. I I'm assuming that like getting some Amazon spend out of the way is not something that sounds too bad. So now we've got all that, we ⁓ we've structured our data, we have classified it, we know okay what are the actual categories that we're buying, how much are we buying, what are the price variations Gives us a baseline, we can now negotiate when we go speak to a supplier. We know how much we order every month across sites. We even have a baseline for when we go into tender to really see what's this problem gonna yeah. Well, actually, what's this project gonna cost us? so all of these things they feel pretty good when you when you get there. ⁓ and so the question that remains is why don't we already have this? Why isn't it not set up yet? And the answer is pretty simple. It's this used to be a very, very tedious manual job. Or as I mentioned before, the the India the India bureaus who do a great job, but also it takes a long, long time, ⁓ 'cause someone has got to do it by hand and it's still expensive, right? It's it's line by line. and you're still paying for someone's hours. So the real reason ⁓ well especially most mid sized FM companies but actually we see it everywhere even ⁓ the bigger enterprise players have like more technology in place but even there the reason why it's not done is because it's really hard to get it right without the correct technology. So teams we're speaking to I'm speaking to big and small teams every day and buyers are managing supplier relationships. They handle exceptions, they support operations, finance closes the month. And the the dedicated data analyst who's sitting around and waiting to classify 40,000 invoice lines. So data stays unstructured. The dashboard is still at a supplier level. ⁓ So the shift that really happened in the last couple of years is that classification. Which is the part that used to require a person reading every single line, can now be done at scale with enough accuracy to be useful. So it's not perfect yet, and you'll see this. ⁓ I really encourage you to, after listening to this, take a small piece of your data set and just upload it to Copilot or you know, whatever the internal tool is that you're using, and see how does that classify. ⁓ then Try the same exercise again and do that a couple of times. And so what we'll see there is it's like already pretty decent. ⁓ and when you use a like a dedicated orchestrated pipeline, that's a that's really the thing that's gonna get you there, which is gonna perform always in the same and consistent way, and just handle of the data quality heavy lifting for you. So, what we've talked about today is what's possible in terms of procurement data, spend data, what do we see across sites? We talked about the blue bucket, and how literally being able to see how many blue buckets have we been buying across sites, at what price, from what vendor, in what quantity, what does that do for us, and what can we learn from that? If any of this today sounded familiar, the supplier level dashboard, the folders full of invoices, no one's looked at line by line. That is exactly what we work on at Pearstop ⁓ we take one quarter of your data and when we get started, we clean it, we classify it, and we give you a clear picture of what's this telling you. There will be a link in the show notes, but more importantly, thank you for listening to today's episode of the Data Edge. If this was useful, Send it to someone who needs to hear it. And I will see you next week.