Erwin De Werd: Hi, in the race to AI, most companies are starting with the flat tire data. You're listening to the data edge. I'm Erwin de Werd from Peerstop and in each episode Stephanie Wiechers and I explore the intersection of data quality, enterprise standardization and the real world value information. Whether you're an asset manager or a bit specialist, we're here to give you the blueprint for a reliable data layer. that fuels smart decisions and next generation technologies. This is the foundation that is working for you. Stephanie in today's episode. Stephanie Wiechers: Thank you, Erwin Erwin De Werd: And today we have an, that's why I have this introduction because we have an interesting topic and that's actually, yeah we talk about AI and the news is all over of course, but let's see how ⁓ the human factor is in getting all the things right. Stephanie Wiechers: extremely ⁓ Erwin De Werd: Yeah, so. us to a project that you're currently running in per stop ⁓ ⁓ that interaction between the human and ⁓ machine is is working out and where are the pitfalls and the benefits? Stephanie Wiechers: The project I want to talk to you about today is procurement related and I'm going to tell you a story. Basically our client came to us with a very complex question around data integrity and data quality in their procurement department. They have a system that's already in place and they have a way of and registering departments, across entities, what exactly their spend is going to. But now. ⁓ Because the organization is very big and they have multiple locations across different countries, actually their goal in order to create enterprise level standardization and to have insights that span across the entire group, they need to move towards one single source of truth around procurement categorization. So. That means that it's really, really important that if one team categorizes something as a contractor work in welding, then the other team can't categorize it as general contracting work. They need to identify that in the same way. And so that poses a complexity for AI. Erwin De Werd: Yeah. Stephanie Wiechers: because AI still needs to know, okay, what's actually going to be the way that we're going to register this internally and what is our preferred way ⁓ logging and registering hours. Erwin De Werd: Okay, makes sense. Stephanie Wiechers: Yeah, so what happens there is, say of the examples that they gave ⁓ there was a real issue around automations is have buyer department that's working on these categorizations. And so the current way working poses that someone has the knowledge of knowing, when a contractor goes to a site at a different location on their invoice, it will state separately the hours that they worked on the project, but it will also have maybe travel expenses, hotel cost. so there's a bunch of other types of costs that are on the invoice of this contractor. So now, ⁓ Every company will have their own practices around how they will book those types of costs. So is all of the contractor costs, including the hotel, something that goes towards human hours ⁓ and logged? Or you decide to split that and categorize the contractor hotel expenses under your hotel and sleeping expenses? That's not a question that any AI can answer for you, it is also a question that any AI ⁓ answer for you. Erwin De Werd: Okay, I'm thinking about this because basically what this kind of company is trying to achieve is to systemize the way they work. And course you can work, you work with lot of people, like you explained, your different departments, it's a big company. So they have different ways of doing and that's what you see with if you do this ⁓ only let's say human power. So it doesn't work when you only hit on discipline and structure from a human point of view. So if you want to systemize this approach, then ⁓ system, in this case, the AI system, need to understand how to do these things, right? So, and basically it means ⁓ humans have to teach them this, right? Stephanie Wiechers: Yeah, it does. Yeah. Well, this specific client already did is they actually had a really, really good and smart way of working already. So someone had designed a system that had a number of rules. So of the costs would come in and they had generic categories to categorize it, to put them down in. Erwin De Werd: ⁓ okay. Stephanie Wiechers: And based on these rules, a number of the items, quite a large number of the items would already get an additional amount of tags. This didn't lead to the level of granularity that they actually required to then drive value from these analysis. So it was okay for reporting purposes to see internally like really, really high level. Okay. What is going on here? but it wasn't okay to have any spend optimization. a ⁓ of the costs were kind of mapped in two generic categories. ⁓ Erwin De Werd: ⁓ yeah, so one of my questions if I hear this, so what did you do? Did you take these rules of engagement, what they already had, and just implement this in the system, or was that not enough? Stephanie Wiechers: Yeah. Well, no, that was not enough. What we did is we actually combined all the possible resources. And this is an advice that I would give to anyone who is looking to expanding their reporting capacity or implementing any type of AI. expect a magic bullet and think about all the possible resources you have available and combine them. Erwin De Werd: Now, OK. ⁓ Stephanie Wiechers: What did that mean practically in this case? It's that in order to categorize this data at this granular level that was required to really drive those spend optimizations, what we did was we made a combination of working with large language model technology. So you always use the enterprise plans so that your data is safe and like guarded. Erwin De Werd: Yeah. Stephanie Wiechers: Then designed a number of machine learning models that were meant to use user input. So the user can now train and the user can say, well, we always categorize the hotel stay of this contractor as like their full scope of work, rather than like booking in another hotels. And so all these types of ways of working can then also be learned to system. such that the next time it's not just generic rule, but it's really okay on this granular level, the system knows and learns and it that knowledge that the people input, make sure there's consistency across and it really like speeds up this process, freeing up the buyers to do what they are real core. work is, which is buying, making contracts, having good relations with the suppliers. Erwin De Werd: Yeah. So can you tell something about the accurate level that is required by companies that because of course the system can do their work and you explained about adding the rules and combining these things. what is the level of accuracy that come ⁓ of this and what are companies typically looking for? Stephanie Wiechers: Yeah. Clients come to us with whole different ranges of maturity levels on this scale, where companies ⁓ some way or form of categorizing their data. They've often started working with co-pilots to create some quicker steps. Many them have an agent working already that, you know, does ⁓ like... does a small bit of the work now. they do have a categorization system. They do add these categories to their spend data. And so either that's just costing them a lot of time. So that's expensive the level of granularity is just not sufficient. And they're looking to get one step deeper and also make sure that there's real consistency across how this is done in a manner. So you don't have to wait for like a month to get your management report, but you just get it straight away. Cause you know, the system never sleeps. I would say that's, that's kind of like average on the spectrum. And then if you're looking towards like, what's the real good situation that you could achieve here. that's when ⁓ actually make this combination of large language models, ⁓ machine learning, ⁓ on all the rules that were already applied and all the knowledge and then taking that into the system. ⁓ yeah. Erwin De Werd: Okay. Stephanie Wiechers: And so when a company already has, when you already have your systems in place and you have a way of doing it by hand, ⁓ the system just learns how that's being done. So that means that just right off the bat, when starting off, you've trained the system and it just starts replicating all of those things that you've already been doing internally. So with a lot of previous input, the ⁓ the accuracy just goes to the roof and gets you the replication straight away. Erwin De Werd: Hmm. Can you give an example for this project that you're currently working on? How the, let's say the review process, where the validation process work for, and how did you set it up ⁓ with the people with all the knowledge and the system that is in learning mode? Stephanie Wiechers: Yeah, yeah. It works like this. So they have different departments and different buyers who are responsible for a different part of the business. So every buyer gets to review only that are relevant to their department. then they'll be presented with an overview where the system automatically recognizes this is a high confidence item and this is a low confidence item. So in the prediction that it makes, Like you can see it as a junior employee. If you ask a junior employee to put your spend in buckets, then they might be able to deduce that an aluminum beam is an aluminum beam. And so if you have that one-on-one mapped, then it will map it correctly. And so you get a high confidence score. But on things like hours and hotel spend, ⁓ you'll automatically see in the system that this is not a straightforward category to map. And so it will highlight it as a category that needs to be reviewed. And so then the user is presented with like a top list of, these items need your manual review. And then... It's really easy for them. just open a menu where they can see, okay, this is the tree structure. So this is the structure of how the categorization here works. You can search and it like makes a number of suggestions. then it also shows you, okay, this is how you've mapped this item previously. ⁓ you can look into history and so that you really guarantee consistency across this categorization. ⁓ then all of these reviews ⁓ ⁓ or however often you do the reviews ⁓ bundled and they feed back into the system and then ⁓ next month we'll know ⁓ okay well the hours of contractors are always categorized like this so I'm going to put them in the bucket of contractor hours. Erwin De Werd: Okay. So that's great. So the whole validation process is built in the right? In the application. Stephanie Wiechers: Yeah. Erwin De Werd: OK, that's great. So that lead me to a last question and that is OK. So now we organize this very nicely and we receive all this knowledge from the local team, the local professionals. I can imagine that the company would like to know what happens with all this IP that they enter into the system. Is that something that will be shared knowledge across application and can be shared with other companies or Stephanie Wiechers: Mm-hmm. Erwin De Werd: Or is it something that is specifically for their project? Stephanie Wiechers: Yeah. Well, that's a really good question. So many AI companies have different policies around this. So I can only tell you our policy, but I know that in many of them, all the data is shared across the platform. What we do is we have an opt-in. And so when you decide we want to participate in the opt-in, that means that all... Erwin De Werd: Exactly. Stephanie Wiechers: information that is non-confidential gets anonymized and then it's used for training. But as easy and a lot of clients use this option, you have own internal trained algorithm and that means none of your information is going back into any system. Erwin De Werd: Okay, that sounds good. So I know Perstop work with quite let's say enterprise companies ⁓ started these projects. So that mean also that comply with their regulations, ⁓ right? it fits in their policies of working, is that correct? Stephanie Wiechers: Yeah, it does. data safety and security is something that you cannot take serious enough. It's really, really important and it's critical to many businesses. And so believe every AI company should have safety and security as one of their top values. Erwin De Werd: That's right. All right, so I think you gave us some more insights ⁓ the level quality and review processes that is a combination of people and the knowledge within the company and the technology. And that's something that is an interesting topic, how the man and machine work together, right, to come to the better results. And I think that's the right way to do it. Okay, so you so much for ⁓ week's ⁓ ⁓ I look forward to the next one next week ⁓ can you maybe already say a little bit what we will do in the next week? Stephanie Wiechers: I can give you a bit of a sneak peek. we will soon be kicking off a series ⁓ interviews with a number of our clients. So I'm really, really excited to share their knowledge as well and to have a broader conversation and really open up the floor to what solutions like this can do and how they've improved their businesses. Erwin De Werd: Wow, yeah, that's great. Yeah, that means really knowledge from the field. And real valuable. We are looking forward to that. So thank you again. And then I hope to see everybody next week. Stephanie Wiechers: Thank you, have an amazing day. Erwin De Werd: Bye bye.