Erwin De Werd: Every year, technical industries lose millions of euros or dollars, not because of strategy, but because of messy data. ⁓ to the Data Edge, the where we turn information into the greatest competitive advantage. I'm Erwin de Werd ⁓ and with Stephanie Wiechers CEO of Pearstop ⁓ we're away the jargon to show you how to build a data foundation that actually works. Stephanie Wiechers: you Erwin De Werd: From Microsoft Fabric to AI Readiness, we're here to help you stop guessing and start leading with data. Let's dive in in today's episode. Welcome Stephanie. Stephanie Wiechers: Thank you, Erin. I'm excited to be here. Erwin De Werd: Yeah, me too. So the topic for today. Well, in the previous episodes, we touched a little bit on the quality of the data and today we extend it a bit more and especially we focus on what to do when the well, the first initial let's say data categorization, data ⁓ quality improvements ⁓ does not work out. way we like to or there are some hiccups that needs to be corrected and how we organize that. So how we tell the models, the agents to do the work better. Stephanie Wiechers: Yeah, well, maybe to start, hiccups are the most normal thing in the world. And so whenever you're not running into them, that is a big, big red flag. Cause then basically you are saying that your first ever recategorization and your first ever processing would be a hundred percent correct. And that all of your AI models without any additional training have done their work perfectly. Erwin De Werd: Yeah, so that's already, let's say, the first misunderstanding that you would like to correct. And well, that's a good one. Because many times we think like errors or hiccups or something negative, but ⁓ we should see it more like it's something that's very common and it's part of the process. That's what you're saying, right? Stephanie Wiechers: Yeah. Yeah. Well, what we see is that, of course, everyone wants this dream scenario of first time, right? We start the data processing and it like works perfectly immediately. And there are two things happening here. The first one is, yeah, you want that perfect quality. And the good thing is it is possible to reach a very, very high data quality. And it's also possible to have that done. almost automated. like almost fully automated, let's say like 95 % of the work can be done with a software. But what we see happening a lot is that now that like people have interacted with Chatt GBT, they've co-pilot and they do see things working quite well is that also sometimes a magic bullet is expected out of AI. And what I mean by that is Erwin De Werd: Let's stop there for a moment. So yeah, that's interesting. So do you really feel and see this with companies and clients that they expect like this high level results out of the door that you have to? Well, let's say reframe the expectations in the first place. Stephanie Wiechers: Some of them do, not all of them. There's also plenty of clients who have very realistic expectations on what can be done and how fast is that going to be done or their timelines. We see, especially with smaller clients, that they do a lot of things themselves internally. So they have multiple responsibilities and then they have such a deep understanding of every single process within their business that when they come to us, they... They don't always communicate about all of those processes. And then, you know, we've like ran an initial data set and then they're like, ⁓ but actually, you know, we always, we always categorize it this way. ⁓ and then they tell us some background information about how they actually offer their projects to clients and how they design installation and some of the just interesting nitty gritty details. that are different in everyone's business. People forget to communicate those upfront. And so as soon as you start operating with an AI system that's going to be learning and training, that's when those kind of contextual bits that people usually don't communicate, they become apparent. And that's actually the great step, right? Because that seems like, oh, you know, it's not working, but it's not working yet as we want because now we've identified what input do we still need to just get this to this fully automated level. Erwin De Werd: Absolutely. Well, I can understand it ⁓ and that people expect such things. So I think it's good at start of the project to be very clear about how these things work and to explain them like that. OK, then there was a little sidestep, but let's then think what to do when the model is not quite doing ⁓ what we like it to do. So what kind of methodologies do we have then to, well, you can see this agent like our new employee. So how we train this new employee actually to do the work better. Stephanie Wiechers: Yeah, I'll just make it a bit more concrete with an example to make it come to life a bit. So ⁓ we're working on a really, really complex project with a client in the construction space on their data quality. And they want their data to be categorized into this new scheme. So there's not a lot of information available yet around how exactly needs to fit into this new mold. Erwin De Werd: Wonderful. Stephanie Wiechers: But they do have a lot of information on how did they do this previously. Then midway through the project, we ⁓ actually got introduced to someone internally on their side, who did this all before in the old scheme. So that was incredibly useful to get all this knowledge on on board, because those are the things we can't guess. ⁓ So now we're at this stage where we are evaluating what's the data quality and what are the last steps that we need to take to really get you to that 95 % desired accuracy level. So the foundation is there and we've been building a number of dashboards and those dashboards have made it so incredibly insightful on what is still missing and what like items are just when you look at it kind of logically going to be off. ⁓ So to give you one example is, they set up this really useful dashboard that lets you walk through all of the invoice lines that are now categorized within families, commodities, classes. So you can walk through the tree view and see what's happening and where have the items been assigned to. And so given that this is construction, we of course would expect a lot of materials. we would expect a lot of contractors. ⁓ But what we also saw that the system had assigned quite a bunch to this medical category. So like, that is not what it should be doing. And so we went into, like we did a deep dive. And so what we saw was, because there's a language issue here, that a bunch of tablets, like tablet computers, iPads, or in this case, they were like Samsung Galaxy, some things, they had all been classified as tablet in the way of ⁓ a lozenge, so like a pill, and which is the Dutch word for a pill. ⁓ Well, I mean, anyone who would look at this, any human who would look at this would obviously classify this as a tablet. Erwin De Werd: out. Stephanie Wiechers: But because this is not a big category, no like initial rules or oral data were available for this. And so the machine learning models didn't pick up on that straight away. Which like sounds a bit counterintuitive because you'd say like an LLM should be able to pick up on this. Yeah. Erwin De Werd: Exactly, I was just thinking like that. I think the telenovels should know this kind of thing, isn't it? Stephanie Wiechers: Yeah. Yeah, well, yeah. And that says something about how our system works, right? Because we don't just use LLM technology, then you just be a rapper around, like chat, GPT or co pilot. And I mean, I believe that then you're better off just building an agent and buying some extra tokens. What our system does is it has first a machine learning layer. So that is like, an algorithm that actually trains and like is confined to an internal state and like learns on how does this client process data and only for that client and then can replicate that. But so these small categories then end up often being like wrongly classified initially. Erwin De Werd: Can you maybe maybe I have a question here because that's something I think you must clarify here because I had some other experience in this field as well. Like everybody knows JetCPT and so on. And they will know, sorry. So they will notice tablets and everything, right? But what we build is basically different. So and. Stephanie Wiechers: Yes, correct. Erwin De Werd: What I understand is one of the reasons is that it must be very consistent doing everything the right way. And we all know that if we ask one and the same question to ChetGPT several times that we will receive different answers. But in our solution, we cannot have that. So can you explain a little bit what is the difference? ⁓ Why we not, ⁓ you cannot compare our solution with an LLM model like in ChetGPT for instance. Stephanie Wiechers: Mm-hmm. Yeah. So it works fundamentally different. It's both AI, but AI artificial intelligence is a broad category of any computer augmented intelligence, which can be as simple actually as a rule set. But when we talk about AI, what we mean is any model or system that learns. When we talk about large language models, LLMs, JCPT, Copilot, they have trained on gigantic databases of exactly that language. And so they are well able to predict on a very generic level what's going to happen. And they're like experts in many different areas, things that are just happening in the world. And they've read lots of scientific literature. They've read lots of forums. They've read, I don't know, cookbooks. So they have this like vast database of knowledge, but all general and publicly available. Now, if we talk about machine learning, we're going to look at, we are going to train a computer software or a system for a very specific task. And so in this case, the task is to assign categories to procurement line items. And there, we don't use the whole internet as a database. We use like a specific subset of input that in this case, the buyers would give to assign categories to the items. And then we train this model to replicate what's happening there. But then to like prevent that you end up just making a number of rules and you have like this elaborate script that goes if banana then fruit, if potato then ⁓ vegetable or Erwin De Werd: Yeah, yeah, Stephanie Wiechers: carb depending on what country you're in. Yeah, and the system does learn and you assign weights to it, which is like, well, that's technically a pass filter. So a percentage passes through a percentage is always random. ⁓ And then after you've done a lot of data input, then you have a reliable system that's able to assign categories. Erwin De Werd: Yeah, yeah, exactly. Yeah, OK. Well, ⁓ yeah, thank you for the clarification because I think it's important to understand these differences. So the other thing what I like you just said about these dashboards in this tablet example, you can obviously find this very quickly that you have to correct something here, but it also say then that. It's really cooperation of the AI and human ⁓ work, right? To get this done. And I like this idea and this whole setup that how we cooperate and actually help humans to do the work more efficiently, better, faster and so on. But this interaction and I think we see this more and more to improve these models that we are needed like a human right to to help them to do the things that we expect from them. So that's an interesting example. Stephanie Wiechers: Yeah. Yeah, I mean, some fantastic things were implemented. So we have a really, really great lead engineer. And even before I jumped on a call with him, he had already read the email. And he was like, I made some changes in the biases. And the bias is basically, where does this model lean towards? And that can be both positive and negative. So we had already implemented a number of positive biases towards more construction, their industry. and then this number of specific types of projects that they work on. So now we also decided to go through that whole categorization and determine these are not so logical, check in with the client, do you agree on this, the client was kind enough to also provide us with a list. Which, you know, not every client does that. That was like really great in this case, which made us able to then install extra positive and negative biases so that the predictions of the model in a very high level are already hinged towards what we would expect just from a normal human kind of industry experience way ⁓ for them to pan out. Erwin De Werd: Nice, very nice. Yeah. Yeah, so, and it also, of course, when you say this, I realized that this model and the machine learning and so on and so forth, that is building a knowledge center for the client and what they really always have available. So that's another, well, maybe it's a side effect, but I think it's a huge benefit also for the company that they working on that and invest in having such a thing. Stephanie Wiechers: Mm-hmm. Yeah, yeah, actually, something that I was thinking about after the well, just basically what happened this week and the updates is we train all these models and they are now able to replicate the internal workings of how the client categorizes. It would be really interesting, maybe for them to have visibility on what are the rules the AI has distilled. So like, I was just playing with the idea of should we implement this feature into our system that just anybody can really clearly see and edit all of the rules that have been imposed now by the system. Erwin De Werd: Yeah, well, it's maybe something to consider, of course, if you want to give that away. Yes or no. But it's true, it's like the human intelligence has. This is then the AI intelligence that is behind this thing. So that's an interesting discussion as well, of course. What belongs to who and how valuable this is? I think that's because it's like an IP, right? Stephanie Wiechers: Yeah. Yeah, yeah it is. Yeah, it is their data. Erwin De Werd: Definitely. Okay, yeah. Yeah, so in this episode, I think there was quite some interesting topics. So first of all, what we discussed is that specific AI, ⁓ many people think AI is AI and that's not the case, of course. So LLM models and what as we know of them in the big audiences. are completely different from this specific solutions. That's one of it. The other thing is the cooperation between AI and human people to bring this to the highest level that we need is 95 % to reach out to that. And the other things, of course, the IP that you built up by doing these things. So it's not only. that you work more efficiently by this data categorization and building the fuel for your AI applications, but also you're building really IP by having this into your organization. So I think there's quite some good takeaways from this week's episode. Thank you very much for, like always, your very valuable inputs. Stephanie Wiechers: Hehehehe. ⁓ Erwin De Werd: And I'm looking forward to the next one in next week's episode. Stephanie Wiechers: Looking forward to it. Thank you. Erwin De Werd: Thank you to everybody, all the listeners and tune in for the next one. See you then, bye bye.