Victoria Lawton: Here's a problem every CTO and CISO grapples with. Your team writes clean, audited code, but under the surface, your application relies on thousands of open source packages your team didn't write. When a CVE hits one of those dependencies, you may be trapped between breaking your application with an unvalidated upgrade or running vulnerable code in production. So today we have Mike McGrath, Vice President of Lightwell Engineering here at Red Hat. to talk about how we're working towards eliminating that choice entirely. Welcome to Technically Speaking, where we explore how open source is shaping the future of technology. I'm your host, Chris Wright. Mike, great to have you here today. ⁓ I thought before we got totally into the details of Whitewell and patching and security vulnerabilities, ⁓ what's your background? ⁓ we we've been working together for a while. ⁓ you have a long history with with Linux. G give us a a taste of of what you've been up to before this. Yeah. I mean, I going back to the beginning, I started in the nineties with the Slackware Bible that I probably got at a Sam Goody or who knows what. and then I started volunteering in the Fedora project right around Fedora Core three and ⁓ Red Hat hired me about nineteen years ago. ⁓ and I've had various jobs working on RHE or Fedora or OpenShift and prior to LIWA I was ⁓ leading core platforms which included RHEL and OpenStack and Satellite and a lot of those support communities. And ⁓ it's been a I've had a I I think by any measure I've had a great and storied career at Red Hat and I've had a good time here. Yeah. I'm Now remembering downloading the entire alphabet of floppies off the internet for Slackware. ⁓ thank you for that. ⁓ so the context here for for Lightwell starts, I'd say back in the spring. ⁓ we had what's often referred to as the Methos moment. ⁓ Anthropic announced Methos and established a program around that called Glasswing and put a bunch of content out there describing A shocking set of vulnerabilities and capabilities that the models and harnesses had to find vulnerabilities and software. ⁓ That was many months ago at this point. And of course, there's many programs out there. It's not just a single ⁓ Frontier lab like Anthropic. There's Daybreak with OpenAI, and there's open source projects in this space like Acreetes. ⁓ Walk us through that, that initial phase of recognizing the ⁓ the reality of something like massive vulnerability scanning and trying to translate that into a a response. I think our ⁓ my main focus on this back then was on RHEL and we had seen this slow but steady increase in vulnerabilities that were coming in from the community and it was not even. Some communities were getting hit very hard, some communities weren't. And as we looked at it, we realized lot of these vulnerabilities, even pre-mythos, were coming from AI. And so for us, ⁓ originally it was less about trying to find vulnerabilities and it was more about how do we respond to this? Because ⁓ you know, RHEL is the enterprise Linux, as far as we're concerned, and ⁓ people expect it to be secure. And so for us it was kind of about managing that increase. ⁓ and then earlier this year, we decided, well, let's just start taking a look at this ourselves and be more proactive about it. That was maybe a month before you started to hear sort of leaks about this mythos thing that Anthropic was working on. And instead of just waiting for it to happen, ⁓ we had some conversations about scanning proactively, ⁓ using our own harnesses and you know, at the time I think Opus was the top dog for that kind of thing. And we just started looking, ⁓ and it did not take us long to find unknown vulnerabilities. And I think the big aha moment for For my team was that we did not have security researchers doing this. These were people that were skilled in AI, they were software developers, but they were not inherently security people. ⁓ and ⁓ in finding these, it was also a matter of where to scan and how to scan. And ⁓ that kind of led into then, you know, mythos being released. And ⁓ I think that was even for us, you know, I had access to to mythos. ⁓ trying figure out what it did and how it did it, ⁓ you know, took only a few weeks and then it's ⁓ now we're in a different world. For the longest time, we've had from a from a Red Hat point of view, a lot of experience working with open source communities and doing ⁓ security vulnerability responses and disclosures responsibly and focusing predominantly on that platform tier where we sit. ⁓ Linux kernel being a prominent example of areas where we pay close attention to security issues. The the rel world has an application runtime, a language, Java, ⁓ Python, et cetera. And we had already started bringing newer versions of this and maintaining ⁓ patched versions of of these language runtimes. But the mythos scanning and the associated kind of cybersecurity vulnerability discovery showed a whole different world up the application layer. ⁓ so h how have we kind of shifted our focus from that very platform-centric view of the world to starting the application down and coming through those dependencies? Yeah, as as we've started to look at this in actual enterprise environments where this code is actually running, ⁓ it is not at all uncommon for more than half of the code that is being run to be in that sort of middle of the sandwich. ⁓ they've got underlying enterprise support at the operating system layer, usually something like RHEL. And then they'll have some platform support higher than that. Yeah, maybe that's a a J Boss or something like that. ⁓ but more than half of the code is kind of running as dependencies in between those stacks or on those stacks. And ⁓ for us, the the real question was understanding and talking with customers that unless that full stack is secured somehow, they're always going to have problems and vulnerabilities. And in the past, we've shied away from going higher in the stack because it is an expensive proposition to do that. ⁓ but it's one of those things that while AI scanning kind of frightened us ⁓ into looking at this space, the capabilities of those large language models have evolved enough that we can also seriously look at evaluating fixing that space. And I think that's where, you know, all of the this culmination of all of these conversations came together to create light well. It's an insane goal, really. If you if you go back in time, scaling with people to do all of this work is a global scale problem. Internally we've kind of referred to it as patching the internet. That's all. No big deal. Patch the internet. ⁓ and yet we need automation, we need tools, we need ⁓ AI LOMs to help us through the this process of not just discovery but remediation. ⁓ I think a really interesting aspect here is we see you know Take take an average enterprise which is built of thousands of applications, each of which has thousands of dependencies. So you get this sort of order of magnitude thing with that that sandwich that you're describing being on the order of hundreds of thousands of of different small language specific modules. ⁓ If you're trying to manage all of that as as an enterprise and think about the change associated with all of that, I think that's by itself overwhelming, just the sheer volume. In this lightwell context, the versions that are being deployed into production are typically quite older than the versions that are being worked on actively in the upstream community. I think part of the critical focus of lightwell is patching those versions that are in production while sort of this parallel life of of bringing it into the upstream so we're making the overall ⁓ internet and open source sa a safer place. Yeah, I think looking at community releases where they're always focused on securing the most recent code ⁓ is is a fine goal for for communities to work on. But the code that is actually running in these production environments is is can be much older. ⁓ in fact, the oldest vulnerability that we've discovered that is in production in a in a customer's environment ⁓ was a vulnerability found in 2001. And so you at that point you have to assume that whoever put that together has probably moved on to another job. It's not their full time job looking at that. And that goes true for hund the hundreds or thousands of applications that these customers are running. And so a focus for us with Lightwell has really been on securing the content that they're actually running. in their environment. And that's the last thing most customers want to do is completely upgrade some application that on its own is running just fine. but because the communities have moved forward and the technologies have moved forward, sometimes it can be weeks, months, maybe even years of work to fully upgrade that application. and as a business, you want to focus on the things that are that matter today. And so by being able to focus just on the Very surgical issues related to security. ⁓ it allows them to keep running application or upgrade it on their terms when they're ready to do it. this is that sort of peace of mind that Lightwell, I think, brings to the the table. The the ripple effect of simply bumping the version number of a single dependency through all the transitive dependencies in your application could effectively touch the entire dependency tree. And then totally destabilize the the application. So I I love that that surgical view. It's a really great way to to think about it. ⁓ what I've seen is ⁓ there's a back to basics here. not all not all businesses have a robust awareness or understanding of what they have deployed where. So sort of like the table stakes is what is your dependency tree? What Yeah. What are your applications? Where are they deployed? Because being public facing versus multiple firewalls into your data center has a very different risk profile. and then what are those application dependencies? So as we produce something, you can make an intelligent choice about how quickly you respond in one context and where you might patch it in another context. ⁓ the the other piece is just good old fashioned CI C D. How quickly can you automate? an update to an application. So again, this back to basics. What do you have? Where is it? And how quickly could you have a push button system that can update an application in production? All right, we got a pretty good view for the problem domain. let's talk a little bit about under the hood, what is light well? We're we're building a lot of content, we're doing the scanning, we're doing the remediation. what does that look like inside the the light well factory, if you will? Yeah, I think most engineers that have worked with Claude or Cursor can kind of work through in their mind how finding and fixing a vulnerability would work on their machine. ⁓ I think ⁓ Lightwell has taken an additional step of building fully agentic harnesses to do this ⁓ sort of headless, if you will. And so ⁓ you know, we are ⁓ sort of taking that ⁓ step in from hunter gatherer society type work with security. into fully industrialized mass production fixes and remediations. What that looks like is it starts with an assessment of what comes in. ⁓ one of the big things I think ⁓ I've enjoyed about Lightwell is ⁓ instead of focused just on how bad a security vulnerability is, ⁓ we also have an assessment of just how complex it is. Because some very serious vulnerabilities could be a simple one line change. And even some very minor vulnerabilities can require an entire application rewrite. Anybody from my team that's about to listen to this will roll their eyes. ⁓ but there there's a famous quote that I think sometimes attributed to Lincoln, sometimes Washington, ⁓ which is if you give me five hours to chop down a tree, I'll spend the first four sharpening my axe. And I think that is very true in Lightwell, where we spend quite a bit of time prepping for the remediation, ⁓ so that we can do it. And so ⁓ that assessment at the early stages of our our intake process ⁓ kind of informs everything else we do with that vulnerability, how much human oversight it will need, and certainly ⁓ the checking and rechecking of the the packages are being built because ⁓ we know that if we just because we fix something does not mean that it has been fixed well. And the last thing we want is for us to ship a a package out that looks like somebody took Thor's hammer to the source code, ⁓ making it unrecognizable. Yeah. Yeah. I think there's a a lesson in there. That goes beyond security, just it's easy to imagine how agents will solve all of our engineering problems, but it turns out there's still a lot of noise in the system and ⁓ getting that noise out of the system so we can focus on the real problem. Key, key here. ⁓ now let's talk a little bit about the scale. ⁓ we you know, we are just patching the internet. ⁓ what does that look like in the in the light well context? We've we've GA'd ⁓ light well, we're we're in the early stages of building out all of this. content and our content repository. What what does that look like? And what are the longer term goals and vision for for Lightwell? Yeah, we're mostly focused on on Java and Python, which is great because we have a lot of expertise in that area. We've got a middleware team that has been building Java packages for decades. And just within Red Hat and some of our internal tooling, and if you've used RHEL you've probably noticed quite a lot is written in Python. So we have a lot of expertise in this area already. ⁓ when it gets to actually ⁓ Trying to bring this flywheel up to full speed, quite a lot of it is built into what we're calling the validated packages. And that is just very simply trying to take what upstream did and get a build out of it. ⁓ we haven't fixed any vulnerabilities in it. It's just about making sure that we know how to build that package, that we can test and make sure that that package built correctly and that we we didn't break anything in the process. That's really the funnel of everything that we're doing. After that, it's a matter of applying patches to it and ⁓ then making sure those patches didn't break anything. And I think one of the surprising things for me has been you've got a package that's got known vulnerabilities in it, maybe it's got some unknown vulnerabilities in it. Going through and fixing each of those one at a time, pretty easy for automation and AI to do. one of the more complicated steps has been how do you smash all of those vulnerabilities into a single build? Which is the customer expects. They want all the vulnerabilities fixed, not just a couple. And that turns out to be a fairly error prone process, even agentically. And so, you know, that's ⁓ I would say the other end of the funnel is once we've got everything built and we know that it can build, how do we then prepare the final deliverable for customers? ⁓ it's another area where we have quite a bit of agentic ⁓ workflow built into it, ⁓ but also a pretty smart ⁓ subject matter expert review process. ⁓ so that human eyes are involved when they're really needed. And that's ⁓ you know, really where we're spending our time on on speed. I think the other part of that too is just the error rate. ⁓ the higher error rate you have in any system, the more human eyes you're going to need. And so that's ⁓ you know, just one of those internal metrics that we really keep an eye on. It's a pretty sophisticated system and ⁓ obviously still ref being refined and and being built. ⁓ that's generating Content, you know, recognizing the inputs, generating the content, delivering that content. So we've talked a lot about the the patching process, you know, discovery, remediation, delivering content that's version specific to what's in production. What's in production, typically trailing pretty, pretty far, what's in the open source community development trees. And so another key aspect of our work is getting those patches that aren't ⁓ simple backports from what's already been fixed upstream into the upstream communities. what is that looking like? That that's a a monumental task. We're talking hundreds of thousands of different open source projects. Yeah, I and I think this is one of the areas where I'm really happy with the business model that Red Hat has come up with, ⁓ because it really does find a way to sort of funnel this enterprise business into bolstering open source communities. And ⁓ I think one of the things, you know, Red Hat's always been very involved and committed to our ⁓ upstream communities. And this is another good example of that, ⁓ especially where we've been able to partner with IBM just to find people to help take these things upstream. while we have this great agentic factory ⁓ to find and fix things ⁓ that does not do us a lot of good in terms of sending things upstream. ⁓ many upstream communities simply would not be ready for an onslaught of content. And I think anybody that's been involved in open source knows how poorly a sort of drive by bug report, ⁓ how poorly that can go. Yeah. And so we've taken a very human approach to this. ⁓ we will ⁓ be sending ⁓ people upstream to work with those communities. If they have a ⁓ embargo policy, we will follow that. if They don't, then we will we know to reach out to somebody and say, hey, I've got a vulnerability. I know better than to just open a bug and and leave it out there. and ⁓ you know, I think there's a there's a sort of rigor and care required ⁓ when taking things upstream, especially because not all upstreams are the same. Some of them are very vibrant and ⁓ we'll say economically sound where they will be able to respond to this very quickly. Some of them are somebody's hobby still, even though they're running a critical piece of of ⁓ of library infrastructure for Java or Python or whatever. And ⁓ we also understand that just because we found this vulnerability doesn't mean that upstream is going to drop everything that they're doing ⁓ to to fix it. And one other policy we have is when you know we we aren't just going to go upstream with a vulnerability. ⁓ we'll go upstream with a fix in mind and we'll work with them on that fix. And for the actually our very first vulnerability that we sent upstream ⁓ they decided not to take the fix that we we provided. ⁓ the approach was the same, but the code was different. And our policy is to then go back and use whatever upstream did. And you know, we'll get rid of our fix and we'll rebuild ⁓ based off of whatever upstream did. ⁓ and then the other thing we're looking at is ⁓ other communities like Accretes, ⁓ which is is coming up. And ⁓ I think we're interested in in seeing how that goes. You know, that has a lot of potential for helping ⁓ Both the future of Lightwell and open source security everywhere, which at the end of the day is the goal here. Yeah. I think it's really important to recognize the the human nature of the trust relationships in communities. So some communities have turned off pull requests from non-identified members of their community. So the the drive-by patcher is just simply ⁓ rejected at the door. ⁓ others are gonna want to see some due diligence put into the process so that you're not just getting garbage. Yeah. ⁓ there's plenty of ⁓ conversations around AI slop and this kind of thing that that really can slow down anybody's ability to even look at a potential issue. And then is it close not a bug? System works as as designed, or is it a a a real issue that needs to be fixed and how do you fix it? All these things I think are really important. Even though those these dependencies are in production and they're older There is a future state where that application could easily bump a version dependency to a newer version. And so part of this effort is broadening just the security stance of all of open source software. And as you go forward to a newer release, it's the security fixes are baked in. I think this is also a really important aspect of how we sort of slowly shift the tide from a bulk of vulnerability discovery to secure development practices and rapid remediation and at the end of the of the day, what do you think? Open source comes out more secure? Yeah, I think I think it has to. And ⁓ well I'm not too s too worried about job security at the moment. ⁓ I do think that ⁓ the future of open source ⁓ is not just more secure but more secure than the proprietary options. And I think a big part of that ⁓ is because ⁓ large language models will always know more about open source communities and projects. 'Cause a lot of times they've been trained on those projects. Right. And so I think that there's a really great ⁓ I I happen to think this is a really great marriage of ⁓ large language models, harnesses and and open source. Yeah. I love that. So we got a a more secure internet and better rested CISOs and security teams. That's right. That's awesome. That's right. Thank you so much, Mike. What a what a great conversation. Yeah. Great to be here. Software maintenance shouldn't force us to choose between innovation speed and security. By decoupling vulnerability patches from forced upgrades, Project Lightwell fixes security flaws without breaking runtime stability. Just as open governance and enterprise Linux brought trust and standardization to operating systems, automated patch factories are now building that same trusted baseline for the era of AI. Thanks for joining the conversation. I'm Chris Wright, and I can't wait to see what we explore next on technically speaking.