Brandon: Hey folks, Brandon Villarolo here with another episode of The Register Kettle and this week there is one obvious story that stuck out to all of us to discuss and that was the leak of A Clawed Code's source code by Anthropic. With me to discuss this mess of a story is senior reporter Tom Claiborne and cybersecurity editor Jessica Lyons. Thanks to both of you for joining me this week. Yeah. Yeah, so I guess all three of us managed to get some mileage out of this one. I wrote the original news story. Tom Claburn: Yeah, thank you. Jessica Lyons: Good to be here. Brandon: kind of breaking the news. Tom, you've been looking at what the source code leak actually kind of revealed to us. And Jess, you've been kind of all over the security part. But before we get into these different aspects of the story, I want to kind of break down what exactly happened. ⁓ Tom, do you kind of want to walk us through what triggered this whole incident? Tom Claburn: Yeah. Yeah. So ⁓ at the end of March, Anthropic pushed, so I think was Claude code 2.1.88 to NPM. And it included a source file map, which is a file that's produced when you're... ⁓ from the links, sort of minified production code, the original readable source, and you're not really supposed to check that in to a public repository. You know, it's not supposed to ship users. And the source file map ⁓ actually pointed to a Brandon: Alright. Tom Claburn: Cloudflare bucket where the source lived and it was available without any kind of authentication or password or anything. And so someone was able to copy the file and ⁓ then the source was all publicly available. And I mean, the thing is that the reverse engineering community has already sort of seen this in minified form. They've already looked at the binary and figured out a lot of this stuff, but the actual availability of the source made it it a lot easier to read. mean, there were more comments and it was just a lot easier for people to decipher. And right away, people actually just dumped the source code into their LLMs and said, analyze this, and then produced reports about this. Brandon: Yeah, yeah, I understand. ⁓ So this is not the first exposure, right, of Claude's code, but it's probably the most thorough and complete, I imagine, right? Tom Claburn: Yeah, I mean it Yeah, exactly. mean, there had been a similar incident that had happened, I think it was February 2025, when another developer found, yet again, a source map file. I think it didn't have as broad exposure. was taken down, I think, like two hours later by Anthrobic. They issued a, whatever it was, a take down request, and I think it was removed. I'm not sure how widely it spread, it's certainly embarrassing to the company. I don't know how much damage it does because these things change all the time and as long as it doesn't expose anything that's of unrepairable or undoable, they can quickly sort of adapt. But it does tell people a lot about what's going on at the company and their current plans. Brandon: Yeah, absolutely. ⁓ But before we get more kind of into what we actually learned in that source code, I just want to shift gears a little bit really quick to talk about the security aspect. ⁓ Jess, I know you've been kind of keeping an eye out the fallout from this. But I mean, before we get even talk about that, this just seems like pretty sloppy, right? Especially like Tom, as you were saying, this isn't the first time that a source map file got left in an anthropic repository. And that seems pretty sloppy from a security perspective. I mean, what does this say about things over at Anthropic Jazz? Jessica Lyons: Well, I think it really highlights the human element in security. even really good developers can make errors or forget to check their build pipelines and expose what you would think would be a private file ⁓ in a public repository. So it's embarrassing from a security perspective for the company. It also shows how humans typically are the weakest link when it comes to security. And that's something that people in the industry have been talking about for a long time. this, in this, right, right, definitely. We see that in a lot of the breaches. wasn't, somebody didn't hack in and steal source code. This wasn't North Korea. This was just a developer at the company who made a mistake. And I think it's also kind of ironic considering we've been talking about Brandon: ⁓ hey, years of hammering that point home, right? Jessica Lyons: AI security risks so much lately, that was a huge topic at RSAC. Everybody is trying to poke holes in the different models and see what they can trick them into doing. And here it's a human mistake. It's not anything that the AI did. It's a human mistake. So you have those issues facing the company, but then you also have other bad guys who are using the leak as a lure to try to download and infect people's computers with malware. We saw that on GitHub, a developer published a malicious repo and it used the cloud code source material as the lure to get people to try to download this. But then of course, when they went to go download this repository, they instead are downloading VDAR, which is an InfoStealer and Go Socks. And that is malware that's used to proxy network traffic. So. These are not things you want on your computer. And we don't know how many downloads, but we know that this particular developer's trojanized repositories had 793 forks and it had 564 stars. So it's safe to assume, I think that hundreds of people probably downloaded malware onto their computer thinking that they're downloading the whole leaked source code. Brandon: ⁓ wow. Yeah, at least a several hundred, right? I mean, speaking of cloned repos and malicious repos, that's an obvious way. Any cyber criminal who is keeping track of current events is going to immediately pivot to start trying to take advantage of that, obviously. But I know that Anthropic has been going behind itself, like trying to, or not behind itself, but kind of going behind itself, guess, trying to close these repos up, take an issue and take down requests to GitHub and stuff. ⁓ Jessica Lyons: Right. Right. Brandon: I don't know, do you know if something like that would get caught up in the mix? would that, you know, because it's claiming to be a legit repo of this, I mean, hopefully that would get caught up in the mix. I don't know, Jess, if you have any. Jessica Lyons: It was, it was still up at publication time. Both of the developers, yes. So I haven't checked again today, but at time of publication of that story, they were still public. Brandon: Okay. Okay, wow. Oh, huh. guess I would have to check too. I haven't checked to see how many repos have been taken down or what have you. But I guess also one of the things I also heard was the way that Anthropic filed this takedown request has led GitHub to removing a whole bunch of legitimate forks of one of its own cloud coder repos. So a of of a mess there too on that side, right? With people with legitimate repositories saying, excuse me, what happened to my? legit Cloud Code repo, but yet this malware repo is still out there. So hopefully that gets fixed. ⁓ So Tom, as we mentioned, this wasn't the first exposure of Cloud Code's source. And I guess coincidentally or not, you were also talking to someone else about this very topic before ⁓ this whole kerfuffle. ⁓ So I guess let's dive a bit more into what previous versions of exposed or leaked or reverse engineered Cloud Code have looked like. versus what was found in this one. ⁓ What's new? What are we learning ⁓ from this actual honest to God plain text source code leak? Tom Claburn: Yeah. Well, the one thing that stood out is ⁓ the researcher I was talking to had pointed out something in an old version of the source that had been refers engineered that didn't have a lot of explanation. It was referred to in the code as I think it was ⁓ Mellon, you know, Mellon mode. And it wasn't explained, but the researcher shared some code that showed that this existed, but there wasn't any documentation of it. And that was absent from the current from the latest release. So Brandon: Right, yeah. Tom Claburn: We don't really know what it was. It was removed and, you know, efforts to get information from... Anthropic didn't really, ⁓ they say that, they work on a lot of things and some, every project makes it into production. So it's really not clear what that was. mean, our speculation was that it was some kind of ⁓ headless ⁓ way of operating. So it operated in the background. And we know that they're developing something similar to that. There's a ⁓ project that was referred to in the current service called, I think it was called Kyros that is a daemon that works in the background and is intended to Brandon: Mm-hmm. Tom Claburn: sort of do work while you're not at the keyboard. ⁓ It's entirely possible. They may have just given it another name. Yeah, exactly. mean, ⁓ but you know, there's all sorts of stuff in the source code that points to their future plans ⁓ and... Brandon: So maybe, maybe melon mode has evolved into Kairos, or maybe... Maybe the melon mode mystery will never be resolved. Jessica Lyons: Thank Tom Claburn: It's interesting from that point of view. mean, I think they lose a little bit of ability to surprise the market based on that. And they're probably scrambling to try and change things up or accelerate releases because now everybody knows that they're working on it. so they're open AI and whoever else will be adjusting their release schedules if that's the case. Brandon: Sure, yeah, mean, Anthropic has kind of been, I don't want to say necessarily like, you know, stealing OpenAI's throne, but you know, from what I've heard, like their use is climbing while OpenAI's is kind of declining. I feel like the stars are starting to align in Anthropic's favor in a lot of ways, and they're, and OpenAI is kind of losing out in that scenario, and I'm sure this is not going to be great for OpenAI. Tom Claburn: Yeah. And that's, and that's been kind of a problem for them because they've had real capacity issues. mean, with their sort of sudden surge of popularity, I think they weren't prepared for it. And so there was all this demand that they can't supply. They just don't have the cloud infrastructure in place to support it all. And so they've had to do things like adjust the, you know, the limits and then, you know, they had a, and they, you know, they kind of have a loss leader system where they just give people ⁓ who buy subscriptions, you know, sort of. Brandon: Mm-hmm. Right. Tom Claburn: you know, unmetered access. mean, it is metered behind the scene, but, basically you have it for a certain period of time and, um, you don't pay the, like a per token cost. You pay a subscription rate and if you buy it per token, ends up being pretty expensive. everyone who wants to try and save a little money does the subscription if, know, and they just can't support all that. And, you know, I know of developers personally who say that they have, you know, spent, you know, hundreds of thousands of dollars on tokens, you know, using a, you know, $200. Brandon: Mm-hmm. Tom Claburn: month subscription. So clearly that's not something that is sustainable. of course, the math is exact. Anthropic doesn't publish their actual token costs. You have no idea how much this is really costing them. But suffice to say, they are taking a loss on heavy users. Brandon: Yeah. Of course, right. Right. But you know, ⁓ one of the things that heavy use in Anthropics' favor has come from, I feel like, has been the sort of privacy concerns and sort of ethical issues around OpenAI. But I think one of the stories you published this week about sort some of the relations in the source code might give people a reason who are considering jumping ship from OpenAI to Anthropic because of privacy concerns. ⁓ They might want to rethink that, right, from what one of your stories seemed to suppose. Tom Claburn: Yeah. Yeah, I mean. Well, I mean, all of these, I mean, just in general, all of these models are just terrible for privacy because all of these, ⁓ you know, stuff gets hoovered up by them and they have all these policies that kind of explain, well, you know, we don't do data retention or whatever. And I think that's, you know, reasonable to say is, know, for like government and enterprise customers, I think they've pretty much established that. they talk about, know, that they don't have, they don't see that that's between the cloud provider and, the government or whatever. they're running, if you're running it on bedrock, it's the rest of, of the world. I mean, all the enterprises. that who are just running it commercially with Anthropic or the consumers and individual developers. All that stuff is exposed to the models and there's just a lot of stuff going into their systems and you have very little ability to control what. exits your system. you know, maybe Anthropic is, you know, totally not doing anything with that information, but having it out there is a risk because, you they could be hacked and ⁓ the kinds of systems that they're working on where they're talking about this thing called Dream Mode, which is essentially like a more limited version of Microsoft Recall. mean, whereas Recall sort of would remember your entire hard disk, Dream Mode would, you know, is focused on specific projects, but it would remember them and, you know, sort of process stuff in the background. Brandon: Mm-hmm. Tom Claburn: They already store a ton of stuff locally and ⁓ in JSONL files and having that stuff available and unprotected is just not great. mean, like one of the developers I talked to, ⁓ you know, talked about losing a bunch of material and then being able to recover it from these JSONL files. ⁓ And I think people just don't expect that there's, you know, all this extra sort of record there that's in, you know, clear texts and plain view for anyone who knows where to grab it. Brandon: Yeah. Yeah, I mean, the fact that your AI could actually recover files for you. Like, ⁓ hey, I think you misplaced that. me give you the copy that I surreptitiously made without asking for any permission. You know, it's just like, that's. Tom Claburn: Right. yeah, if they've seen your, know, and these things have probably seen, you know, environmental variables and sensitive API keys and things like that. And so all it takes is, you know, one, you know, clever kind of malware to just read all that stuff to, to expose a lot of data you thought might've not been exposed. Brandon: Mm-hmm. Jess, are there any signs in the security world that this kind of malware is out there? Or that these sorts of capabilities exist? Should people be worried about this? Or is this an edge case right now? Jessica Lyons: Not that I haven't heard of any specific malware examples because right now it's pretty easy to trick the AIs into giving you information they're not supposed to without using fancy malware. You can just ask them and different workarounds where they're likely to do that. it's exactly right. Brandon: The old human factor, right? Again, it work, right? Let's rely on the fact that the AI is essentially programmed to behave like a human and con it into giving us information. Jessica Lyons: Right, the two non-deterministic things in your environment are your AI and your humans. so those are probably the most easily manipulated. So criminals don't really need to take that extra step yet. I'm sure that some people are working on it. But why spend that effort and money right now if you don't need to, if there's easier ways to get the data you're trying to steal? Brandon: Mm-hmm. Yeah, guess it remains to be seen. If a proof of concept can be discovered, I'm sure someone will come out with one relatively shortly that would be able to exploit a lot of that data collection that you've been talking about, Tom, especially if it's being stored locally in plain text. Someone's got to figure out a way to get a local LLM to spit that back out ⁓ to them. Is there anything else we didn't know about ⁓ Cloud Code's design or anything that we learned from this exposure, Tom? I think maybe something else about safety rule? Jessica Lyons: Right. Tom Claburn: Yeah, well, one of the, one of the sort of embarrassing things was that they, they had a directive in there to hide the fact that, ⁓ Claude code had authored code contributions. like anthropic people who have, you who are submitting to, ⁓ you know, ⁓ public repositories, you know, often if you make up, if you have an AI model, make a pull request, it'll by default say authored by flood code or whatever. There's a ⁓ suppression line in there that, you know, they're not supposed to reveal that they are, you that this was done with AI. And that's kind of, you know, a little bit, a little bit disappointing. mean, if you're going to sell this stuff, you know, being embarrassed about it a little bit, you know, I don't know. And it's, there are lot of Brandon: Yeah. Yeah, it's supposed to be so good, right? Yeah. Tom Claburn: Yeah, there are a lot of, you know, open source projects that have just taken a stance. Like we can't deal with open, you know, with AI contributions for various reasons, either the volume or the quality or whatever. Brandon: Yeah, I mean, that's just, I don't know. It almost feels like it's a recipe for, ⁓ for, know, illegitimate bad code commits that are, that are not being properly reviewed because someone doesn't know they're written by an AI. I mean, hopefully Claude code isn't that bad, right? But if they're, if they're hiding the fact that it's making code contributions, that's not exactly encouraging. ⁓ What about the, the safety rule bypasses? There was something to do with that being found in there too, right? Tom Claburn: ⁓ yeah, they had, ⁓ there was another thing that came up out of the source code where they had, they were checking, ⁓ they would check ⁓ stuff entered in the command line for, know, to make sure the commands were safe. you you don't want your model to, ⁓ you know, make use of the curl tool to like make remote network calls. So you might put, disallow that by putting a rule into the system. And they, ⁓ capped that check at 50 because it was just ⁓ it would have been a lot of I think for whatever reason was a lot of work for them but ⁓ Israeli company called Adversa just spotted the that comment in there and just said, well, it's really easy then to overload that by creating a sort of compound command that has a bunch of no ops and the curl command that will, then it will, you know, that the system will execute despite the allow rule. And it's not a total, um, sort of, um, takeover right away because the machine will still revert to his default behavior and ask you, do you want to use that tool? But so many people use these things in ways that where they just sort of auto allow stuff or they just click yes, yes, yes. You know, after doing it. 20 times that it opens the system up. And they patch that after that was disclosed. But again, it's not a great ⁓ approach. Brandon: They increased the max to 55. Yeah, we patched it. We patched it. ⁓ Yeah, I don't know. It's interesting. It kind of makes you realize that even a company this high profile is doing some sort of sloppy things with their code that are enabling people to potentially take advantage. I mean, it really makes me wonder what we're going to see come out of this in the coming weeks. Like you said, Tom, it's entirely possible that this will prompt them to Tom Claburn: Yeah, yeah, exactly. Yes. Brandon: accelerate some timelines and even possibly rewrite some codes so that this stuff is less relevant. ⁓ But yeah, we'll see what comes of it in the next few weeks. don't know, does anyone, do you guys particularly need any zero days or anything being abused coming out of this? Tom Claburn: ⁓ the one that may have impact is there was reference in the code to the use of anti-distillation tools. Model distillation is basically just a way of copying the model by querying it a lot. ⁓ we're storing the senses. So ⁓ it out that ⁓ they were injecting a bunch of fake tool data as sort of a form of competitive data poisoning so that people would have a harder time copying the models. And now that that's disclosed, they'll have to think of another way to do that because I think people who are trying to distill their models, i.e. model makers elsewhere, will be wise to that. Brandon: bit rich honestly coming from an AI company. Hey we distilled, we built these models using all this public information that people didn't want us to copy but ⁓ can you do us a favor and not copy our model? Tom Claburn: Yeah. Jessica Lyons: Well, yeah, think along those lines too, mean, that's something that's also been brought up is the potential for national security risks due to this. Like you just said, distillation of models has been an issue that Google and Anthropic have talked about. They've specifically accused China of doing this to train its model like DeepSeq. So we did see a house congressman, ⁓ democratic congressman, Josh. Gottheimer from New Jersey earlier this week sending a letter to Anthropic saying, hey, this poses a national security risk. He said, Claude is part of our national security operations. If replicated, we sacrifice the competitive edge that we've worked to maintain. And this again points back to the whole issue of stealing people's models by distilling them. Brandon: Yeah, I wonder if this ⁓ gives fuel to the Pentagon's fire that anthropic is a supply chain risk. They may have been before this, but I feel like letting the model get out through, I guess what, the second instance of a bad map file or a map file getting left in might give the Pentagon fuel to say, hey, look, we're not saying this just because we're petty and we are seeking revenge on this company that told us we can't misuse its model. They're actually doing some... of bad stuff in terms of being good custodians. I guess closing thoughts ⁓ to run us out here, Anthropix has been really pushing for an IPO lately. They're going hard at this. I mean, is this the kind of thing that could harm it, along with maybe the over ⁓ pull request or the overly kind of... ⁓ Oh gosh, we're out of blank. The overly judicious use of takedown requests on GitHub. Is that the kind of thing that, that, you know, that plus this plus the Pentagon stuff, is this gonna harm their IPO chances? Tom Claburn: I mean, I think it should make people think. ⁓ mean, certainly the capacity problems are kind of the, think the bigger financial one. ⁓ And if they can't figure out a way to offer the service at a cost that makes, that works from their revenue perspective, ⁓ that's something investors are going to be aware of. The leak, remember, cloud code is really just sort of a front end harness on access to the actual model. And so the, the It remains to be seen, I think, how much value you can really put in a front end client before people just say, you we don't need to use Cloud Code. We can just use somebody else's service, like Cursor, and then maybe pay through the API costs and get a better experience without having to worry about what kind of telemetry is in the client and things like that. So I don't know. mean, I think it'll make people look more closely. Brandon: Mm-hmm. Yeah, Jess, any thoughts? Jessica Lyons: I agree with Tom. think that the capacity issue is probably the bigger problem, but I think anything like this would tend to make investors a bit jumpy. I don't know. We might see an IPO pushed out a bit. But again, we'll all keep an eye on that, I guess, and see what happens. Brandon: Yeah, we'll see. guess just more growing pains in the AI industry. I just feel like in general there's signs that foundations are starting to crack and problems are starting to appear in a lot of these different companies. I guess this is just another example of a bit of a mess, for one, to pull themselves out of. Well, hey, thanks for joining me this week to talk about this issue. And whatever breaks next week, we'll be here to talk about it on The Kettle. Tom Claburn: Thank you. Jessica Lyons: See ya.