Ravid Shwartz-Ziv: Hi everyone and welcome back to the Information Battleneck. And today we are talking with Nathan Lambert. He's he was a lead the post-training at the Homo at the LN Institute for AI, the author of the RAGF book, and the writer behind Interconnects. Hi, thanks for joining us. Nathan Lambert: Hey, thanks for having me. Fun to see some more researchers get into this space. It takes a long time to compound. Ravid Shwartz-Ziv: We are trying, we are trying. Not like you, but yeah, we are trying as our best. Yeah. Maybe you can give us some tips after it. so I think I want to start with you you were really vocal about open open models versus closed ones. So I want to hear your current takeaway. Like wha what do you think we'll we will see? Do you think like the future is open source models or Maybe like close ones. Nathan Lambert: I think in many ways there's a default where people want to use open models, but it's very unlikely that the capabilities keep up. I think the the primary reason is that the people building the best models there's in order to get like economic feedback loops as being a platform, the open model builders have a very indirect capture of of value from their models back to their business versus open AI and Anthropic have the fastest growing businesses of all time because they're integrated from model to harness a product. And this is very hard to overcome. And at the same time, I thought that the effects of this would be seen sooner, which is like the open models have keep up to a remarkable degree. I think a lot of this is just psychology and competition of the fact that you once you know a model is buildable, it's a lot easier to chase it down. And like obviously distillation and some other things help. But I think we'll be seeing some more strong US models soon, potentially even Before this is actually released. I think there's like it's just a wild dynamic of how many people can build strong models. But the value capture is the fundamental thing. And the closed market is very well suited to like sell super high margin products to individuals. And there's this amorphous open model economy that is brewing with the fine-tuning services, the inference providers, the open model companies that don't necessarily have a great way to monetize. They're trying to like serve all the other issues. use cases when I think of like business automations and things with private data and it just takes way longer for those deployments to grow than every software engineer and knowledge work is like cloud code is my number one tool. I can never live without it. It's like I setting everything up takes a lot longer and potentially the open market could be bigger. But I think without this kind of centralization of feedback from profits to reinvest in the models, it's like really tricky to know exactly what's gonna happen because Like this is how big companies work so well for so long. Like Google went through this, which is like they made a new business model and they made so much money they could invest in infrastructure. And I think the bad advantage for open AI and Enthropic is gonna be hard to overcome. And that's and I think it's like a a scary world where AI, if you think it's gonna be as powerful as it can be, to only have two people that make all most of the decisions and control who has access to it. Like That's mostly what I'm trying to hedge against and make a future that things are more diffuse where more people can have access and there's kind of more diversity of AI and more people can understand what's happening. So I th 'cause I think that'll be safer on net. And that explains a lot of my various ebbs and flows of what some would call activism, some would call promotion of open models over the last few years. Ravid Shwartz-Ziv: So but w what is the solution? Like at the end, if if if open open models is not a solution, what do you think is the solution? like the optimistic solution will look like. Nathan Lambert: To which problem? Ravid Shwartz-Ziv: to the problem of like that open AI and entr an entropic just will increase the gap. Nathan Lambert: There's a potential s it's some of it is still unlike unknown epistemics of how AI will unfold. If the AI boom is just so big that all of these inference and fine-tuning companies become big established successful businesses, then actually it might just work out. So it's like if the AI boom is so so big that there's so many verticals to capture, there will be many companies. And a lot of companies are by their nature incentivized to want to support open the In that world it kind of works out. But if it's like not this explosive market, I think it's it's really hard to overcome these economics in the two to five year range. I think in the immediate range there's so much fundraising and very few people are training fable size models. I think none of the open weight models are scaling training investment to match that scale. Even they will there will be a lot of hype online that's like, ZAI and Kimmy are gonna release a mythoscale three trillion parameter model. I would still say that that's I don't know, like two to five ten X less pre training compute or something than Anthropic likely spent on on this model. So it's a there's a there's a strong incentive to hype them. But I think like that's the point where we start to see a bit more bifurcation in the ecosystem. And I i it's hard. I think a cons like consortiums are possible, but those would also take years to play out. So it's like it's it's really in this medium term that we can I think more of this will be clear. Allen Roush: So, I you know, I I think that some of the main sources or flywheels of data that that as you observed earlier, you know, allow anthropic and open AI to pull ahead of everybody is codex and clawed code. but now we're seeing Chinese companies releasing their own harnesses, as well as, you know, open code powering big banana as they call it, it's like top four models or whatever that you get for free are all Chinese models, and they're very open about, yes, we are training on your usage. And so do you do you think that China will be unable to establish the data flywheel needed to to kind of stay in that, let's say, seven to ten months gap that the anthropic folks themselves regard to be the the current difference? Right? You seem to think it's growing. Nathan Lambert: It's there's a lot of like noisy factors to consider. Like when I went to China in April, there was like almost no data industry. Most of the labs were like, we'd build our environments in-house and things like this. And we're starting I'm starting to hear more cracks of like, they're reaching out to X startup for environments, them being some Chinese lab to some US environment building lab. And there's a lot of rumors of like exclusivity clauses where Like some people are convinced that it's like anthropic buy as an environment and then they have a three month exclusivity window and then the price drops five X and the Chinese labs buy it. And a lot of this Rumor mill is sourced of types of truth, but we don't know the exact dynamic of the these type of training environments. But I think of this type of training d data being much more where the leverage is than the actual user data. Like there's some in the user data. But if you're like on a critical path and the pace of progress is so high, that's normally because like you know exactly what you need to train on and it's buying this environment for ten million dollars from five startups that that makes your model better in two months. I think it's like it seems like the training dynamic is much more like that, where it's like, Like there's a lot of low-hanging fruit. You go and you acquire a bunch of things that are pretty close to it, and model gets better. Whereas a a bit longer term cloud code and these harnesses might look something like what Cursor is doing, where they can feed the real data in kind of into this slow RL loop that makes it better. But I think like I don't think these frontier models are just like lacking for things to do. Like the next fable's probably mostly done training and same goes with all these other models. And they're not waiting around for the human data to do this. And a lot of the enterprises have zero data retention. But and I like I think that's fairly similar, but across all the labs in terms of the like very high-end knowledge work stuff that's happening right now. They're probably In the past, like R L H F eras, there was definitely like we'll fine tune on some user feedback data and make the model slightly better over time. But I just think the pace of change is super high, which is The models seem so much smarter, but they also like are a little bit cooked, which I think is just a sign that the the pipelines are so rich, they really want to get the model out, but they're not so developed that all the like trade offs are ameliorated. So you get these like w super powerful models with like really weird quirks. Like GBT five point six, I don't know. I shouldn't remember what it is, but it's like had some b has not as like for example it'd be like GPT 5.6 has some weird like git problem that it used to not have and it's like why does this have this like git problem again? Or like people online like Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: the extreme version is complaining about 5.6 deleting their whole database, which is like I don't like it's not a good example, but it's just like the model's like, it's better, but like what is this weird behavior I to deal with? Like this is because the mod the labs are just churning out new recipes. Allen Roush: yeah, like like like like the goblins thing that they had a little while ago. Yeah, and and I I do want to ask, so so I do th I I'm convinced that sovereign compute and basically that that every country that And you know, even if they're two or three years behind on foundation model development, they have a national security reason to kind of keep it up. So I always use this as an example, and I feel bad the Kohere folks maybe don't like it, but Kohere, right, in Canada, right? They're, you know, at least probably more like two or three years behind state of the art now with whatever their best Command R plus model is. Maybe they've pivoted exclusively to re-ranking. But the point being that nobody else in Canada makes it, and America could just turn off access to any of their best closed source models. And possibly China, you know, there's I've been hearing stuff about China also possibly very much restricting access to models. So if we go into this scenario, do you then think that there's like national security reasons for everybody to develop models, even if they're pretty bad? Nathan Lambert: Yeah. I th I think so. And it's like the rich countries will actually not make that bad models because the the entry point is like orders of a hundred million dollars, which for rich countries in a national security priority is like still pretty cheap. It's just like more of a political cost to like really push some of your best researchers to be like, Hey, you gotta work on this. And it's like how do you navigate that? Which is like a little bit more tricky than than the money at this point. And I think there's a subtle interplay between this sovereign AI and open models thing where the countries want control and you like only can have control if you could kind of distribute the weights and they I think there's a incentive where these company countries are gonna want to work with domestic industry and be able to share these weights in a fluid way and more of the like off frontier weights you'll just o release openly 'cause the countries will go through the same path that the Chinese labs did, which is like no one will care if our s about our AI unless they could poke around with it and try it. Cause nobody cares about the twelfth rank model being an API that you pay for it's like nobody's gonna look at it and that's how the Chinese labs were like we have to just release these openly so the Bay Area people talk about it and I think there will be more sovereigns like this and I I mean it seems likely that Cohere and Mistral and them are taking the path of getting like getting business and funding through this. I mean even Reflection pretty much seems to be doing this which is like will be the lab that has the most time and willingness to cozy up to the government and talk about their sovereign AI story. I don't I don't necess it's like not a per business I'm personally interested in, but I think there's a lot of success to be had there. Allen Roush: And then and then j just maybe probably the meat and potatoes question on and and to be clear, I feel almost exactly like you about open source models, you know, information wants to be free, Aaron Schwartz kind of seeing the the modern crusade for open models in a similar vein to previous Crusades for open information. But what is the answer to okay, GLM 5.2 comes out? Philip Emmanuel Wideman or whatever, you know, the heretic creator goes and his code gets used to obliterate it, orthogonalize it, you know, to to uncensor the model. or or maybe, you know, now it's GLM five point five or something, like a fable class version, and suddenly, you know, it really does become a lot easier for lone wolf terrorists to like manufacture bioweapons or whatever the scenario is. Like Do you do you do you believe that we have, you know, a way to stop this or or mitigate this? Or like what's the answer to this problem? Nathan Lambert: Well we have to look at bottlenecks that aren't the model because like there's so much access to AI and as AI proliferates it's gonna not like AI access is going up and up and it's like increasingly hard to restrict that to people. And any sort of policy action that you can take on open models right now is not gonna stop these bad actors from doing this. For one, it's like I'm always amazed that the story doesn't get more coverage, but it's like when Mythos was in super super secret like hundred companies, there's like a random Discord server that had access to Mythos and was playing with it because they just like got the credentials of somebody at a company and were they were a contractor of it and they like had mythos access on their Discord server. Which is like there's still like very low hanging fruit like that that isn't addressed, that's much easier to get access to than manipulate a one trillion parameter model. But I think it's like in we just need way more independent capacity and measurement. Because even the claims of like the cyber risk and all of this are not well documented publicly and directionally, there's definitely some meaningful risk there, but it's hard to know the magnitude without independent entities. looking at it. That things like bio, it's like you need to have other bottlenecks and like bio has to interface with the physical world. And it's almost always gonna be easier to bottleneck the physical world than bottleneck some digital product. Because I think most most of the restrictions are just like anti-good actors. And the the bad actor prevention requires international coordination. It's just like the US and China would need to agree on what models cannot be released and things. And I just think it's like gonna be this messy middle of you need to have very what is the big picture, very rapid institutional and social change around AI. Which is why like I personally tried to get the word out. independently. I think this is like an existential problem and the big tech has backed or like backed themselves into a corner, which we're seeing with the data centers as the current issue. But there's gonna be more I think there's gonna be more severe challenges socially in the next two to five years. Whether it's like data centers with the first flare up, there will probably be something with jobs and Ravid Shwartz-Ziv: Yeah. Nathan Lambert: maybe something with more extreme risks around bio or cyber Ravid Shwartz-Ziv: So it looks at like Entropic is like Entropic and Open AI like has like kind of like mm like opposite approach, right? Like entropic basically say, yeah, like it's very dangerous, like we can't like right, release metals and like we we put a lot of garlas on the top of it, like both like legitimate ones and also like yeah, you can't develop the the ML infrastructure to to compete against us. versus open eye that are like more let's say like more like open about it and like say that the they are less try to to to scare you that the their model will control the world. what do you think is the is a better approach? Nathan Lambert: Well, I think it like anthropic is just too far along in their safety id ideology and it's a superpower for their business, but it's definitely an ideology that like that's just what they're gonna keep doing. Like it's how their whole company and leadership is formed and they're the most successful business ever through that culture. So I don't see it going away anytime soon. And opening eye historically has been much more fluid in culture and much more chaotic and internally disruptive and I see them they'll be like a mirror to anthropic stronger culture if it serves them. Like right now it serves them. And I think that's good. But I think like anthropic is much more of a defined entity in this space where open AI kind Ravid Shwartz-Ziv: Hm. Nathan Lambert: of will move back and forth. I mean, Greg Brockman donating twenty-five million to the Trump campaign seems like like I mean, I I I don't endorse this personally, but as a business move it seems like a total genius play for him. And I'm like This is just like a crazy world to live in, but they that's like their short term plan seems to be that and it's working out for them and figure out the other details when they have low hanging fruit. Allen Roush: Well but but isn't the government I mean I claim Democrats would be probably regulating even faster, so doesn't it make sense to cozy up with them no matter what? And if anthropic is going to stick to their guns even when it gets them in trouble, doesn't that imply possibly the death of the company due to r their idealism go you know? Nathan Lambert: Well, Anthropic's ideology is that they want to be nationalized, but they want to be nationalized in a democratic, insane way, which the US political system is unfortunately going through an by its nature unstable era, and we don't know exactly how long it'll last. So it's just kinda like their ideology has set them up for this a bit. They've talked about nationalization for a long time. And I do think there's a there's a portion of you have to play with who is in power and Anthropic and open AI I think are like pretty clearly on different sides of this by a substantial margin. I think you could do both could be more neutral and still engage. Ravid Shwartz-Ziv: But do you think like Entropic really believes in all these like I don't know risk issues and and that they basically save the world or this is just I don't know PR or like just to try to to prevent from a competition? Nathan Lambert: I think they fully believe it, but it's easier to believe it when it supports your business. So it's like a lot easier to change their mind if it's like you're burning tens of billions of revenue to the ground based on your ideological position. But it's like creating the mass market share, standing up to the government is politically popular with their user base. Like it's helping them. Allen Roush: But but okay, so do you think anthropic cares about some of the other negative externalities? So for example, rampant linguistic colonialism. everybody now talks using phrases of the form it's not x, it's y, con c I think it's called conditional negation, because they're reading it from LLMs and they're overusing dash and semicolons. So d d you know, in in terms of you know, risk to like diversity. of humanity at large. and and, you know, e not even just cognitive atrophy here, but like, do you think that they're mindful of these? Because it's really easy to take this opinion of, we don't want the bioweapons, but what if you mind virus everybody? Nathan Lambert: There are people working there that think about this, but that they're not on the decision making radar. The decision making radar is is the government gonna ban us or do we have an existential risk in our next remodel release? And I actually think those things like I mean language changes over time. We're in a particularly fast rate of change. But I'm not that concerned about language change and it's like it's gonna slopify the whole internet and you have to make new structures for how you how content gets filtered and blah blah blah. But that stuff I think is like much easier to fix. I think like meta is gonna solve that immediately because their ad business is on the line and they're just gonna get rid of slop if people don't like it. But a lot of people like slop, so it's fine and engagement is up and like that equilibrium will work out. I think it's hard like where you want to look at that right now is the impact on like it like adolescent development and all the other like adult internet economy stuff is fine because that's been hor kind of horrible for our brains for a long time. And I don't think it's gonna be like that different for adults. They we just need to figure it out. Ravid Shwartz-Ziv: Do you think like Enthropic like is happy with the decision with Mythos because it looks right like the government banned it for a few months. and and now they extend like the the the ability to to use like the the paid plans, right? every week. Do you think they're happy because they had some like advantage and now almost it's gone. Nathan Lambert: It's very messy. So like the perpetual delays type thing. So this is referring to like Claude Fable Five. We keep getting it a week longer than they say. And this paints Ravid Shwartz-Ziv: Yeah. Nathan Lambert: to some weird internal like leadership chaos that we don't normally see. Like it's a it's a much more open AI coded type of thing. But now like open AI is kind of easy of like you can use our stuff, we'll re open source our harness. You could use our model in claude code. And I think this is a lot of this is just like they OpenAI don't really care about that type of strategy. They're just kind of like throwing shit at the wall and they're like, well, there's some low-hanging fruit. We'll take that win. We have the compute to burn. Like the change of the ChatGPT app is crazy. Like, OpenAI is kind of just like out there doing shit. Like, this is the Mac app that used to be ChatGPT and now it's just Codex on the ChatGPT app. And it's like Doc, didn't you have users that want to chat with this Ravid Shwartz-Ziv: Ha ha ha. Nathan Lambert: model? Like like I don't think if I was a like my par like I don't think my parents use a chat GPT Mac app, but if they did they'd be like, what the fuck is this? And like it's a Ravid Shwartz-Ziv: Mm. Nathan Lambert: very open AI, but like Anthropic always seems a bit more buttoned up. So I I I think this is like the type it just looks like a crack in terms of decision making, but I don't know exactly how to interpret it. And it's like the question is like, are they happy with how mythos got treated? I think they're happy on the increased regulation and the press that shows them as having the the best model. So like initially, in some ways it was good for them. I think they hate the government pretty clearly. I think they've they've wanted to make some amount of political wins, which I don't think is the best business decision. it's a pretty bad thing to like like to do is one of the biggest companies in the US to try to stick it to the government. That was a mistake. But them being s labeled as having the best model, that's like a huge branding asset. So it's it's it's too complicated to be like, are they happy with the situation? I think like it's they can't on net they can't be happy. It's like it's been so painful to go through the whole process. Ravid Shwartz-Ziv: Yeah. Yeah, but like do do you believe so like th they they tried to release like several like apps or attempts, right? Like there was like the design cloud and like now there is like science c cloud or whatever. Like do you think or do you believe it looked at like outside of this wall, like everyone thinks that like yeah, if Entropic now released like a app in your area, like that's it, you're done. Do you believe the that this is the case or? Or not. Nathan Lambert: It's a very Bay Area mentality. So I actually think most people don't think like that. And if Ravid Shwartz-Ziv: Okay. Nathan Lambert: people will try the tool and use it, but it's like super San Francisco brain right now to be like once a Frontier Lab touches your domain, you're screwed. Where I think in reality the dominant thing is if you have users that actually use your product and you can leverage that. to make it sticky or underst better tune the model to their use cases or in the best case do the cursor thing. And it's like engagement is the king in AI right now. And if you have engagement, you'll likely figure out how to make money. Or if you have GPUs to sell Allen Roush: O okay, okay. I'm gonna I'm gonna challenge that because there are two companies that I think kind of c slightly go against this. One, Hugging Face is only worth, you know, what is Yeah. Well Nathan Lambert: Okay, you don't we don't need to go down the whole hugging face rabbit hole. They're intentionally independent and their intentionally independentness makes their revenue lower. Allen Roush: okay, l let's go deeper. civit.ai is worth like thirty million or less despite their user base because of what they're forced to do because of that user base. Nathan Lambert: I don't Okay, look yes, there are obviously counterexamples. The core thing is like if you're d making a product that drives cutting edge knowledge work and these things that are actually on the frontier of AI progress, then it'll monetize. Of course there's decades of AI startups that have users and make more money because they're designed for a previous era. And I would have liked Hugging Face to do the thing like what Reflection is saying to do, but the founders didn't want to bet the company on being able to build and monetize open models, which is a sign that they're happy with. their neutral, high relevance, high influence role, and they didn't want to risk something they really much like to ten or a hundred X company. Which like as somebody who wants to see more open models in the world, it's easy to criticize it. But also like as a human you can understand the decision in some capacity. Like they don't want to get involved inquired Allen Roush: Which makes bigger Nathan Lambert: by NVIDIA or Microsoft because then they're no longer viewed as independent. Allen Roush: Th this makes a big bloom that much more impressive than it ever got made. Nathan Lambert: Yeah, all these community projects are super impressive. Everybody does one of them and then they're like, I'll never do this again, it was too hard, and it's a miracle that it exists. Ravid Shwartz-Ziv: So so what is what what is your ta take on like the the common like Silicon Valley like product? Like AGI is coming and self improvement is coming and basically entropic and open AI will control everything. Nathan Lambert: pr well, i it's on its face, obviously lame, but also really sad because it constrains the talent flow a lot. So I'm I'm personally very sad about like the top the the incentive structure is for all the top individuals graduating with like PhDs and a lot of them have safety nets already from family and stuff, they're like all I can do is go join anthropic. And in the like long term, these are the people that you want to go and take wild bets on building new things and new types of things and that's what would make like an interesting diverse AI economy. And they're all just like, this is the next 18 months I need to make my few millions of dollars or now, or else I may never never may. And they end up getting sucked into these labs. And I think that I don't fault any of them. Like I've gone through some of this process. And it's like if I was a PhD student or advisor's like, you can graduate early and join anthropic. It's like, yeah I've rec I I would recommend a lot of people I mentor to do this. But it's just a not a great incentive system because it collapses the diversity of what people work on and the diverse like how much people can share ideas because there's so many you've surely have seen a lot of people just fall off the grid, which like, yeah, this friend that used to be like a vibrant open thinker and now they're like no a no comment person. You talk you ask them Ravid Shwartz-Ziv: Yeah, it it's Allen Roush: Okay. Nathan Lambert: and they're like no comment Allen Roush: So so so so do you do well, because the only people I'm hearing that would do anything to support this, and I don't support it just because I know what it would do to t you know, productive like like basically kneecap us, right? But Bernie Sanders and Elizabeth Warren are talking about antitrust and going after and breaking up open AI and anthropic. Do you support something like that as kind of a means to i you know, it's like sorry, person earning five million dollars to shut up, you're gonna go back to having to you know You know, do open more like better work for the world. Nathan Lambert: No, I think that's stupid. You have to create its financial incentives for people to do otherwise, which is like personally I'm invested in like trying to create like new AI nonprofits where I put that in quote because like the AI nonprofits are not like standard nonprofits, which is it's like a org with a ten to a hundred million dollar budget and I don't know, twenty ten to twenty employees, where normally a nonprofit at those scales have hundreds of employees. Ravid Shwartz-Ziv: Ha ha ha. Nathan Lambert: Where it's only nonprofit because that is the governance mechanism that is in law that can let people take tax deductible money and put it into your AI organization. Where in reality these like 51c3s are not good governance mechanisms for these orgs and like opening eye drama aside. Allen Roush: But but but but I want to point out a lot of people, Circa Rockefeller and some of these other, you know, Gilded Age people were probably saying the same thing when the government talked about antitrust on big oil and big railroads. Why why is this different? Like it's stupid Nathan Lambert: What do mean? I feel like that was Allen Roush: to break them up because they get efficiency gains from being centralized and winner take most. Like that might be true, but what if the public gain from from breaking them up outways? Right, like do you think that that's possible at all? Nathan Lambert: I feel like there's a lot of context that has not been stated as the assumption of this comparison, which makes it very messy. Because like modern antitrust law is about consumer harm. And I think that these AI companies are not talking about consumer harm. And that antitrust lever doesn't really apply to talent harm. Like we were talking about talent disparity and concentration of knowledge and like a chilling effect on the diffusion of the economy. And like antitrust is all about consumers. So I think this is like a really different situation. And I think at the time of the railroads all the antitrust law wasn't really formed. And I think there's like a whole bunch of financial Ravid Shwartz-Ziv: So Nathan Lambert: crises that need to happen before that could actually kick in. Ravid Shwartz-Ziv: So wh why all these like famous people are joining in a tropic? Do you think like probably they have enough money, right? Right, but like probably Nathan Lambert: They pay twice as much as anyone else. And it's fun. Like it's a the culture there is very open. So like you hear from people, they're onboard and they're like, I got my laptop, I ran the script, it c set everything up for me locally in the cloud, it worked perfectly. The Slack is fully open and so many super smart people work there and you're working on the most impactful technology of the time. And they're like, Yes, I can take the trade-off of like a lot. There's a lot of there are It's a self selecting nature 'cause they have this culture interview that you need to go through. And I'm sure a lot of people are able to fake it. Like you can fake a lot when you have a five to ten million dollar payout on the line. It's not that hard, but like like it's somewhat self selecting in that capacity. But I I really think that in the next year it's gonna start to Fade and then inevitably be leaks, and this open culture is not going to be able to last. So it's like Facebook went through this where everything used to be open and they had their decade of leaks, and now they're like any other big company. I think Anthropic inevitably will go through this, and it's painful. Like I'm not I'm not ex- It's not a fun thing for any company that is successful to go through, but there's a lot of people in tech who are like, I used to work at a PM at some like data app layer company, and OpenAI Anthropic offered me the 4X, I'll go. B a PM on their shit. Like like obviously, but like those people don't immediately care about AI safety in the same way that a lot of the researchers that they hire who have been slowly monitoring this for a long time actually care about. But I'm I'm impressed when I hear how open their Slack is. I'm like, this is wild. Allen Roush: Well well, how is it so open when when they're not allowed to be open with anybody at NURIPS, for example, or all the other conferences? Right. Nathan Lambert: It's a cultural strength. I'm baffled by it. It's crazy. I like I'm like a such a high amount of people that I've respected know very closely go there and I ask them something and they're like, no comment. Some of them are still off the record of like, yeah, obviously I can't say anything to you in public. This is like a if I'm having a very private conversation to them and they're like, I'm happy you're doing what you're doing. Like there are still people that are like this, but I the cultural strength and the financial squeeze that they have on people approaching IPO is really big and maybe after the IPO lockup periods, there's a lot of there's like a much more turbulent time for the company. Cause I I know tens of people that are gonna make tens of millions of dollars off of this IPO. And I think that a lot of them are gonna be like, I done that kind of I had to hold my nose a bit at the end there, but I got it done. It's just a lot of money. It's pretty like you guys could do like we could all do a lot of things if we were like in a position and all we had to do was hold it out. To see it out. Like I'm 100% sure if I was in that position, I would be totally silent and be like, Godspeed. Ravid Shwartz-Ziv: Like but I don't know, like a person like like like Karpatis, right? Like probably like he doesn't need the money, right? So like d this is only because like now like he he wants like infinite amount of of compute. Nathan Lambert: There's a lot of paths to it. So I think like yes, like it's there's a technological allure to it. Like you are going to have direct and direct marginal impact on this technology that is transformative and you will live through this transformation firsthand. I think like n like I remember this like Noam Noam Brown poking carpethee on this. And I think it's a super valid point. Which is like I tell people if they're gonna be at a frontier lab, you should probably just go to open AI or anthropic. Because the likes of like Reflection, Gemini, MAI, Meta, Meta maybe people are optimistic about it right now, but none of them are gonna be as close to the technology. Like the the scale, the incentives, the competition. It's like if I want to work at the pinnacle of the technology, I think those are the only places to work. And I think that's a valid reason as a scientist. Like this is what your trade is. Your trade is to build these things. And that is the best place to do it and the best time of your life to do it. So like Allen Roush: I I I I I I don't know. I claim that anything that's safety sensitive they will intentionally suppress. So for example, LLM sampling. tail free sampling was invented by the current head of mechanistic interpretability back in twenty nineteen, and that's a really good sampler because it uses a full distribution. They are still using top P and Top K because it suppresses high temperature, because high temperature leads to safety and alignment problems. If you're working on LLM sampling, I claim ThoughtWorks, where I'm at right now, is probably the the best lab in the world because you you will in be suppressed because anything that unlocks more creative outputs is dangerous by definition. Like do you disagree with such a view? Nathan Lambert: I don't, I would just work on something else that's closer to the model. Because I and I also think a lot of the work at the labs Allen Roush: Ha ha. Nathan Lambert: is really boring. So like a lot of researcher there's like Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: the two there's like for simple initial, there's like two tiers. The one tier is like maybe in post training it makes more sense. Pre training like scientific and really cool and really cool engineering problems, but post training there's like the type of people who are putting the recipe together. Which is like there's all these substream sub teams and data processes and compute and timeline of release that you have to solve this complex org problem and like do this like empirical research. But a lot of people that feed into this are like, you join and you're like for six months, you're gonna improve this benchmark. And you do some working with data vendors and you do some synthetic data and you compare it to other valves of the company. And like that job to me sounds so boring since I've done the like cool post training research direction thing. So like I have a much more intellectually stimulating job than I think I would at OpenAI or Anthropic. But most people I don't think have this optionality and I think it's really valid to be like, I'm just gonna take the job and figure it out later and just figure out what I'm gonna work on. And I don't have a good I don't have enough samples of like how that works out for people at OpenAI are anthropic. And I I just think it but it's like it's kind of like why I went to my PhD. Like I wasn't in AI. I like, I'm just gonna go see if I could get into AI and figure it out. If you're like, I'm just gonna go try to contribute to this like metaphorical God machine, like I'm not gonna blame somebody. Like it's it's a cool Ravid Shwartz-Ziv: So Nathan Lambert: thing. It's such a cool thing to be able to know and be a part of in five to ten years. Ravid Shwartz-Ziv: So what do you think are the the most excited directions and most exciting like next next thing to work on? Nathan Lambert: Depends who you're asking. I think it's it's hard for me to ascribe to Ravid Shwartz-Ziv: You. Nathan Lambert: open AI and Anthropic or doing the I think I have my own list of things. Like I'm personally want and planning to in the next 12 months as I set up a new or like merge a bunch of the more like alignment and like or nascent empirical work of character training into like a full Omo style post training recipe and just like understand how model character impacts leading vowels, leading agentic behavior behaviors, how trade-offs and defining the character and like what subtleties in the spec can unfold into the model. And I think it's just like an interesting challenge to learn a bit more about the nature of these models once you've done the hill climbing thing so much. Like once you're like in a l of lab and you're like, we want to make these numbers go up, you do it for six months and then you choose two new targets. Like it's a serious grind and it's really hard work to keep making benchmarks go up. But it is like I think like a lot, a lot of people can do that type of work if they're put into the environment. It's just like, it's like a little bit, it's like a new thing of like trying to understand a bit more of the darky dark messiness of these models. So that's like one side. And then also I just want to build the science of big RL runs because over time I'm pretty confident in a few years it'll look more like pre-training scaling ladder type thing, which is like there's a empirical science, which is the industry standards type of methods that you do when you're trying to scale up your RL runs. And it's baffling to me that there's no common language on that. I think there's a scale RL paper from Meta, which is like the starting point. But I think there's a whole literature there that's going to be built and kind of defined a define a lasting language of scaling post-training. And I was like, okay, why would I not do that? And if I keep listening, Ravid Shwartz-Ziv: Right. Nathan Lambert: I think like on policy distillation is an interesting new tool, but I'm not as like time urgent to do this. It's just like a new post training tool has emerged. It's fun to understand how that works. And this is like paints pretty clearly of how I think about the problems that I personally want to work on. Ravid Shwartz-Ziv: And but do you think like at if in the future like the frontier like models will still be like yeah, you have the p pre training and then you have like post training on a lot of different environments or like problems, separate ones? Nathan Lambert: We could bet on it indefinitely. There's a like we're in a bubble or boom, whatever language we want to use. the investment is there. There's like much higher likelihood that somebody solves something like continual learning now and the investment is just so so high. But it's a fundamentally a scientific problem to me to understand h how you would modify model weights on the fly for a general purpose model. And I think it's really, really hard. I think It's more likely to figure out things like what Cursor did, which is just constantly iterate on variant distribution data and it can get better. And in context learning and like this harn science of harnesses will become a really, really big field. I think the status quo is that the big tech companies know they can get so far without these major innovations that most of my probability mass is that the the methods look fairly similar, which is like pre-training and post-training and RL is the like the RL type of like learning from trials as a way to like it's just a way of getting different experience and feedback into the model, which is it's very sparse. People love to like shit on the information density of RL. But it's like the only way to extract information from an environment and put it into model weights. Just like you give it an environment and it manages to learn about it into the weights. Like that's way more complicated than a pre-training document, which is like just like pure information, obviously. But I think that these things are most likely to stay the same. I think it'll be fun if they change. Like it'll be a whole new if there's any change, it's like, we have like two more years of the AI boom for sure. Which is like the O three type moment, or like, tool use is a new thing. Let's go. We got a long time to hill climb on that. Allen Roush: D d do you think you mentioned continual learning, you know, dw Dario has gone and said on the dwar Dwarkesh dwarfkesh, whatever his name is podcast, that he doesn't think continual learning matters or is important. Nathan Lambert: That's that's my argument of like big tech is gonna build these machines and they can get so much better. And it's not really important because the runway is already there. But if you do make like there's people in the labs that are like RSI is downstream of you make a hundred X and you you start making 10 to 100 X efficiency improvements and applying them to your own models and they don't leak and then you have RSI. And I think like that's like a little overblown. And it's like if you make one of those. It'll just trans slowly transition to the whole industry in a few months and everybody will copy it. so most of these people are scientists Ravid Shwartz-Ziv: Why? Why you think like you see this? Nathan Lambert: and they're proud of their ideas. And I think the people's ability to lock things up in the Bay Area is historically awful, at least until the labs are truly nationalized and shipped off to the middle of the desert. I I just think it's like if you're a scientist at OpenAI and you make a literal 10X improvement, like people are gonna want their glory of like, I made AI ten times better and Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: They're gonna I think they're gonna want this. This is just human nature. Because everybody Allen Roush: Well Nathan Lambert: in OpenAI is gonna know it eventually. Like they're gonna know the person that made their models way better. Allen Roush: It it seems like this is a perfect place for espionage, but it seems like if that's happening, I I guess it appears that China, you know, adversaries that would want a fable or mythos class model don't have one or at least have not shown their hands, so why why hasn't this happened yet if they're so bad at keeping secrets? Nathan Lambert: I think the next models are mostly a recipe Ravid Shwartz-Ziv: Mm. It's far. Nathan Lambert: of compute and data and like the actual training recipe, which is like a long sequence of complex buttons that you have to push and draw on proprietary data and compute. And it's just like really hard to exfiltrate all like you're exfiltrating a company. And most of the algorithmic ideas they probably can take, but like none of them are so big that they obviously steal it. I've I I've been reflecting on the distillation thing, and there's an interesting debate of like now that you know the AI labs literally like jailbreak or the Chinese labs, like a lot of them literally are jailbreaking the APIs to extract the reasoning trace. Like that seems like the core function of what is distillation is called distillation, and it's talked about in very vague and confusing ways. But it's like were did they do that to meaningfully speed up Like it like the Deep Seek R one release and Quinn had this model at the time. Like, did they really like I wonder if they did just extract some bunch of reasoning traces from O one like within a week and they're like, we can train on this and do RL like that type of historical counterfactual would be pretty interesting to see. I think it's like impossible to ever actually prove. But now I think distillation's impact is much more diffuse, where it's like the models are actually more similar. It's a lot harder to zero to one, that type of training. It's like how are you gonna do this? But I don't know. That's kind of a ramble, but I think it's the espionage thing is hard unless you could steal the model weights. But I'll say it's you need to forecast it out two years down the line. I think it th the pressure will continue to increase and different types of walls will keep going up between who builds the models and Whatnot, 'cause it's it's very obviously the case that if you're at one of these companies, you can't just like click around and see the training code or see any of the pre training data or something. Like there's already definitely firewalls like that. But there's not the like social chatting about what you do firewall. Ravid Shwartz-Ziv: So i if you need to guess like how much we can push like current metals and pipelines, do you think like we have like acceleration and like kind of like I don't know, self improvement thing or like we kind of like saturated? Nathan Lambert: I think it's all downstream of monetization. I think the models are gonna keep getting meaningfully better. I d I don't think it's this like recursive idea, but the pace of progress is really high. Ravid Shwartz-Ziv: You don't think it's recursive? Nathan Lambert: No, I think it's like if you if you zoom out of the pace of progress of the last three years, it's been pretty wild and I think it's gonna keep being like that. But it's not gonna be like humans stop being involved. I I don't remember who said this, but I think it's a great counterpoint. It's like if you're gonna have recursive self-improvement of like the training teams and these orgs should be able to just like shed people and get smaller and smaller. And they're all just getting way bigger and they're all getting way more complex data ingestion, like acquisition and more vendors. So it's just like I don't really think that that is the case, but it's a very, very effective way at spending inference compute to make your training pipeline more efficient, which is like you clean data simply All of your kernels are way better. You need less ha individuals to handhold your experiments. Like the experiments work. You have a lower failure rate. And all these things are meaningful, like single digits to double digit percentage add-ons, which is like really, really wild to say. But it's not a 10x year over year model improvement. Because this is like I think it's bottlenecked on kind of fundamental science in some capacity. That is much you can't brute force it. Ravid Shwartz-Ziv: Lake what? Like what? Nathan Lambert: Like it whatever it says continual learning thing or dramatic new architecture change. And this is the output of like very deep intuitions at somebody trying it. And I don't think you could short circuit that. Like the all the background work goes much faster, but like it still takes people's intuitions time to update. And those intuitions are how you keep seeing what the low hanging fruit is for your experiments. Allen Roush: do you do you think that we're in an like AI bubble in terms of investment? I mean you mentioned boom or bubble, like do you think that most of it's like wasted or do you think that like basically people who are long on NVIDIA and these other hardware companies will continue to make money? Nathan Lambert: it's starting to transition a bit towards bubble as there's talks of like big tech companies taking out debt. Once those all the big tech companies just on free clash flow, I think like a boom is a very apt way to describe it. But I I'm pretty optimistic in the value of tokens. I think the fact that H-100s have like doubled in price over the last fourteen months is like mind-boggling. Like I kind of bought into the three to four year appreciation cycle, like ooh, a little bit scary. But it like but like that Allen Roush: you the Michael Berry stuff. Ooh. Nathan Lambert: no, not like that. But I've just bought in of like prices will go down. But it's way harder to get an H one hundred reservation now than it was like fourteen months ago. Like it's literally like the Allen Roush: E even A one hundreds are like that and those are seven years old now. Nathan Lambert: Yeah, so it's like so long as that is the demand profile, I don't think anything particularly like there's gonna be some adjustments along the way. Like there's gonna be some some days with I don't know, 10 to 20 percent drop in valuations. But I don't I th I see it's still chugging along all demand just like this. The really, really hard thing to measure is how much of that demand is just VC propped up like self-funding type of thing. And some of it is there's just so many companies and so little supply that some of these c like some of the companies are so fake and they're gonna get got and it's gonna be good. Like in in some ways that's the healing process. But I think like selling tokens is a much better business than selling GPUs and the demand for tokens is high and you can make a lot higher margins on it. It's like these inference companies are in a very, very good place in that capacity. So so Ravid Shwartz-Ziv: But it looked at like No, but like it looked at like right, like around like January, February, like right, all the companies basically try to to maximize the the co the the tokens that the their developers used, but now it looked that they are Kind of like they didn't see a big improvement, like they see some of it in some cases and but not across all the board, and now they are starting to to restrict it. Do you think this will continue? And actually like let's start from the beginning, like why we actually don't see this like 10x 100x improvement in the efficiency of the or like the quality of the products. Nathan Lambert: 'Cause it's like a let's see. The way that I think the tokens can be useful is if you have a very small organization, a lot of products are now your product quality is proportional to token spend and just people poking around the app and see like this issue, this issue. This is like the early stage app process. And for somebody like me that has like we have some proprietary data on interconnects and we can make dashboards and mini products and things with it. And it's like Unbelievably high ROI. Which is like I don't sell our API right now, but I could set up an API for a data in half a day and make some dashboard in a few hours. And like I could conceivably sell this for five to six figures to hedge funds or whatever. And like that's incredible value capture for like my my like fun blog side project type thing. And that's like mass value. But all the companies that like tech companies have already been way bigger than they needed to be. So there's a lot of organizational and personal inertia between these like power fiefdoms of what gets merged and what gets updated. And I think this is gonna be like very disruptive to the tech industry, who is by their nature ruthless. I think they're the the most effective money making industry of all time. I think eventually their their like their like teeth will show and this ruthlessness will be like a reduction in headcount in big tech. So it's like peak size of the big tech companies. But there's just so much organizational just like friction that makes ta the apps getting better take a long time. And there's a way I think there is then excess spend because a lot of these people are spending on things that conceivably could work, but you just like can't push everything into the same pipe in the way that the orgs are now set up. And it's just kinda Ravid Shwartz-Ziv: But how do you but how you explain it? Like even like entropic apps, right? It like they have like they're quite small, at least they they were quite small, and they have like infinite amount of tokens, right? Even their apps are not so good and quite buggy and and they're like they still have so much place to improvement. Nathan Lambert: Yeah, it's just like it's weird. And I think it's just there's is a lot of inefficiency in it, which is why th the usage will come down. And there's a lot of pressure to shift that to open models and cheaper models. Because anyone that looks is like, this is silly. Like most of it is very silly. There are people who get a lot, a lot out of it. So there are like It's like the high pr the people who are high performers now have high outputs of their models, and the people who are low performers, which is most people, just burned more money. And we're now in the slow reckoning of how do you measure that? And the default policy will probably be less spend or cheaper models, and then the special approvals will be for the good people. Allen Roush: But it it seems inevitable in your mind that humans retain control. Why doesn't AI prompt itself? Nathan Lambert: AI has generally a pretty poor track record of prompting itself. If you put the models in loops, they kind of become diluted. But I think it's like socially, politically, whatever, all the Like like power structures are designed to maintain themselves. And most of the all the the existing companies will therefore maintain some level of human control. It's like the government will maintain human control. The lawyers will keep their jobs because they write the laws. stuff like this. But I do think there will be new like there could be new companies that are formed without humans. I think the bottleneck right now is you need a If you want to use Stripe Atlas, you need a registered agent, which is a human in your state. So it's like there's very clear legal bottlenecks on entities needing humans in this capacity. But I think the future is supportive of very small businesses, which is a few few a few people and proprietary information and many agents. Which is why I think that the job displacement is real, which is like we need to shift the economy from You go through the education path and you get a good job if you're a high perfor high performer at a slowly growing industry with different tiers too. Everybody needs a real liberal arts education and to find something they're interested in, a rabbit hole, and then to build a small business around it and the most successful ones will scale over time. And that's just like a total Like we have to sell tell a totally different story on how people get jobs and navigate their early career. So the in between is where I see mass job displacement and unrest of some capacity. Ravid Shwartz-Ziv: Yeah. We had like several weeks ago, we had Alex Imos in the podcast. he's like he was a professor at Chicago and now he's director of AGI economics at Google. And he said basically that like we had some jobs that like we like we already fully automated sixty years ago, right? And like the the best example is yeah our brokers, right? Like You don't need, right? Like we all know that like that the best thing to do with your money is to put it in a in index funds. But still we have someone that manage our money, right? And because we we need someone that now we are I'm going to a cocktail and I want to tell my friend that yeah I am managing my money in with this person or this person and not just index fund. so I think like Like the job market is much more complicated than everyone are thinking and especially in the Silicon Valley. Nathan Lambert: Yeah. Yeah. I think it's like my projections I feel like are mostly for big tech. And it's like I understand the industry and the big tech is obviously bloated and stuff like this. But the long and there's probably things like this that are older and very successful industries. So maybe like finance. Like finance will do a good job at slowing the rate of junior analysts and only taking the absolute best ones that can leverage AI and otherwise they can use AI and like turn the tap down. But there's still like it's not an elimination. It's just this like onboarding and development and growth is really, really different at the early stages. There's things like hospitality, like that's not gonna change at all. It's like you still gonna want a human at the front desk of your hotel and Most hotels have one or two people. Like I don't think that's gonna change. And like the whole robotics thing I feel behind on the debate and I think that there's a lot of optimism there. Like that job displacement is really, really different of a factor. And but it's like I it Alex is the right person to do this. So it's like what is a percentage of employees of different industries and what are they gonna take? Like self driving cars probably actually more job disruptive than LLMs. Like if you actually get all the truckers and all the taxi drivers away, which my opinion would be great. But like Allen Roush: So so so I have to I have to I have to I Ravid Shwartz-Ziv: But Allen Roush: have to point out that there's a very old Simpsons episode about this exact thing where Homer becomes a trucker and finds out that they had automated the truckers had quietly automated their whole industry with self-driving and then Homer blurbs about it to the public and they all try to stop him and and I I th thought this was very prophetic for what you just said. Ravid Shwartz-Ziv: But okay, let's talk maybe a bit about robotics more. do you like what do you think about it? Are you optimistic that this will be the next like it will have like the the GPT moment soon? Nathan Lambert: I think it'll take at least as long as self-driving cars. Self-driving cars are in a controlled environment, which is roads, which gives them an advantage on deployment. And the thing that Amazon and all of them do is they build new they build new fulfillment centers designed to be from robotics from the ground up. So this type of thing where you can control the environment is really, really good for robotics. But the home robot and or like robot like the time to walk seeing me just like walking around and seeing just like robots walking around on the sidewalk. Like in mass, kinda like we have Weibos everywhere now, is like kinda like I feel like it's it's gonna take the same amount of time. But on the intelligence side, I do think that you can make mass breakthroughs by scaling data and compute. I think this is like it's hard to bet against this. And the argument that I used to be like way more pessimistic, but the argument that I hear for humanoids that is very good is that. We now live in a human bipedal shaped world. And therefore, if you want to scale automation intelligence, the fastest way to do this is through a human form factor. Which is just like it's on its base very reasonable. But the the like timelines that are like within eight years we have robots building data centers and robot factories within the US seems extremely implausible and kind of fake. Cause it'll take a long time to go through all the displacement and You think data centers are hard, like w wait until a robotics factory that is a bunch of robots trying to churn out more robots shows up to a rural town in the US. Like people aren't gonna be happy about that. Ravid Shwartz-Ziv: But w what do you think like are the the bottlenecks at the end? Like do you think like if we have more data we will solve it? Nathan Lambert: That and there's a lot of just development on the actuators. So a lot of these robots are extremely unsafe because their actuators are so powerful that if they make a movement around a human, they'll just catapult you across the room. They're like, are the human form factor but dramatically higher force? And I would need to catch up on like what is this, like one of the companies that's like, we're gonna start selling home robots. Like, are they have they solved this problem? I don't really know. So I think a lot of it is just like that. flywheel, it will take time. It's like Unitry's really far down that path for like the quadrupeds and stuff that they have, and like their economics to produce them is a lot lower and they get a lot more data. But I think my argument is mostly on just like it takes time to do the real world deployment of physical goods that is a lot slower than this LLM transformation is. Allen Roush: what do you think of of these other promising like non-LLM technologies or I guess you know LM's part of it, like JEPA, world models, I guess world models as a path towards robotics. Do you have opinions about these things? Nathan Lambert: I wouldn't say they're super developed. I think this is like a much it's it's much less clear if the timing and data is right for them. It's like and what so I I don't really know. I think it's like I'm I'm again a little bit surprised by the quantity of them. Even just base language models that have vision are still so rough and I don't know how trying to apply that to this new type of paradigm is gonna work for people. It's like one of my big takeaways from the talking to like Mulmo people at AI two is they're like it just seems like in the maturity of development of the methods for multimodal models, they're like a year or a year or more behind the language backbones. Which is really, really a long time in the language modeling industry. So I'm I'm a little s skeptical, but mostly I like have just abstracted it away as information I haven't had to keep up with yet. Ravid Shwartz-Ziv: But so do you have like a guess why? Like even like multi model models, like why they are like at the end we have separate encoders today for vision and text and like we we didn't even like success to to merge them properly. Nathan Lambert: Probably some representation thing, which is what people have told me about. But I like I don't really know. Like Gemini is pretty good, Ravid Shwartz-Ziv: Okay, yeah. Nathan Lambert: but I don't think it's a priority. It's like as it's like RSI style arguments excuse me, become Ravid Shwartz-Ziv: Yeah. Nathan Lambert: top of mind again, like the it like relevance of multimodal goes down which just like it's like part of your timeline and like what is AGI. So it like Demis is like a Like it's really interesting to see like what is going on at Gemini right now as they kind of are waiting to do anything. Like they're just kind of doing nothing. And it's like it'll be interesting to see if they come out swinging with a bet that like just is again very similar. to open AI anthropic or if they're like, well, we're gonna do our thing for a while and take a longer term play, which would they would get absolutely eviscerated. They would be so so eviscerated. So I I I don't really know. I'm not I've I have not done a lot of research in multimodal models. It's generally like of interest, but haven't had an entry point. Ravid Shwartz-Ziv: Okay, great. what about so like how do you see so like we talked about a bit about it the beginning about open ecosystem, but what do you think we actually should do? what you you try to do? Nathan Lambert: Yeah, so I think like most of this is downstream of investment. I think like having more open model providers in the US right now is very, very existential. And I think NVIDIA NVIDIA is trying to bake roll the whole ecosystem in the US, pretty much. And without them things would be a lot a lot worse. But it's very hard when AI is becoming more political and more under more regulatory scrutiny when the simplest thing that people actually associate is like open source AI is Chinese AI, which is just kind of like non-tenable. So this is why it's like break that mental association is the quickest thing that we can do. And then otherwise a lot of people in the open space can look into more diverse ideas and do things that are more of like, how do you make an open model that complements the closed agent? Because Closed agents have product market fit. And to the extent that you can do research and build models that have genuine use cases within that sphere, it's like the longevity will be there and there will be way of like economic value to making this. I think it's really important to like understand the role you're playing when you're building open models and the the like kimmies and zais of the world, which are building businesses on near frontier open models. I don't know their exact business plan, but that niche is that they're gonna offer something that is close but not as good as the open AI's in Anthropic for a much lower price, and they figure out how to get a margin. And that's like one niche that could get in, but you need to be really well capitalized to do this. And then the other niche is like you work at a company that has a special domain and you can make a medium-sized niche model that unlock certain use cases in that, whether it's like on prem or low latency or something that kind of like can make AI move further forward in this domain of interest and like open models therefore could be like efficient or have more distribution. So I I'm kind of trying to make people be better at describing like what they're actually doing with open models and they're not just like open models for vibes is like not a indefinite type of p plan. Ravid Shwartz-Ziv: And do you think we'll see both? Nathan Lambert: I think so. I think there's still a lot of economic value to just saying we're close to the frontier. And this is why like Reflection hasn't released the model yet. It's like their business model is bent on like their first model being good. And I'm like, I don't I don't know what they're actually doing. Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: But a lot of people are kind of hands tied. In that capacity, I'm very excited about the like inference and fine-tuning category supporting more companies interested in open models. It's like Thinking Machines stumbled into a multi-hundred million dollar ARR business with Tinker. And Tinker is entire entirely a business that is predicated on open models. So it's like that business grows, they're very clearly incentivized to like build their own model because their own model could be like integrated and have better margins and work better in their system. And things like that are really good. Same thing with the inference companies. It's like the Together's base tens, fireworks, they're not selling open AI and anthropic tokens. And their demand is very good. And all of these are now figuring out of like how do you create positive sum flows of RD money from these companies to make the ecosystem more functional. And the hard link is like How do these become base models that are better? And like that's still a lot of the bottleneck. But I'm a bit more optimistic because these agents and like GLM 5.2 have like in my mind made those like some proportion of that revenue is like very real productive work. And that's still pretty new in my opinion. Like I think a Ravid Shwartz-Ziv: Yeah. Nathan Lambert: lot of these open models were kind of bad and like kind of fake for a long time. And we're d we're early in that. And that gives me a bit more, like as I said, optimism, but it's still Hard to get there. I think personally I'm more invested in just like the open science version, which is trying to activate as many scientists as possible, which is like a five to ten year technological diffusion type thing and a safety thing by getting more eyes on target. But for the Oaken ecosystem to work, the existential thing is the economics, which is like this one to three time year timeframe. So there's like There's a lot of issues at play with open models, and I think sometimes people get confused because they're just like open models solve all of these things. And I think even AI two was very much like, we're gonna let people own their economics by having a fully open state of the art model. And really open o AI two was like, we're helping with technological diffusion on a ten year timeline by activating all the academics that are studying language models on their niche problems. And It's really good to do, but I feel like it was like not even communicated that well because businesses have incentives to say a simple story that has an ROI. Ravid Shwartz-Ziv: So you you're you're not considered like AI T U as like a failure at the end? Nathan Lambert: No, I think it's like like the AI two as a org is going through a governance change, which is like refocusing it onto different things. But this is what I mean, like non talk I talked about this before, like nonprofits are really imperf like imperfect instruments. And that's a type of imperfection that emerged in AI two. But I think it's like Olmo is a massively successful project and a r I I think of it as like research infrastructure and a massive success in this and the like the checkpoints validate this. But no one's gonna remember Olmo as like a competitor to Quen in terms of like a small model that you deploy into an application. And I think that's fine. Like I don't I don't think it was ever at its soul trying to do this, but it was a helpful way to potentially like fundraise and be like, we're gonna beat Quen with an American Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: model type thing. Ravid Shwartz-Ziv: So so why you think like actually like NVIDIA is not like building a frontier model, right? Like they recently they started, right? Like with Nemotron, but it's still not there. Why do you think is the reason? Nathan Lambert: I think they're very good at it, but it's not their core business. And building language models is hard. So I think they're like a very much a they're an infrastructure provider and they're very they're the best in the world at it and making a lot of money at it. And the people who you're competing with and building models, it's like having the best model is existential to their startup. So it's very hard for NVIDIA to have the type of motivation that the people at Kimmy and ZAI have on their models, where like those people obviously have fewer resources, but it's like They're all like, we need to do this or we die. And that Ravid Shwartz-Ziv: Yeah. Nathan Lambert: competitive pressure is very effective for making people have like innovations and stuff. And like NVIDIA is gonna create good models, but it's it's like they're not gonna develop that culture. So it's like they need to have an impact, which they're trying to do, which is like we do models and other stuff. Like we have open data, we support people doing other stuff with their models. And it's like I think that's kind of how they can make up for it, is that they are rich and they can just keep investing in order to make it a bit easier to use their models and have more impact. And I think so much of building models is competition and just human drive because it is like very hard work and massive disparities in compute and resources can create similar models if people have a like a bit better execution. Which is really hard to price in. It's really like people need to understand that like just research or skill alignment and org chart can have like multiples on the output of the model. Ravid Shwartz-Ziv: but but do you think like but do you think Nathan Lambert: It's it's hard for regulators to understand. Ravid Shwartz-Ziv: like e even like besides like open open models, do you think like in the future we will see more more companies are trying to compete with the the frontier labs? Or to see more models? Nathan Lambert: There's still a lot more models coming, and I'm always surprised by how many more model companies there can be. So I need to weigh my like I've been wrong a bit on this because I have expected consolidation over time. So I s I my gut still says that consolidation will come and there can't just be more companies all the time. But if to selling tokens is such a valuable business and therefore people will figure out how to monetize, like building the Building the token machine. I think it's like people are gonna figure that out, which therefore the answer would be yes. So I I I'm in two minds, but it's it's such a expensive activity that if the entry point is a mythoscale model, which is like a billion dollars of CapEx or a billion dollars of R and D investment, like eventually like some people can't do that or can't fundraise to do that. Allen Roush: Yeah, I agree with you on expecting consolidation in the space for sure. for do you do you think that there's a significant distinction between open source as freely available and open source as in the code and possibly even data to reproduce with models? Like do you think that you know going all the way is more like so-called morally good? or is it really just around that like scientific diffusion stuff you've talked about earlier? Nathan Lambert: If I mean like if a bet if the best model in the world was fully open, there would definitely be like I think people would figure out how to use it, but it's kind of competitively damaging. So I think the open artifact if you compare it to an operating system is like the training recipe, because you can keep improving the training recipe with like small contributions and then when you press go every three months you get a checkpoint, which is like an output of it. Which like that's like the new part that's complicated with previous open source pathways. And I I think it's like at the simplest level, like open source is what people agree it to be, which is kind of these weights and all the dynamics are a little bit different. I I don't expect it to change a ton on if there's fully open things. It's just that like you need the fully open things to do a lot of types of research that is really deep in the understanding how the model and training works. And that that's not gonna change. I think we've seen the distribution of research done change a lot towards prompting avows and things like this, but that's just downstream of resource availability to a large extent, I think. Like people don't people would rather be doing training and stuff that's like more ditty gritty into this, but the average researcher doesn't have the ability to do that. You guys are both on mute now. Allen Roush: No, sorry. Ravid Shwartz-Ziv: Maybe let's talk a bit about your incoming book, RLH. So yeah, like tell us why the like okay first like why to to write a book about RLH and Nathan Lambert: So I bought the domain and wanted to do this like years ago. Mostly because when I was learning RL H F and post training, there was literally no book. So I was like, I'll take practice taking notes and studying fundamentals to like make the book I would want to have. And if I were to do this full time, I definitely like sometime last year I would have done a ton of work to refactor it to just call it a post training book, which it kinda is. It's like Ravid Shwartz-Ziv: Yeah. Nathan Lambert: a post training book through the lens of RL H F. Although as a scientist, I'm like fine with the fact that it's called R L H F 'cause R L H F is such an important technology that is like it's it's it's fine. It's not the optimal book tell title for selling on Amazon, but that's not really why I did it. But it's it's just like a nice it's the structure that the book follows is like how to understand post training in the canonical recipes that emerged and those emerged when RLHF was popular. But in reality it's like it's that plus the long tail of post training discussions, like other synthetic data methods and evaluation and character training and these things that get interweaven into it. So that's why I joke I refer to it as like, yeah, I also wrote a post-training book. And it's very much about trying to communicate intuitions on how model training works through the scaffolding of like the math and the formalization that people use. So you have like different optimization tools and it's explaining what these tools are and then how I try to think about them being useful in a post-training recipe. So it's very much not a like You're building this from scratch. So I would recommend people to be comfortable with like a Sebastian Rasco type of book and then pick it up to be like, okay, like I've seen the code. You need to read the code once in your life, Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: and now the coding agents will do it. But it's like, what are the people actually doing and thinking about when they're applying this to a real model? Which is like the stuff that nobody will really ever really tell you. So in that nature, it's a somewhat advanced book. Like I think of it as like, very good for a grad student that's just getting like tr out of undergrad and wants to learn about what post training is and how to think about these models and things like this. So I'm I'm happy about it in that regard, but it's definitely a niche thing in some ways. Ravid Shwartz-Ziv: And so wh so first of all do you think like RHF like is still like an important method these days? Do you think this is like it was very important, right? Like at the beginning. Nathan Lambert: I think p labs are still doing preference tuning in some capacity. Like there is a reward model that is downstream of preferences and these training pipelines, I'm very d sure about. But it's like, are people still talking about RLHF? Probably not. I think there's still human data getting plugged into the models and there's still reward models being trained. But it's not that like it's not the kind of like cultural center point that it once was, which is like when the word post training emerged, like That was very much immediately downstream of like you did the RLHF pipeline. That's like that's the era when I wrote the like structure of the book. And like and then I was just like, I have too many jobs to go back and essentially the thing that I would consider doing in a second edition is I rewrite every chapter's introduction and the introduction to be like this in the post-training book. And if I did this, Ravid Shwartz-Ziv: Mm-hmm. Nathan Lambert: it would work. It's just like 35. It's just like tens of hours of time that I don't have with other things in my life to do. But it's like the actual content would be like almost exactly the same. I think I need to do another week on tool use because that's blown up. But like there's always gonna be stuff like that. Like it's always evolving pretty quick. Ravid Shwartz-Ziv: And what what do you think about like so it looks at like the the last like JLM, like they train it with critic models versus the GRPO. What do you think about it? Do you think this is the future? Or like the past and the Nathan Lambert: It's I think GRPO at its Ravid Shwartz-Ziv: future? Nathan Lambert: GRPO is very much of like a scaling pill algorithm, which is you have a lot of compute lying around, it's pretty easy to find great find ad like advantage gradients in it. But like the Chinese labs are very, very training compute limited. And therefore if you could switch to PPO and have it work somewhat similarly, you can do many more experiments on your recipe and keep hill climbing. But I do not cast lasting aspersions on like which is better. I think the value functions are very compelling, which is you can have per token credit assignment across complex behavior. And it seems like the on-policy distillation methods are much closer to supplying that than PPO, because there are many, many more models that are like we use multi-teacher on-policy distillation to ship our org chart into a model because we can have teams working on specialists and then distillation gives a dense feedback on each token and I I see like that I the idea is this like way to get dense feedback and I I think I and many people were surprised they used PPO but not like super surprised. Ravid Shwartz-Ziv: Mm-hmm. So do you think like as we have like longer trajectories we will see more and more PPO style algorithms? Like now we have two usage, right? It it's much longer. Nathan Lambert: Yeah. I mean there's definitely the economic incentive and the it's like it might not like it might not be feasible to do multiple rollouts in every environment and therefore you need a different way to integrate this. Or you like you might not be able to do rollouts that actually are like differentiable in this like outcome only way when tasks get sufficiently complicated. I think more of our training today is still on you have these like intermediate building block tasks that are like pieces of the complicated tasks that the models can do in I dunno, like ten minutes or less and a lot of the training is still probably closer to this than the you don't have like an eight hour training task in your RL loop. That's just too it would be too off policy or too unstable or too something. Ravid Shwartz-Ziv: Okay. y I think we are almost like out of time. do you have anything else that you want to add? Nathan Lambert: No, I think it's thanks for having me. I think it's good to have one of my goals is that more researchers could define the conversations that are happening because there's kind of like a a civic independence or neutrality to it that's like very grounded in a way that is useful to the world. That's part of what I think of what I am doing and all my various endeavors from writing this book, which is more f more new people get into it or interconnects is grounded in this. So that's why I'm like that's why I said yes, I want to support. people doing stuff like this. So keep going. Happy to see you guys. Ravid Shwartz-Ziv: Thank you. Thank you so much and thank you for coming. Allen Roush: Yeah. Yeah. It was a pleasure to meet you. I mean I've been seeing your stuff fighting the good fight on open models for years now. Nathan Lambert: Yeah, thanks. Ravid Shwartz-Ziv: Thank you.