Ravid Shwartz-Ziv: Hi everyone and welcome back to the information button podcast. And today we are hosting Alexia Julicar Martin Yu. She's a principal researcher at Microsoft and the author of a the paper of tiny recursive models. paper I think it's called Less Is More, right? Recursive reasoning, reasoning with tiny networks, that actually like achieved like really high performance in only fraction of the number of parameters with compared to to much bigger models and it achieved like a 45% on a IRC AGI1 and 8% on a IRC AGI2. so hi Alexia thank you for joining us Alexia: Hi, thanks for having me. Ravid Shwartz-Ziv: And as always, Ailen. Allen Roush: Hi, happy to be here again. Ravid Shwartz-Ziv: So I want to start like you know from the beginning because like I think it's it's very interesting to to understand how you actually like end up with developing tiny models, you worked a lot on diffusion models and like on gans and how to to to make them better. So maybe like you c can you start like with maybe like why to work on GANs and what make you make this shift to to language models and then to tiny recursive models? Alexia: Yeah, well I've always been It started a bit with GANs I would say. it's when I first saw what the first e text to image GAN, it was the in the very, very early stuff. It's when I saw this that I got really excited about the field and I wanted to transition because I was in biostatistics and statisticians don't like it when I say this, but but I feel the f like the field of statistics have stagnated a lot. in the the recent year and there wasn't there's not it's more there's not that much big progress but in computer science things were crazy. It was about scaling stuff, trying crazy stuff, like generating images. So that's when I got excited about Galleries. Ravid Shwartz-Ziv: Wha why you think it's that? Why you think like statistics is because like a dance this is kind of like data scientists, right? Like take data, calculate some things, but like i I again I'm not an expert, but like no one work on statistics anymore, you know? Alexia: Yeah, well t there's two things there's it's well one thing with statistics is that you have in the in the university it's all like theoretical statistics. So very, very theoretical, very specific solution to specific problems. And they kind of live in their la la lands at the end. Not in the real world. So especially on the causality side there's always this assumption that are never true in the real world, never apply. But they develop this method out in this Box, but then it's misapplied always in practice. Sorry, I had to say something about causality because it's a very messy film. But but okay, but then there's the practitioner side, which I was in. The problem is the data is so bad in quality. Just so bad, so limited. You have basically let's say you have for us we had a court of children and their mother as the child gets born. We have a test to see how the mother is is behaving with the child, how the child is developing, stuff like this, right? So you can evaluate like understand how how the the the the children evolve over time based on their environment and stuff. Well that kind of data is super hard to get you get missing appointment, missing tests, you have insane amount of missing data and then you don't have that much participant and it it it it's insane. So the kind of data that the psychologists or psychiatrists or in in medicine also you're dealing with small per small data set with a lot of missing data so so a lot of the deep learning stuff don't even work at all like no you have to rely on very very simple like linear reg regression and stuff so the thing is I think the data really limits what what they can do. They're kind of stuck in a very limited box because of that. So I think that's the main thing, the data. Well in AI it's like, we c let's get more data. Let's make more data. Let's make everyone get synthetic data. Let's build crazy amount of data and scale things. So that's not what we see really and in the the in like biostatistics and It's it's very small scale, so you have to be very careful. So that's the thing, that there's a big stagnation, I think. No no ambition to to to do crazy stuff like generative models. Allen Roush: Do you think that this was Ravid Shwartz-Ziv: Okay. Allen Roush: because of chat GPT getting so good and everybody leaving to work on LLMs? Alexia: I don't think it has to do with it was before L L I saw a big stagnation. Like some of the cool techniques like decision trees were invented by statistician, but since then it's been Very limited. But no, but I don't know what the effect. That's the thing. I wasn't there when the big L LM craze arrived. I wasn't working as a statistician. So I cannot say how it is today. I bet they use a lot of L LM to do the data science for sure. I haven't s been there to see this but Ravid Shwartz-Ziv: Okay, so you end up with moving from from statistics and then what? Then like why why it was so interesting and like what w what problem did you try to solve? Alexia: Yeah, I mean it it's just the idea that you could generate something instead of just predicting, classifying and you could really create something new, create novel data and that was mind blowing to me. And I'll be honest, also there was a video game I played about ascension AI and a computer scientist and I thought, Hey, I wanna be the computer scientist who create Allen Roush: Wha w what game? Alexia: I don't remember the name, it was a point and click kind of game. But Allen Roush: Wait, and what year was this? Alexia: maybe two years before I got into AI, something like that. Like that got me hooked and then thinking, yeah, maybe computer scientists would be cool. And then I saw the gun stuff and that's where I was so excited, so so I jumped on the on the train then I learned everything I could about AI and did some some work on GAN and did Meow Generator, the first generator of of cat pictures. That was fun. Then relativistic GAN, so I've done a lot of work on on guns before. Because yeah, it was tricky time. I c I couldn't get in the PhD. I applied to the the PhD at in at Miller and also at Megan I think and I was rejected from bold. So that's where I did some research solo and I got published and things moved along really fast after that. Ravid Shwartz-Ziv: D do you like to make a solar research? Alexia: yeah, yeah. It worked well for me. Ravid Shwartz-Ziv: Why because like like at least like in in you know recent years it looks at like the the field moves to to this like big lab but like with a lot of like people that are working on a project from different perspectives and there is like the these like single author papers are so rare these days. Alexia: Yeah, terror but but I like it 'cause you kind of have no constraint, nobody to tell you add more this and my guy do this, do that. I don't know. I I I like it and and I feel like for for creativity it's a good it's a good place. Like of course for large scale stuff it's better to have big teams. But honestly, if you're especially smaller scale experiments, generally it's like one or two authors that do the the most of the work, right? That's the reality, especially in academia. But I think even in the industry. It's just of course for big models you need cra you need so much stuff. So you have out our list of two hundred people. That's inevitable 'cause it's such a big endeavor to do the whole pre training, post training, all this Ravid Shwartz-Ziv: Okay, so you you you chose to do like guns and and and all these things and then what? Alexia: and then yeah there was diffusion right. I I noticed it the first paper by Yang Song right when it came out. I s I saw the potential so I went and did some diffusion stuff before it was called that. And I thought it was super cool. But now everybody does diffusion, so it's not as cool anymore. Ravid Shwartz-Ziv: We had several episodes about effusion with like Hermond came here and also like others. Alexia: cool. Ravid Shwartz-Ziv: Yeah. Okay, so it's not so it it was not so cool and then Alexia: Yeah, well and then I've done some work on reasoning with the time. Well I did there was a lot more stuff. I went the thing is from PhD to to my first research position at Samsung, there wasn't jump in compute actually. It was no jump in compute. I had the same con the same compute which was academic compute. So I was very, very, very limited in compute. So 'cause I I I mean I I finished my PhD with work on video diffusion and the the idea was to scale this up to audio and stuff like this to add more. But I never got to do this because I was so limited in compute with getting four GPU get maybe the GPUs. So so yeah I've been very limited. So so that's the the the the main thing. So limitation made me try a lot of of different ideas that could work at small scale. So like for it Allen Roush: Yeah. Alexia: like decision trees diffusion. I did the first case of this it works very well. For tabular data Which is more in line with my my work on on as biostatistician. I did a lot of tabler data. But I tried off a bunch of stuff basically and at the end TRM was one of these exciting stuff. I saw the archical reasoning paper. It was very exciting. I thought there's a big potential there. But the paper is a bit horrible cause everything Ravid Shwartz-Ziv: Why? Alexia: is everything was explained based on biology and it was all said this is clearly this all come from biology and and everything was given in one go without explanation how they they got to dare accept that it made sense based on the brain of of mouses. I thought it was a bit a bit crazy. Allen Roush: So do do you have opinions about like the diffusion versus standard autoregressive methods? In particular, I've noticed that diffusion models seem to handle creativity better than competitor autoregressive models on image tasks at least. I don't know if that's architectural or based on the data. I was just curious if you have opinions on this. Alexia: it's probably inductive biases, but in the end I have no preference views both basically for molecule for example I I found that from literature like simple smiles it's the the the the format of the the molecule, it's just characters from left to right. And And this just autoregressive generation of that is is honestly better. I I found it was better or as good as no better than diffusion. Because diffusion would working on graph of molecule always miss some connection between the graph always cause cause some bugs and problems but you didn't have that with auto regression. So it's really case dependent depending on your problem, I think both auto regression and diffusion are both too too powerful technique, you just have to pick the best one based on your problem, I think. Ravid Shwartz-Ziv: So okay, I so I I want to like understand like so you you read this paper about like I don't recursive models. and and and what? You think that like w we should try it in different like settings or like they didn't like did good enough job or like I can extend it, like how you Alexia: It's the explanation, explanation. It made no sense to me. When people start I mean I'm a d no bullshit person, you know. I don't I don't like a fancy word for n for no reason. I I like in biostatistics it's f it's filled with that. so many complex explanation of simple things. I I don't like this kind of stuff so and there was no ablation in the paper. I thought there's big potential but no ablation all based on on biology. No no but I I'm excited about it. So I I I wanted to explore it and see can I understand it in a simpler way. And yes there's a much simpler way to understand it. You don't need biology, you don't need anything like this. And once you understand it this way, everything s clicks and I found ways to to make it more efficient and better. Ravid Shwartz-Ziv: So tell us. Tell us about the model. And t in the simple Alexia: Yes. Ravid Shwartz-Ziv: way. Alexia: So in the simple way, the way it works is your nor normally you go from you you go from from the your past context to from question to answer for a problem it's in general like you have a question you have an answer or you have a text prompt and then you have an image it's always kind of the same thing you start from input to output now the the the thing is you can try it to do it in one model that takes that goes from from the input and that makes the output you can have one model that learned this whole transformation. And it's different basically whether you have diffusion, if you have diffusion it's noise to image. If you have autoregression it's one word to the next. But it's always a transformation from input to output. Now the thing is this this is if the transformation is sever, massive, it is a big it's it's something very difficult to learn. You need a big, big, big model for this. But the idea here is that you can you can basically have a smaller model that that work through its hidden state. So normally it's input and then there's a long series of hidden state that goes through many layer and then you have the output. But instead you you you go through a small network with a placeholder answer. And it thinks a little bit, so it does a few hidden state iteration and then it come up with a new updated answer. Then it takes its old answer and then it does the same thing again to update it. So it's keeping track of both its hidden state, which we don't know, it's like a latent, we don't know exactly how to interpret it, but it's really going through you can see it as going through a large. network of transformation it might store a lot of information to to to continue its processing but yeah basically it it keeps track of its hidden state and then the answer the current answer and it updates the answer once in a while after thinking through through a few steps so Ravid Shwartz-Ziv: So you have right, like you have two networks, right? Alexia: you can have one also, that's w one of the thing Karaama showed like you can have just one network. You can learn to diff distinguish the two. But yeah, you can have two networks. It's it's it's a choice basically. hyperprimage, let's say. Ravid Shwartz-Ziv: yeah, and the idea that like you can create like these layers like basically on the fly, right? so so w why you think like what do you think is like I don't know the the important takeaway here? Like what what do you think we can actually like use in from this? Like do you think this is like that yeah you actually need to to compress models as much as possible or like it's more like this like kind of like you know like building on the fly the layer? Like what do you think is Alexia: Well, I like the idea of having smaller model, more like we don't need all these crazy tons of trillions of parameter. There are ways to be more efficient. Like the loop the loop kind of transformer, there's a big big direction there's a big Big subfield there and basically a lot of them show like you can just half the number of layer and do two loops and boom you have same performance or better. So there's kind of a trade-off. So you can you can get more compression. You don't need as much parameter. There are some and one way is to keep track of your answer. So keeping track of your answer as you move along is a nice way to not need as big networks because you you you have a separate memory if you want. So you only need to learn a smaller, simpler function for. So that's one thing. the other thing was the deep supervision. In the original paper, which was also in the universal transformer, I think, older paper. But this idea was to it's it's it's it's very nice because it's to to say that you have a long process, so you don't want to backdrop to so many so many layers. So if you if you do some loops and then you you say my push, push. push my answer to be the correct answer then I do a bit more reasoning more loops and then I keep push then the new updated answer more so you don't need to backprop to the whole graph Like people try that before with like network, like trying to like a layer one or two, try to predict the output. There was some work in that vein before, but now it's just you can just loop the idea of looping the same model and at every loop you you have a supervision step but you truncate the path. So you remove the gradient. So you don't need to backprop to this crazy thing. Ravid Shwartz-Ziv: And and so okay, w where do you think like is the where it's success and like where are the the fatal cases, you know, for for these types of networks? Alexia: right now it's it's very Ravid Shwartz-Ziv: We know that for IRC it's good, right? For the Alexia: Sorry, can you repeat? Ravid Shwartz-Ziv: Yeah, so like we we we we know that like for the for the for the ILC AGI it's good. It works really well. what other types or like maybe you have intuition why it works there and like and other places. Alexia: Right, the thing is puzzle task like sudoku or whatever, it's it's good to keep track of the whole context. The thing is autoregressive model, they kind of struggle because yeah, if you if you try to fail some parts of the sudoku, let's say autoregressively, but then you made a mistake. Then you cannot go back and fix it. So this simple thing causes a lot of issues. Well, if you take event diffusion model, now to show Sudoku as a nice example for why it's better, because if you take the whole Sudoku and whether you use diffusion or you use some recursive model like HRM, TRM, then if you have the whole thing and it's bi-directional attention over the whole the the whole thing then it's easier you can iterate on the the puzzle until you kind of converge to something okay then it f it fits the constraint it's much easier than trying to self-improve like this while taking the whole thing in context versus auto regressive left to right. So some stuff are good left to right some stuff are better as a as a whole like Ravid Shwartz-Ziv: And and and do you think like there are some like cases that like you you just need to to scale it up, you know, like it's not even like related to to to the specific structure or learning, you know, like the bitter lesson tells us that it doesn't really matter, right? And like everyone or so like build on yeah, just s scale Alexia: Yeah, the the bare lesson is is Ravid Shwartz-Ziv: it up. Alexia: yeah, scaling is very important, but But I feel like the bitter lesson is not always a a true bitter lesson. 'Cause there's there's if you look at animals you have natural selection and then in and then as human we we learn from each other and we evolve as a species. But but you have the same thing with architecture, like we're we're picking the good architecture, we're leaving the bad one, we're picking the right data. We're d even the data like everything is handmade to be as good as possible. But scaling is not guaranteed always like you have to pick the exact perfect recipe and things get unstable very quickly. let's say RL and stuff, we we all know this. So yeah scaling is is is very important. But but it's not always super straightforward and and yeah we need more scaling for for this kind of technique like recursive models and yeah to to to get to the large scale L L or multimodal Allen Roush: So so I i I know you're into small models. What what's your opinion on things like parameter efficient fine tuning? do you think it's promising technique and have you worked in it? Alexia: Yeah, I've used LoRa before or small there's a whole zoo also of LoRa like method. Yeah, this this is clearly promising because you don't need you don't need it's more more at the pre training level that you need the the to learn to learn at a very w h wide how to say I forgot the name but what when your matrix is as it needs to be full ranked, yeah you need to be to have very high rank in your matrices when you're learning you have you need to have a lot of space for things to to be learned. But once you have a good good good place when you have trained for a while then then yeah it's easy to to do small low rank transformation, low rank fine tuning, low rank stuff. Or even averaging averaging weights work very well once you have a model that is pre-trained. Ravid Shwartz-Ziv: So d d do you think like at the end how to say like if I have like infinite compute and infinite data, do you think there is like cases that I will want to use RLMs or like any other more smaller models? Alexia: If you but yeah, if you have infinite compute, infinite but still you you always have a latency bottleneck. So processing the your data takes takes a while. So that's definitively but yeah, if you have infinite everything you can go as big as you need, but but still you need something still smarter I think than just auto regression. Ravid Shwartz-Ziv: What? What is the solution? Alexia: Like this. We haven't found it yet, obviously, but Ravid Shwartz-Ziv: How how you see like what is the guess the like either the principles, the design principles of the of the solution will look like? Alexia: Well for me I always look at animals like as examples. Like we have to think we we got there, we were very smart animals, so how do we get as close as possible to to us? And Ravid Shwartz-Ziv: Why why you think like animals are like good good models? You know, like artificial like models are very different, right? Like the way that like we train them, the data, like everything is so different. Like why you think it's a good I don't know, it's a good thing to to look on animals and to to to try to to make it work? Alexia: Because animal are very efficient in their data processing and very extremely smart and like models can be done in very I know they're very extremely different. But my interest has always been to make AI that gets closer to human at some point. We're not there yet because we have handcrafted data and and the training, everything is very different. But I think that emulating human is is a way to get something a bit more intelligent. Allen Roush: Well and I'll I'll point out animals are very good at exploring their state space and often making few assumptions about it. And and in fact, you know, I I would also answer by saying dogs, in particular that I have experience with, as smart as they are, they are also highly predictable in and I mean some dogs more than others, and so I find them almost sometimes analogous to very chaotic computer programs. Alexia: Yeah then dogs also sometimes they have the pat pot problem, like there's there's there's a grid in front of my dog and and so I just just just go around and and doesn't understand but but every animal has a stuff. Ravid Shwartz-Ziv: Well but w what do you think like are the missing factors in order that like tiny models will no like will reach to I don't know current exota models? Alexia: Well t there need to be more research on on this the this subject for sure. That that's one thing. Because everybody kind of there's a s paradigm that works. So I feel like there's too much time spent on that single paradigm. Wh instead of looking for different ways to to solve the to solve problem. But but yeah, I I I think we need more inspiration. From biology I But like in yeah in a smart way. I think animals are a good example to to follow. 'Cause everything is uncrafted right now. It's like and then we're we're fishing for the like people are fishing for the data that gives you the best downstream performance on specific math and coding tasks. And it's like very it all comes from the the human, like we try to optimize for specific things and Ravid Shwartz-Ziv: W and what do you think about self improvement? Like the model will improve itself. Do you believe in this direction? But it looked at like Alexia: Yeah, that that that's a dream goal. Ravid Shwartz-Ziv: I don't know, at least like in the Silicon Valley it's very like a trendy term. Do do you agree with that? Like Yeah. Alexia: Yeah, I I think so well we have to work on that because it's like you have to have a system that can self evolve. So that it's not so that it's really a bitter lesson, it's really exploring and learning on its own so that it's not just us feeding it, the right data, the right things or everything. Like because of course as we there's so many like data people working on data, selling data and say of course the models are getting better but they need to learn to get better on their own some some way, for sure. Ravid Shwartz-Ziv: How? Alexia: Well that's that that's a difficult question. If you find the answer can come come talk to her. Ravid Shwartz-Ziv: You just need like one trillion dollars and many researchers. Allen Roush: I if you had that kind of money, where would you invest it? Alexia: All the that money for for yeah, I would go for for definitively so self learning and and more human like AI for sure. I think that's that's what that's what is missing. Like these are just question answer books. They they don't have they're missing something that that we have like in nature. Ravid Shwartz-Ziv: But but w what do you think is the or are the missing components? You know, is it the data, the like the the transformer, the basic component itself, like the amount of compute, the communication, multi agent? Alexia: Yeah, one thing is the the memory f because AI model kind of live in a bubble. It starts from from nothing and then then you have some memory through through tokens and then we We try to compress a bit the token and stuff and and we save some of the the past interaction as text but but it's like the it's perfect it's like it's like when at some point I was I I knew a guy who who had no memory. So as he would go to sleep he wouldn't remember things the next day. So he he counted his cigarettes and stuff, I remember Just to make sure it doesn't smoke too much. But this is kind of like AI right now. it's it's it can look at at its notes to so I did this, okay we did okay. But doesn't have life experience or something. It's very it has very it lives in this small bubble, it doesn't have memory of of its past or so Ravid Shwartz-Ziv: And and do you think K R can be a good solution for continual learning? Alexia: R and N, yeah, or linear attention, this kind of technique could be possible, but we need different training paradigm, I think rather than because we're just feeding a ton of random data, a ton of random shit to to them, so it's not it's not conductive to long term memorization and and experience. Ravid Shwartz-Ziv: And do you think so like right, like everyone are talking about like test time compute, right? And kind of like theorem is a test time compute with just tiny model instead of like of a giant one. So why do you think like that that the the feel like default is still the giant version? Alexia: the giant model they work very well. This just they cost million millions of dollars, that's a thing. They work very well but they're slow. They and for some tasks, like I said, there's some tasks where auto regression is not ideal. Like if you have a puzzle or something, it's easier to work through the the the whole context. But now there's diffusion, like diffusion LLM, so that's a cool direction. But but yeah Ravid Shwartz-Ziv: So so like do you think like we will see like different like scaling like test time scaling? Alexia: Alright, that's test time. Sorry, I had forgot what was the the question. Yeah, there's there's many many ways to do test time compute, so chain of thought is just just one of them, right? But now There's were there's so many ways of of trying to do test time compute within looping, looping, all kinds of looping in different completely different ways. So and you I there was so much work on TRM with test time compute also like there's just adding noise, for example, just adding noise into the hidden state so that it thinks different solutions and you pick out what You think is the best. There there's just so many ways to do test time compute. It's not clear what is the the best approach. But yeah, writing down things for for like math or something, it's super helpful because if you have a problem to solve, a math problem, I mean I don't know for you but my teacher would give me zero if I just gave the answer, right? I needed to give the whole the whole explanation of how I got there. So this kind of thinking is extremely helpful but but there's it doesn't need to be for different problem might have different ways where to to Ravid Shwartz-Ziv: But d do you think like it should be like in the token space level, do you think? You know, like there is like Alexia: It's not clear, there's there's various ways. Token level works great, for some tasks, but it's not clear that it would always be the case, right? For Ravid Shwartz-Ziv: Yeah, there was the coconut paper, right? Like maybe two years ago already. With the first one. Alexia: Yeah, that's the the first one that started it, but there's been there's so many like so many like alternative way Ravid Shwartz-Ziv: But i it looked at it like it's like people didn't like success really to to s scale it up, right? And like to use it like as a I don't know, a real like in in in real use cases. do you think this is something like fundamental there or like this is just we didn't understand how to how to use it? Alexia: Well there's a lot of research where it's still early but and then things have trade off and then like is it better to only do C O T or on like if you do only C O T versus only coconut, maybe COT is better. But then maybe if you did a little bit of looping internally and then a little bit of strain of thought maybe it's better. There's like ways to It's like you can you can take advantage of all this technique, maybe even together. But yeah, it's n it's not clear right now what is the best Ravid Shwartz-Ziv: But but do you think like it's actually necessary because like I don't know in other like domains we are not doing this something like that, right? Like in tabular data. We don't have that this like hidden like like chain of thought about it that we are like iterate and trying to to run them or something like that. And here we are using it because these like words, right? Like we understand the language, we understand it. yeah, the model is actually like he's thinking about something, so let's try to steer this thinking. Alexia: All right, it's Ravid Shwartz-Ziv: So do you think this is like it's useful? Okay, so there are several options, right? Like it can be that like it's only useful in language, it can be that like it's not useful both in language in other domains, and it can be that like it's also useful in in like other domains. What what what what do you think? Alexia: Only only specific domains because you have worked a lot on on molecule and the LLM are horrible at this. Horrible, horrible, horrible. Even the one train on on molecule stuff or chemistry. It's all bad. It's all the baseline for any kind of L L or people have tried C O T. I remember we had a project before where Or t they wanted to try C O T, chain of thought for for molecule and stuff like that. Didn't work, just didn't work. and I've seen Ravid Shwartz-Ziv: Why? Alexia: so many 'cause yeah, it doesn't Ravid Shwartz-Ziv: Why I think. Alexia: seem like that reasoning about the molecule, let me add this part to it, let me add this part to it, because it does this or that and then it should increase this property. It in general it just just doesn't will work right now. It's much better to have specialized model only on working directly with atoms and building the building the thing. But the why you're saying why? it's just I don't think text is is that helpful. But it it's still chamels, they have some intuition. This they have some intuition. They will say, Well, I wanna add this kind of group to this atom because I think it will bring this property. So they do have this kind of insight. But it doesn't seem like it translate with LLM. So right now, maybe more data, maybe Ravid Shwartz-Ziv: But but do you think this is something like kind of like what so we talked with Frank Kutter about like a Tap Pfn like like two or three weeks ago. And he basically said look like like these models like didn't train to to good to to have good predictions. So let's train them from scratch, right? Let's try them on like some random data and then like fit everything in the context. and then it works really well. Like the transformer architecture is like is is good enough and give you like enough power in order to do all of that. do you agree with that? Do you think like this is something that is like not fundamental in the architecture but like the only the way that we train these models Alexia: It's really data dependent. It's it's data dependent. Like anything that is chemistry or physics, like astrophysics, I think there was a paper on this. Also at some point in this kind of domain I can tell you that scaling doesn't really work. Like scaling gets you extremely, extremely small gains. Like tiny tiny some so tiny that you shouldn't do this. At this for this kind of problem using RL with a small model, small small specialized model is better than building these big pre-training model. it's it seems that for for a very For stuff that is not well defined in language, that is more physical, like physics, b and chemistry. Like these smaller models are mu much, much, much specialized model are much better. Text doesn't just doesn't bring much games. Ravid Shwartz-Ziv: do you think like we can like we will have models like even not LMs, we will have models that will give us like good predictions for molecules? Alexia: And so far we don't see this so much, but yeah, as as they get better they should be a bit better, but it's very abstract thinking and for physical property about it's very complicated with quantum mechanics and stuff, so so using text is not a good medium for that I think. Allen Roush: So I I'm surprised, you know, earlier this discussion about LLMs being bad on on certain domains that seem amicable to next token prediction, do you think it's really just like a data problem or do you think there are certain types of problems that are fundamentally not amicable to like the the methods used with today's generative models? Alexia: I think right yeah, there's there's definitively like different data are not not good for but it's not the objective per se, it's the whole And like every data set need a different representation. There are better representations for different problems. Right now we don't have a universal representation. But text is our universal representation. Seems to work well for a lot of stuff. But not everything. Ravid Shwartz-Ziv: So do you think we will have unified representation that we work for everything? And what what is the the property what are the properties that we actually need from such a representation? Alexia: I I I'm not sure honestly, Ravid Shwartz-Ziv: What's the problem? Solve the solve all the the different tasks, right? Alexia: Yeah, you could try to get latent representation for I mean there's been work on on like let's say clip clip for so many different things. So you can do image, audio, video, or MRI, wh whatever you can do a bit of everything. So I think that's that's cool, but it's still you're s it's still limited like Because it's all it's it's how you you ingest the the data. it it's very like a latent representation is always gonna be imperfect, right? It's never it always has problem. It's not something so so right now it's really depend on the problem. Text is good for some stuff and I don't know if we can find a universal representation that is for everything. Ravid Shwartz-Ziv: But but do you think like we should try to look for a unified representation? Do you think like our models will work better if we have kind of like one global model or Alexia: it's nice to have unify. I I really like the kind of direction that we're removing let's say images instead of images, you work directly on the bytes. So that's cool because you have GPEGs. You just take the bytes of GPEG instead of images, then it can take the bytes of of raw audio, raw video. And that that makes things a lot cleaner if you're working on text like kind of models. But but there's no universal there's no no universal perfect representation. Like I don't even know if it's possible. I don't think it's possible. Ravid Shwartz-Ziv: okay I I want to to to go back a bit like to TRM maybe to to to talk a bit more. So you like we know that like right like tiny recursive models generalize past puzzle domains. Okay, but like what is the first real application you will bet on? Alexia: Well there's there's already a lot of people doing application. It's just not not advertised as much, but I had a lot of people contact me for for example, one was doing protein, so protein how they bind to receptors. You know, it's a classic key in the lock, so it has to fit. So one person was doing TRM to make sure that the the key fits in the lock. And there was some people doing power grid security stuff, to prevent D DOS on the power grid. There's a there's a lot of like time application like this. But of course I I know you're talking about the big language and multimodal imaging stuff. So so yeah, for so for that we just it's just about finding the right recipe and and building this way because it's it's not always clear the recipe. And one thing also is memorization versus like generalization like tool tool use. Like previous LLM were a lot memorizing memorizing stuff a lot. So so there was a lot of hallucination and problem but now they're using tools. So that's a good way to make the model smaller. So so if you have a more specialized model that can use tools Then they can get their answer from Google, from from anything, which which means that you don't need as much I don't feel I feel like with agentic work you don't need as big model as before because you don't need to store all the knowledge. It's just about leveraging the tools and under being able to reason properly about the output and how to use the different tools. Ravid Shwartz-Ziv: And w what about scanning loss? do you like know do you add aware to some scanning loss of TRM? does more recursive recursion keep helping and what breaks it first? Alexia: Yeah, it would be interesting to do the I've never done scaling loss 'cause you need you need crazy compute for that stuff. So I've always been more small scale but but yeah of course it's important to know how things scale at different levels and how things evolve. So we need recipe. Now there's a lot of like looping models. So now it's just about the different techniques, the different bags of tricks and looping is one, there's deep recursion is one. And so it's just about finding exploring this and understanding scaling a different level. Allen Roush: Do you do you think that different types of reasoning can be more helpful? Like do you think like for example are like debate strategies are an effective way to reduce syncophancy when you're doing like test time scaling type stuff? Alexia: So the debate between models, for example, you're saying population of models thinking together. Allen Roush: Yeah, yeah. Alexia: So yeah, for sure 'cause It's assuming it I mean population in general helps. We know that taking just the average across different prediction or the mode as the classic way to do test time compute, write a bunch of stuff, the most common answer that's the mode, that's the one you take. So reasoning across different model could help for diversity because there's always if you take the average prediction or the mode of of the answer then the thing is if you don't have a lot of diversity within your model then you're limited. So so there's always some kind of balance between diversity. So diversity it can help to have different model training on different things so that it's a bit more diverse even if less accurate through the the population, averaging to the population it will be it will converge to a a better estimate of the the transfer. So this kind of stuff yeah can definitively help. Ravid Shwartz-Ziv: okay we're we're almost out of time. let's go to several questions from the audience. what do you think mo the most interesting projects you have seen with people that taking the framework and like using it? Alexia: The in the recent time the most interesting project the there's the most interesting pro well yeah, hierarchical reasoning was super exciting to me. But otherwise I'm not so big on the the the next big L and the next b it's always it's always small tweaks, small difference, Ravid Shwartz-Ziv: But what about like a projects that like related to like to to RLM? Do you think like there is like like there were like cool TRM, sorry, a cool project that like used T RM for I don't know, different domains, different like cool ideals that you thought are cool very nice and interesting and Alexia: Yeah, there's a lot of cool stuff on test time compute for example with TRM and pushing pushing accuracy a lot. I've seen one work using math, so and then there's the HRM text also, but this one is less it's less novel because it's more like a classic LLM stuff. But but yeah, I'm I'm still looking for the next the next cool thing. Ravid Shwartz-Ziv: can you compare finite recurs recursion to fixed point equilibrium models like HRM TRM to a DQ or like EQR or like F PRM Alexia: Yeah, there's a lot of similarity in the origin like in the original HRM paper they were claiming to reach a fixed point. Yeah yeah, I find it's not always the case but it's not really reaching a fixed point but you can interpret it in this way and there has been work since then on on like making Like like doing loop until you have a fixed point ish and then then you move on. So you don't even need right now the meters a Q head that determines am I at the solution or not. But I know that this work was saying no you don't need that, you just you loop through you reach a fixed point. And that that's that's the end basically. And yeah. So so fixed point is big big big links to to to this kind of approach, recursion and stuff. Ravid Shwartz-Ziv: So thank you so much for joining us. It was a real pleasure. Allen Roush: Yeah, it was a real pleasure. Alexia: Thanks for having me. Ravid Shwartz-Ziv: Came here. Alexia: Alright, ciao ciao.