speaker-0: inference for transformer probabilistic models. And actually kind of the tagline is kind of, is very short. the idea is that, okay, so I discussed before how it's important that all these models kind of preserve this set property. ⁓ So let's say, you can kind of move around your data points, but let's say your prediction doesn't depend on the ordering of your data because it's a set. Now it turns out that that's all very good and nice, but it makes really nasty effectively expanding this set. So if you have got more data, since it's a set and you choose this self-attention, which is between all the points, every point that you add, need to recompute everything from scratch. So there is no concept of, for example, KV caching, which is a staple of modern transformers and modern LMs. The idea that you can store some computation you've done so far, and you only need to compute the additional tokens that you learn. So the idea is that, of course, we could do that, but we would... possibly lose something. So the idea is that, okay, so what if we use this idea of a buffer in the sense that, so we keep the context fully set-based, fully self-attention, then we can add something which doesn't do full attention with the entire set, which would require us to recompute everything, but effectively only does what's called causal attention, which means that each tokens only watches the previous tokens. So now this is not permutation invariant anymore. So we are kind of breaking this permutation invariance, but we are doing it only for kind of this additional points that you add. So say, the idea is that you do kind of add new points very quickly and you can make predictions very quickly for new points. And then, know, once you're done with that, you can kind of, okay, you know, now you have time, you can put them back into your set. And now you can do, you can use low update in which you update everything. So kind of, you know, so, and this actually has a lot of usages because say, because it's kind of, so I'll say, you can keep your context, know, permutation invariant, then, you know, you can do a lot of kind of, even in parallel, you can literally sample, you know, thousands of points, you know, using this technique. And then, you know, let's say you can kind of merge it back, all the information back into your context. Anyways, so it's maybe in a way it's kind of more, yeah, a kind of architecture, more engineering in a way, let's say there was a lot of points that need to be solved, engineering wise, but to say, yes, it's something that we're kind of building on because say, it kind of lets us, again, it's an unlock because we can do a lot of things with this now. speaker-1: Yeah, yeah. Yeah, it's we definitely need the the links to these papers in the show notes. in so what what are the the cool things you we can do now with that? Like I what are you working on right now based on these papers? I'm guessing you're trying to add some of these ⁓ possibilities to ⁓ to some packages. So yeah, what ⁓ what are you planning around that to help people use it? speaker-0: Well, so the grand vision here is that we've been discussing, say my group, let's say, know, I have the, you I can show the logs. I've been talking about, you know, foundation model for inference for many, many, years. And ⁓ so that's something, you know, we've been kind of discussing for a long time. And that's always been kind of the, you the north star for us, for my group, especially, you know, building exactly this function model for inference. And in order to get there, yeah, we need to solve a lot of problems. Shout out of course goes to what have become foundation models for inference recently, mean, recently past few years, all the tabular foundation models have kind of, have shown us the way how you can build a foundation model for inference. ⁓ Of course, I they apply to tabular data, but I mean, a lot of things can be put in a table. So from my perspective, let's say they're kind of, they're general purpose foundation models for inference. Let's say, even though people kind of call them tabular data for historical reasons, I guess, but let's say, you can apply them to a lot of things and, know, they've been kind of doing amazing. We are very interested essentially in that direction, to expand in that direction. And there are lots of things that you can do, let's say, but you know, say, so of course, say there is a lot of new ideas that can be tried in the space. And luckily these models are still fairly small. let's say compared to LLMs. So they're still nicely within the reach of what we can do again, within academia. So we don't need to spend millions or hundreds of millions or billions to get training done. So that's quite exciting because we can still contribute. And for example, just to mention the buffer that I just mentioned, yeah, it can be applied to transformer, neural process, et cetera, but also to Tabla foundation models. One particular thing that you can do with this kind of fast prediction that doesn't require you to recompute everything ⁓ is unrolls, which are very useful for if you're doing reinforcement learning. So if you kind of want to do fast predictions to see what happens as you add more data points and you want to do RL on ⁓ this, what essentially would be built ⁓ allows us to make it like a hundred times faster or something like that. So it's, you know, with much, you know. with 100 times fewer memories. It's super useful. So that's something we're going to use definitely. mean, also others can use. just mention a direction. So a lot of the things we are doing are part of a grand plan. And we're building one step at a time. there is an arc we're going towards.