speaker-0: Yeah, now we're getting into the thought questions. Now, hierarchical models, you can treat them as the reward at the end of the final level of SPI. Almost all SPI, with very few exceptions, is dumb and flat, aka non-hierarchical models. There's a good reason for that. Hierarchical models are super challenging. To begin with, what is the... simplest hierarchical model that you can have. It's a two level model. The more, let me, what is your favorite two level model? speaker-1: So you mean, a kind of model on a kind of dataset? Okay. you know, one I really like and I started working with was something doing electoral forecasting in France. here the hierarchy is very interesting because you've got cities inside things that we call departments, but they are like, you know, regions. And then you've got the whole country. So you have that pyramid of three levels. I think it's an interesting one. speaker-0: Yeah, so let's stick with that. So we have to simulate this model across three levels, which are now, well, one level is always there. That's a flat model. But the new levels that you added, let's say the region and the country, now you have two more dimensions that you have to simulate, which may not seem very problematic if your simulator is fast, but suppose that your simulator is slow, like it's already pretty hard to simulate even one instance. ⁓ Suppose you're modeling the brain, right? And you have a brain emulator that takes a few minutes to generate one sequence. Now suppose you want to hundreds of brains at the same time, you have hundreds of days to simulate, and this is just one training instance. So there's no way you can train this model efficiently if you proceed like that. Okay. And this is also bearing memory issues and the need to design specialized networks. Now, what do I mean here? Right. When you have hierarchical models, you're modeling two different categories, at least two different categories of parameters. Like you have the local parameters, which vary. for example, stick with your case again by location, but just for global practice that capture what is shared among the locations. And you don't want to it separately, but what you really care about, if you are a proper patient, you want the joint distribution of all these runs conditioned on all the available data to get this precious shrinkage that you're after in the hierarchical model. This puts you in a tricky situation. when designing your neural networks. Because now you can't just say, okay, I'm gonna take all my parameters and put them as a single vector and say, okay, that's a high dimensional parameter space and neural networks couldn't deal with higher dimensional parameters spaces, images. No, no, no, you can't do this because your problem has a symmetry, right? So these, this joint posterior factorizes in a particularly nice way. If you think about it. And so you can actually ⁓ pose this problem as estimating each of the local parameters conditioned on each of its local data, but also conditioned on the global parameters.