speaker-0: Hello and welcome to the Robotics Tag podcast. My name is Leo. And today we have a lot on our plate today for you guys. First and foremost, we are going to dive into the top robotics companies by revenue, and they may not be the one you are thinking about. Then you have probably seen the videos of people plugging Cloud or ChatGPT into a robot and wonder if that works. We are going to do a deep dive on that as well. After that, we are looking into speaker-1: My name is Oscar. And speaker-0: Google Gemini Robotics 2, their latest model for robotics. And finally we are getting into the exciting robot fights. They are all over the place, all over X recently. So we are going to check what's up with that. You ready? speaker-1: I'm ready, let's go. Maybe it's time to make a little top 10 list of the robotics companies by revenue. Just to give like our audience and also ourselves an overview. Because spoiler alert, it's not humanoids at all. These are 2025 full revenue numbers, but I also looked up the newer numbers. If they were published, they mostly paint the same picture here. Some of these manufacturers they also diversified, so I tried to separately count the disclosed robotics business revenue. And the first one is one of Leo's favorite. I mean, we had this one before. It is Intuitive Surgical with ten billion dollars of revenue. We know them for their DaVinci robot used for surgery and they have a bunch of different instruments and accessories for it. Because these instruments are often consumed or replaced as the procedure is performed, Intuitive has a strong recurring revenue model. rather than relying on one-time robot cells. Now the Da Vinci surgical system is primarily for minimally invasive surgery, but they also sell something called the Ion Endoluminal System to navigate the lung for minimally invasive biopsy. Currently they're rolling out the Da Vinci 5. The company mostly uses AI for imaging or navigation or training and the workflows rather than complete autonomous surgery. We discussed it. For the DaVinci, for example, they have the booth, and there's still a surgeon required to operate the robot, but it's still a robot. We had 11,106 installed DaVinci systems at the end of 2025. Globally, around 20 million patients have been involved in procedures with the systems. But we also discussed in one of our episodes that some people in academia and research are using actual unitry robots in combination with the booth, not the full surgical robot. And are almost able to replace some parts of the surgeries that are performed by the Da Vinci usual usually at a much lower price point. speaker-0: So it's like 2,000 surgeries per machine on average. It means you can kind of compute some kind of profitability index per machine at some point for the clinical hospital who buys it. And do you know if most of them are deployed in the US or because I mean we mentioned that the last time, I believe most of them are very much in the US. speaker-1: Yeah. I mean they do have some in Europe, definitely. In in Berlin for example, they have one. Globally US is focused area still. speaker-0: The most interesting part to me here is that approximately 10,000 machines and the company is those 10 billion in revenues. It means they do 1 million of recurring revenue per year per machine. So once you have a machine deployed, it's 1 million guaranteed revenue per year. Pretty good business. speaker-1: Yeah. It is. But now go on to number two, and that is EvoBax Robotics from China. And they're doing approximately two point six five billion in revenue. So it's a brand mostly known for cleaning service robots. And also the broader group owns Teneco, which sells wet dry floor cleaners and other intelligent appliances. The major product families though are the DBot, which is a vacuum and mob robot. And then the Winbot, which is a window cleaning robot, or the GOAT, which is a robotic lawnmower. And recently they also announced a new flagship, the DBot X11, and that introduced power boost charging, which allows the robot to restore the battery during mob cleaning intervals to let the robot cover very large areas without a lengthy charging stop. They're also expanding beyond this floor cleaning market. They have entered the pool cleaning. with the ultramarine robot and are also expanding into professional cleaning through the DBot Pro M1 and other models that they have for warehouses, factories, supermarkets. You know, I'm a big fan of cleaning robots. They are just so useful and they're really tied to this one use case because the third one is actually almost an identical company. It's called Roborock. And they're also from China and they're also doing these kind of cleaning robots. And they are doing 2.6 billion. So that is spot number three on the list. And for Roborock, the revenue grew by 50% just in 2025. Even more, 56%. And this is really the highest in the company's history. Now, how do they make money? Of course, also these cleaning robots, vacuum and knob systems, just like the one before, multi-function dogs and wet, dry floor cleaners. And they also move into robotic lawn care. Recently, though, they had a new innovation. Which is called the Zaros Z70. And this one incorporates a foldable five axis omni grip robotic arm. So it's basically this robot, this vacuum robot, and then an arm comes out of the robot. And the arm is able to pick up socks or move stuff out of the way. So this is more a robot in a traditional sense. speaker-0: Yeah, I've seen that before. Yeah, it's pretty useful. speaker-1: Yeah. It it also looks quite funny in my opinion. They spend quite a bit on R and D and they have really fast innovation cycles. So almost every year they put out a bunch of different products. We know this from Chinese companies. They're no exception here. At CES 2026, they also presented the Saros rover concept, which is a wheel lag architecture, and this helps the robot climb stairs and clean across multiple levels because Obviously, this is still a limitation for a lot of these cleaning robots that they cannot get above the ground level. Place number two and three, basically cleaning robots. Number four is the Fanuk Robotic Division. And they approximately do 2.5 billion. It's a Japanese company. They sell more traditional robotics, so industrial robot arms, controllers for welding, for example, and a bunch of other stuff. Also, automotive manufacturing is very important for them. But they also serve electronics or logistics. Now, what's interesting about them, they have a sort of a competitive advantage because they also have their own CNC systems that they provide, or servo motors, or other robot machines. And these go hand in hand with the robot arms, for example. For clients, it's quite useful. They can go to them and they can basically get it all from one supplier. Historically, they have been quite closed in terms of ecosystems, so they had a Very proprietary environment. But now recently they have been opening up a bit more, for example, to the robot operating system too to support that, or some Python interfaces. They also expanded their work with NVIDIA on simulation. And recently they announced a ninety million dollar Michigan factory and distribution center project, expanding towards the US more. And they have a new R2000 generation robot coming up, or I think it's already released, which has lower energy consumption. And a bunch of other improvements. These companies are not stopping either, you know. We see humanoids, new product category coming up on this side, and also these traditional robotic companies being like right on the matter of what is truly being used. And that's why they have these revenue numbers, right? Number five is ABB robotics from Switzerland. And they do 2.3 billion, a little bit more than that. And their business, this is quite interesting, is now classified as discontinued. Because they are having a planned sale to Softbank for five billion. So this is still subject to regulatory approval, but the completion is expected for this year. And we know SoftBank has been heavily investing in robotics companies. They're everywhere. Now ABB, what do they do? They sell industrial and collaborative robots, complete application sales, engineering software, operating software. So they're a bit of all over the place. And they have been, for example, integrating stuff from NVIDIA. For example, the Omniverse libraries into their robot studio hyper-reality simulation platform. They also introduced their high-speed power collaborative robot and a simpler robot-picking software called Pickmaster Lite. Also a trend that I've noticed in robotics, which is making robots easier to deploy, even for smaller manufacturers, even for less specialized teams. Then we have number six, which is Symbiotic from the US. Even though they reported a $91 million net loss, still over $2 billion in revenue. It comes mostly from warehouse automation systems. And they also have some smaller recurring revenue streams from software and maintenance. Now they use fleets of autonomous robots, high-speed lifts, vision systems to move cases through large distribution centers. Walmart originally selected them for their automation across all of their 42 regional distribution centers. And in 2025, Symbodic actually acquired Walmart's advanced systems and robotics operation. They also acquired ARMS innovations. So yeah, they have acquired quite a bit. Number seven is Yaskava Robotics from Japan, doing 1.65 billion. Yaskava sells the Motoman industrial robots. They also announced working with Google Gemini Robotics together, which we will come to in this episode as well, and a separate physical AI development with Softbank. And number eight is Cougar Robotics Division, 1.14 billion. It was originally a German company, but it has essentially been bought by China. It's part of the broader Cougar group. Now Kooker is especially associated with these orange arms. You might have seen them. Six axis industrial arms for automotive welding and assembly. Really heavy payload robots. And they also have a neural product which is more compact, and that's called the KR Agolus Ultra. And then we have number nine. Which is Kawasaki Robotics, with approximately six hundred and twenty-one million. So we're out of the billion range. Kawasaki, of course, from Japan traditionally had their strengths in welding, again, painting, pelletizing, but also clean room and semiconductor transfer and heavy payload robots. And a recent addition that they had is the CP11L high speed palletizing robot. So really interesting systems that all these companies have. And they're also expanding into healthcare, Kawasaki and Foxgone. Have developed something called the new robot, a nursing assistant robot. And they're also building solutions around the Hinotori surgical system. So this seems to be an area that they really want to pursue. They're also working with NVIDIA on a digital shipyard, which is quite interesting. I've never seen that before. And then finally we have number 10, and that is Autostore. Autostore is from Norway. They do 538 million. They're especially known for robots that travel across the top. of a dense aluminum grid and retrieve inventory bins stacked underneath. So it's almost like a warehouse or a storage solution that they provide. They said they have almost installed 2,000 of these systems and this underlying cube storage system concept that they have. This dates actually back to nineteen ninety six. So this makes Autostore one of the older companies in modern warehouse robotics or if not the oldest, I don't know. Yeah, these are the top ten. So what do you think, Leo? Any surprises? Yeah. speaker-0: I mean not surprised, but it's mostly companies you don't hear often about. What I'm surprised about is that Japan, you know, was very early on robotics. They are still very strong at industrial robots. And they were early on humanoids. They started making humanoids in the 90s and the early 2000s. SouthBank was investing into Aldebaran Robotics, French company who was making also humanoids, ⁓ who completely closed down last year in 2025. So I'm surprised that Japan especially considering its aging population is not heavily investing into humanoids for its domestic market right now. speaker-1: Let's see where they go on. I think Japan definitely has demonstrated that they have the knowledge. So maybe they have something cooking up which we just have not seen yet. speaker-0: Maybe. Alright. Are you ready for the next topic? Yeah. So this time I wanted to talk about wiring a foundation model into a robot. You have probably seen videos where people claim they control a robot or a robotic arm using cloud or chat GPT, right? Have you seen that before? Yeah. Yeah. The thing is, these videos are often a bit misleading and often quite often they cut out the thinking time. So you see a robot like stacking objects together like cubes and it does so in like 30 seconds but if you cut out the thinking phases that are in between each movement it actually takes like 15 minutes so it's not really a reliable option. So what I was wondering about is where does the intelligence sit and what runs at control rates. And in this case the foundational model is not necessarily in the loop in this case and that's what I want to address today. So Looking at the literature and everything I read on a daily basis, I could find three patterns. The first one is when the model is the controller. So you you get a model and pixels and the robot state goes in, and the model outputs basically a motor command, like ⁓ a robot movement, if you want. So like the model is basically the central brain and there is nothing else. Another one is a dual system. So you get a slow model, quite often a VLM. for vision language model who is planning at the top and a small fast policy that executes underneath, you know? Because the robots needs two at the same time. The robots needs fast command, fast action to react on. And the third pattern, which is more recent, is using a frozen prior and a tiny head, which means you use an existing model, you freeze it, you hope that it understands physics kind of. And you train a small decoder to read its latent state and emit action chunks to fit to the robot, right? So basically the first pattern, which is a large model controlling the robot end-to-end, is not really reliable and not really working anymore. Pattern number two, dual system, is what currently works, but the new research is heavily going into the third pattern, which is a frozen model and a tiny head, ⁓ you know, fine-tuned on top of it. speaker-1: And how is that different from from the dual system exactly? speaker-0: The dual system is one large model that will, you know, plan tasks and emit actions, and then it transfers that to the smaller model that acts in real time. The third pattern is when you use an existing frozen model that is basically a bit more passive, right? And then you only use the decoder to run live. Okay. So there is only the small part. I mean, yeah, you could see that as In the third case, the large model is passive, the tiny model is the only one active. But we get into that. Basically the question is where does the model sit in the control loop? And if you look at the first pattern, direct control, which is what we have seen ⁓ advertised by Waddle or Robocurve, for example on X, ⁓ which is more of a marketing play. Even even they recognize that this does not really work. It's more of a you know, a good way to make content. It's fun, it's catchy, but it's not really reliable. And even anthropic. In July, they published you know an article called Cloud Plays Robotics. I'm sure you have seen it around. They say that LLMs are not really good at controlling robots, but LLMs can be very good at helping a human write a controller for a robot. And that's about it. And the main frontier is in the latency, because a model like Cloud can give you an output at 0.2 to 0.4 Hz, while a robot on average needs 80 Hertz. So it's like two orders of magnitude of difference. So I mean 80 Earth depends on the robot, depends on the model, don't 100% cross beyond that. The need for robots is two orders of magnitude above what a model can output. So you cannot like plug into plug an LLM into a robot and hoping it works, right? And besides that, no LLM was ever able to make a humanoid stand-up. So just that is already enough to say. When it comes to object manipulation, best performance. From an LLM on Libero is 5.5%. LLMs always also fails at self-localization. When I say LLM here, I mean foundational models, readily available like ChatGPT or Cloud. And what's interesting is that larger reasoning does not help increase performance. So it means the model is fundamentally unable to perform this task, right? One thing though that can help the model perform better is to give it tools like a cursor to improve manipulations or a compass. To improve navigation, which is interesting because it means if you give it the actual data it needs to complete the task, it will perform better. You know? Second pattern, so two models, two systems, one large model, one small model, one slow, one fast. It's actually what Figure AI does with its helix system. So they have system two, the slow model. It's a seven billion parameter VLM, works at seven to nine Hertz for scene and language understanding. And System One, which is the first one, is an eighty million parameter model which works at 200 Hertz. And this one we would put wrist poses, finger flexion, abduction, torso, head orientation and everything. And as you notice it's two orders of magnitude, you know, difference again. So it kind of makes sense with the difference I was exposing earlier. But that that's what FigureI does. And it's also what Gemini Robotics kind of does. I mean you will get Deep into that later. Nvidia growth is also the same. It's a pair of model, a 3 billion parameter model, like they call it the backbone, and they have a 32-layer diffusion transformer action head. So it's like a much smaller model on top of it. And ⁓ physical intelligence, in one of their recent research called reinforcement learning with experience and corrections via advantage condition policies, said that flow matching models cannot give you policy gradients. So it means that. When you train a model, it can only get as good as the best example you have in the dataset. And you don't necessarily know which examples are bad because you don't get, you know, per support gradient, for example. So system one, which is the first one, can only be as good as what it is shown. So again, data is very important here. And that's where also physical intelligence is investing a lot of research. If we look at the third pattern now, which is you stop fine-tuning a giant VLA, you freeze Or reuse an existing model and you mostly train the small fast one. One approach is called patch policy, so it's a research from ⁓ Ian Lequin again and Pintor. And it's a frozen vision transformer with dense patch tokens that are fed directly into another transformer policy. And they use block causal attention masks, which means Over time, the model gets the current example of the timeline and also the previous ones. So it can deduct causalities over time, but it doesn't get the future one in training. This model it beats a fine-tuned Open VLA by 18%. So it's 18% better than Open VLA with 0.7% of the parameter. Yeah. And it's also the same thing that Flux from Black Forest Labs and Mimic Robotics is a Swiss company. speaker-1: That's not bad. speaker-0: Who makes gloves and hands. So they have a partnership. They made some research using Flux 3 video prediction model as you know a backbone. And they have a lightweight action decoder that reads the intermediate features of the video model. So it never renders the frames, they just get the latent features as I can understand. And they go from 55 to 95% success performance on the solved body kitting task. So it is pretty good, right? speaker-1: Yeah, it sounds definitely like a new approach, although I I don't really yet grasp the frozen model. Is it like a VLA or is it a a foundation type model? speaker-0: It is a VLA in this case, in most cases. So it is a VLA or it is sometimes first of all in the example of flux and mimic, it is an image and video generation model, basically. So the idea is that this model you use it as a backbone that understands physics and then you only train a small model to actually control the robots. Yeah. I see. While in the second pattern, the large model kind of makes action planning, very complex, ⁓ compute intensive tasks. speaker-1: Okay. speaker-0: And the second one is kind of okay, I have to act at two hundred Eaths per second. So I need to, you know, translate this large and sometimes abstract movement into concrete task that I have to do now, right? While in the other one, the larger model is much more passive. speaker-1: Interesting. Yeah, so these are kind of subtle differences in architecture that then cause actually quite a difference when you apply it to certain tasks. speaker-0: And what's not subtle about it is that for example, in the second pattern, you have to train both models to work together. In the third pattern, you only train the small one. And you can use pretty much anything as your backbone. So in two patterns, yes, you have one large model, one small model. It's actually the small model that controls the robots in real time. But in the third pattern, it's much less compute intensive at training. speaker-1: No, that's really interesting. Because bringing these LMs into robotics would be nice. But at the end of the day, we see with these types of approaches that's often better to not do that actually. And rather go the other route. Even though it's tempting because yeah the marketing and everything around it says, Okay, these models are capable. But no, when it comes to robotics, they're actually not that useful. So Yeah, exactly. speaker-0: Yeah, like you said, it's tempting, but it's not really the right way. Because what we learn from this approach is that feeding the action decover ground truth future video would increase performance to almost perfection. So the day we have a model that can accurately predict the future, you know, in video, so like a world model as a backbone, we know for a fact that this would pretty much solve robotics interactions. And Physical intelligence is kind of doing that as well. You know, they are doing real-time chunking and they generate the next action while the robot is still executing the current one. They don't want to change the model, they just skip video generation, shrink the decision head, and get a faster inference layer to get smooth execution. Some axes of research still remain, obviously. For example, force, the force channel is completely ignored. Contact reach is completely missing. We see it in GMDI robotics. You're probably going to cover that, but unscrewing a light bulb is 92% success. Screwing it is 36%. Anything that requires contact, touch, and force management is much more difficult because it's completely ignored so far. Tacti and Mimic, they try to fix that with a glove. Also, another interesting piece of research I've seen, I'm preparing a post on Twitter about that, is from Stanford IPLR Lab. It's called Fact. And basically a big chunk of the fourth wall is a training bug. It's not because it's a data or sensor problem. So basically we are usually under training the high precision part of the task and they fix it by reallocating the gradients where it's needed. If you want a robot to plug a phone charger into a wall plug, for example, maybe it will use all the gradients and all the precision to approach the plug and then fail once it's very close to to the wall. Because it's lacking precision there. While it actually needs some fine tuning or some fine grained understanding of the task where precision is needed. It doesn't need it when it's approaching, you know, the wall. And they are pretty much reallocating part of the training to where the task needs the most interesting. speaker-1: I think the wall plug is an interesting example. Let's say a part of the plug where you plug it in is covered, right? Or it's not even visible. Then sometimes as a human you might do 80% of the tasks visually by looking where you go with your hand, right? And then the rest you're basically just feeling out, okay, this is where the plug fits, and then you plug it in. Robotics really ⁓ forces you to think about these mundane, easy tasks in a very ⁓ specific way and to question them, are they actually that easy or is it actually a very sophisticated physical actor? speaker-0: Exactly. Something is that I'm still thinking, even after reading all this, is that the numbers you see they cannot be trusted in the end. Because you can see success rate, but you don't know what's underneath the original independent benchmark. Success is binary, it's not a good metric, you don't know what's under it. I still think we badly need a an independent benchmarker somewhere in the robotic. speaker-1: process. Yeah. Maybe that's the point where we can come to the Gemini Robotics models because obviously they also published their own benchmarks and their own success rates. Google just announced Gemini Robotics 2, a family of three models, obviously building on the earlier Gemini Robotics releases, which we discussed before in another episode. This is more the generalist robotics model approach that we see. You know, we've seen it with a couple of other Companies there are specialized models for specialized tasks, but the idea is of course that you have a generalist model or suite of models that work in conjunction and that can basically, without specific training, do certain tasks at a very high success, ideally full success, of course. In one of the videos on the website, they said they want to focus on three things. So the first thing is whole body control. The second thing is dexterous hands, so manipulation of objects beyond The simple pick and place task and collaboration. And collaboration is actually something that I haven't heard so much so far. And with this new Google model, I think this is gonna be a good setup for other companies to chime in and try to have multiple robot collaboration as part of their model because it's quite an interesting area of research, in my opinion, and underserved. So what kind of models do they have? So they have three models in total. The first one is the main VLA, and this is kind of like the big player. It's a vision language action model. It's called Gemini Robotics 2. And this one takes vision and language, so image and text, and then it outputs actions, aka motor controls of the robot. It is able to control four humanoids. In their examples, they mostly used the Aptronic Apollo 2 robot, which looks actually pretty nice. And they also applied to biarm platforms, so different embodiments. For example, they demonstrated that one checkpoint works across the Apollo 2. With a sharper wave, then the Apollo 2 with the inspire hands, and then also the Frankadoo with a robotic gripper. And Frankadoo is like two arms, it's not a humanoid. And the model is able to generate commands for this advanced dexterity as well. For example, if you take the sharper hands, these are pretty good hands, you know, multi-finger hands, a lot of degrees of freedom. And they really show complex tasks. For example, tying knots of a trash bag, sealing ziplocks, screwing and unscrewing bulbs. They truly had quite the impressive demo that they showed there. Just a demo, but it's still impressive. And also they talk about these whole body intelligence tasks, because as you just pointed out, even walking for some models is not something that you can just do out of the box, right? LMs are not even able to walk. So apparently this model is able to help the robot walk, but also crouch and stretch and do a bunch of other tasks. Now the performance is again self-reported. So for example, if they take the Apollo with the Inspire hands and they do a pickup from table task, they were able to achieve sixty-eight percent success rate. Or a pickup from floor task, they achieved forty-five percent, or a pickup from shelf, they achieved a seventy-six percent success rate. And that was interesting to me because you know, pickup from shelf sounds higher than for example, pickup from table, and pickup from table is obviously higher than pickup from floor. So it requires the robot to bend over, you know, and we know that bending over causes a lot of balance stabilization requirements that you have to the robot. So it's funny that you have something on the shelf and the success rate is higher because the robot basically doesn't have to bend over so much and destabilize itself. But that's just speculation on my side. Then for the Apollo with a sharper hands, and this is where they did the very fine tasks, which you mentioned before, the screwing of a light bulb. And there they achieved thirty six percent success rate on the screwing. unscrewing the light bulb they achieved 92%. So imagine that you asked your robot, hey, can you screw that light bulb for me? And the robot says, no, I can only unscrew it for you because that's the only thing that I learned. And then for the Frankaduo, which are these air arms, they achieved a pick and place success rate of 74%. I mean these are still good percentages. Again, they're self-reported, but ⁓ let's trust Google on this side. Matter of fact, you cannot use the model yet. Unless you have early access, you are an early access partner because it's still private for the Gemini Robotics 2. Now the other model, which is the second model that they have, is called Gemini Robotics ER2, and that stands for embodied reasoning. And this one is actually available already. You can find it in the AI Studio, for example, and it's a VLM. So this is exactly what you talked about before. This is the brain part, actually based on an LLM that's based on Gemini 3.5 flash. And It is supposed to control the VIA I mentioned before. So this two-level architecture, that's why they have this embodied reasoning, which the name already says, and then the VLA for the execution of the task. Now the embodied reasoning model takes much more input. It takes text, it takes image, it takes video, it takes audio, and then it outputs text plus tools. By tools I mean obviously it can control the VLA or other robot APIs. They also have a streaming variant of this one. For real-time interaction so that this model can really think while it's acting, which is quite interesting. You can also see the pricing on the Google AI Studio per token. It's not cheap, but I mean it's there, you can use it. So what is this model exactly for? As I said, it's for the reasoning. It breaks down the tasks, it plans multiple steps, it also tracks the progress from continuous video and also does self-correction if it's needed. Now what's interesting, they also showed The multi-robot collaboration with this one. So they had different robots. They had one task, and then they had the humanoid and then they had the arms, and they were working together, completing this task, coordinating the steps and handing off work, which is really, really impressive. And for this one, they have reported metrics as well. 57% on progress classification. So you see these tasks are pretty different. Progress classification, what does this mean? How good is the model at Saying, should I stop now or should I keep going? You know? Imagine pouring in water into a cup. If the robot doesn't know when to stop, it will just keep on pouring. And they also tracked another thing, which is moment finding. And that is specifically for that. 91.3% on moment finding. And moment finding again means like, okay, imagine a ball is being thrown up the air and you've got to catch the ball, right? Or even with a water example. When is the exact moment and before you have the progress? as it's going up towards filling the glass or whatever. And then they have the third model. And the third model is Gemini Robotics 2, but on device. So this is the VLA again, but just a version of it optimized for local on robot inference. No network dependency. We're gonna see what's more useful. In many cases, you can think about it being very useful to have low latency on device stuff, but maybe for more complex things you want to outsource stuff to the cloud and do more complex compute operations there. Now the on-device model inherents the motion transfer techniques as well. So even when the shapes, sensors, and degrees of freedom differ significantly. And for this, the status is unfortunately also early access partners or trusted testers. So since it's just a version of the VIA, you cannot yet use it if you're not a trusted tester. They also introduced another benchmark which is called Azimov Agenc Hugging Face. So it's basically an open test for agentic safety. If you give a BIA a task, is it refusing an unsafe tool call, for example? Or how good is it at assessing uncertainty? Can I even do this task or not? And also how good is it at proactively requesting human help? And that's why I think it makes sense to think about collaboration as well. Maybe robots in some cases are not 100% there, but if the robot can open the bottle, if the robot can place the bottle, And maybe even pour the water in. Well, you're still human, you can still drink the water yourself, you know. So you don't have to have the robot do everything for you. That's just what I'm saying. Yeah, they definitely also acknowledge the remaining gaps here. So, for example, the movement speed when it comes to movement is not human like, and also the multifinger success rates and very hard tasks specifically. As we pointed out before, independent testing, this has not really been done. There were some people on Twitter who tested it, specifically the public model, so the embodied reasoning model, not the full VOA. There was one guy from Robocurve, and in he ⁓ tried to replicate a test that Google had published on this new models, where they basically take a clapboard and they look how good is the robot at clapping the clapboard. And the results from this person on Twitter were that zero out of five tries it had success. The model hallucinated without touching the board and and then one time it even flipped an arm off the table, it damaged the camera, so it really didn't work out that well. They compared to Opus, interestingly, Opus five, which apparently scored one out of five and never hallucinated. But the catch for this test is that they had the embodied reasoning model and they tested that one and not the BLA. And the clapboard for example that Google showed was actually using the BIA. And there was one researcher also in the comments who pointed that out. So I just want to be clear, it's not the same one to one comparison here. But yeah, there were also other people commenting on this. They were playing around with the models and they said, for example, the embody reasoning could potentially be used as an auto annotator for robot data sets. So three new models from Google. Let's hope the two that are still private get released soon. And I think it's good to hear something from Google again. speaker-0: Yeah, definitely. I I'm very curious because obviously they did not communicate that much on it, but I'm very curious about the future of robot collaboration and h how is it going to be managed? ⁓ is it going to require like a shared thinking space or something between robots? Because in the end you want it to you want the robots to collaborate as if they were human. If you are three humans working on a task all together, suddenly one of the one of the person leaves, you know what you have to do to adapt, you know? It's not like you are losing knowledge with one person living and the other two can adapt and, you know, resume the task with only two people. And I think I would be very impressed to see robots reach that point as well. speaker-1: Yeah. And it will be necessary as well. Because imagine being a human. A human is not doing everything on its own as well. speaker-0: Yeah, so I guess it's time for the final topic of the day. Robot fights, right? And ready. So I'm not asking you because ⁓ mean you you don't want to fight a robot, basically. ⁓ as I such with the point here. Stephen Hawking said it that you don't want to fight a robot because you are pretty much sure to die. So right now we are talking about robots fighting other robots. And whatever you may think about it, we have to address it because you will see there are already several organizations. doing this and the audience is real showing up and paying for it, so it is still pretty relevant. Just to give you an idea, there is an organization called Rec R E K in San Francisco. They sold 3400 tickets into a venue that sold 2500. It was the biggest draw into that place entire history. And they hosted an event involving a Twitch co-founder and a UFC fighter at the same time. Another Institution called UFB sold out several events in San Francisco and Las Vegas. CES 2026 featured a BattleBots arena with a human referee. URKL, which is the Chinese equivalent, also hosted an event with two T eight hundred terminator-like humanoids from Ngin AI fighting each other. A lot is happening in this space. And it's not only humanoids, you will see. So the first one I'm gonna talk about, an organization I didn't know. much about before. It's called UFB and it's for Ultimate Fighting Pots. I'm surprised they did not really get sued by UFC already, right? And they were founded in early 2025 in San Francisco. You literally became an official hardware partner in November 2025. Like I said, the sold-out events in San Francisco, Las Vegas. And right now they are starting season one in the Bay Area. So basically like a season of different fights going on. They have what they called Ultimate Boat Studio which translates uploaded video of humans movement into robot motion. I've seen footage, but maybe it's older and not relevant of humans wearing a VR headset and holding, you know, controllers in the hand, and they were able to fight from outside of the ring to control the robots. Because yeah, in this case I did not specify that robots are inside an octagon, like in the UFC MMA events. And it's like this for all humanoids fights I've seen. All organization uses this setup. Same setup as human MMA. And more recently they switched to handheld controllers. The robots in this case they wear gloves. They used also to have some kind of head protections. So you know like amateur boxing. I'm guessing it's to damage the robot a bit less and to focus on reusability. And they introduced a new feature this season, which is pretty funny. It's like the robots have like some kind of balloon over their head which is illuminated and breakable. And the goal is actually to break this. So it's actually how you win the fight. It's not to destroy the robots, it's just to break this head that is completely replaceable. ⁓ Depends who you ask. But what they focus on in this case is they partner with creators. So most of them I I didn't know, I think they are mostly famous in the US. But for example, in one of the events, they had the guy from Linus Tech Tips, which made me know they do tech reviews. Basically, the fight is like a creator against another creator, each of them controlling a robot. speaker-1: It's boring. speaker-0: Which is the same robot, Unit 3G1, blue team, red team, and that's about it. And they focus a lot on yeah, influencer marketing. And then second organization, also in the Bay Area, is called Rec, R-E-K for Robot Entertainment Combat. I've also seen instances of robot embodied combat, but I guess you see the idea there. They were founded in 2022 by VR entrepreneur named Six, finally at Six Live on X. They raised Five millions with RobotStrategy, of course, because Robot Strategy is everywhere. And what they do is full contact where the robots are actually trying to destroy each other. So much more fun in your scale, I'm guessing. And robots are teleoperated by humans using a handheld controllers. They make use of a VR software developed by a UK studio called Reflex Arc. And they are also developing a simulator. That is going to be available on Steam very soon. So it's basically a video game. Play fat with your robots and the control in the simulator will be exactly the same control as the one in the real world. That's cool. Pretty interesting. And what's interesting for both Rec and UFB is that it's teleoperation, but with some kind of assistance. It's not like you control the joints individually. It's more like a video game where you press a button and the robot will punch or the robot will jump or the robot will kick. Because if you had to control the joints individually, it would be much more complicated and much less fun to play with it overall. So it's kind of teleoperation, remote control with assistance. And they hosted, you know, the events that was sold out that I mentioned. It was Justin Cannes, so Twitch co-founder against UFC veteran Hider Amil. It was an event in San Francisco. After that, they did a tour across five cities in the US. And recently, like over the past few days, maybe they are still there. They were in Tokyo. doing robot fights. And apparently they had a lot of success there as well. If I trust the video footage I've seen. And again, that depends on Unitary hardware. They mostly use Unitary G1. They apparently became experts at Unitary G1 maintenance and repair because they had to repair so many of them that were broken after the fight. So I think it's pretty funny. When you see the robots, most of them are pretty used up. They have crotches over because you know they just keep fighting. And ⁓ they also have one T eight hundred. But they mostly use it for show. They don't really make it fight. I mean sometimes they make it kick another G one just for the fun of it. But there has not been like an actual fight of a G one against a T hater, right? Because I'm guessing it will be, you know, way too unfair. speaker-1: I mean the T eight hundred has has a lot of torque. speaker-0: Yeah, T Angel Red has as much torque as a small diesel engine. It will not be fair at all. And also it's much more expensive. T Angel Red starts at twenty five K and if you get a good one, I guess it can go to fifty K pretty quick. A G1 is sixteen K. One issue though for both of these, UFB and REC, is that they depend on Unitary hardware quite a lot. Unitary, there's been kind of a ban, you know, not only on Unitary, but on foray made robots in the US recently. I'm wondering how these companies would do. ⁓ maybe we will get an Optimus Against Figure fight eventually. I don't know. Or against the 1X. But I don't know. I feel like the One X Neo would get, you know, beat up by pretty much everyone. It looks like a kind of robot that would get bullied by the other ones. I don't know what you think, but that's my feeling. So these are the two that are present in the US and there is a Chinese competitor that entered the ring literally recently and obviously based out of Shenzhen. Because that's where the Eastern Tech World art is right now. It's called URKL for Ultimate Robot Knockout Legend. Quite a loaded name, right? And it's actually openly organized and owned by Engine AI. So you see where this is going, right? Yeah. They organized one event that attracted a lot of attention in the robotic space. So many people I know in China have been. Apparently it was great to see it live. And I kind of feel sad. Have missed it. We are not sure whether it was a one-off event or will there be other future events? I hope there would be other future events. And it had like UFC level of organization with speakers, cameras, video recording, live shows, fireworks, everything, you know, just like an actual fighting event. And it was opposing two different T800, so basically two different teams. I have some issues about how they communicated about it. They said, okay, we are selecting like 32 finalist teams, and in the end. Only two teams, the one with the best algorithm, will fight. Because supposedly the robots are autonomous running an AI algorithm that fight autonomously. And I'm having a hard time believing that, and I'm not the only one to be honest. Engine AI, they have raised $340 million total. And their value that's more than $1.5 billion. And they file for Hong Kong IPO discreetly recently. So it feels like pre-IPO marketing. I'm not sure at all about the fully autonomous no human involvement claim because first it would be historic. It has never been done before. I covered previously in this podcast the difference between large models that doesn't show in real time and small models that runs in real time. And even that is difficult to have to make in real time happen pick and place or organizing ⁓ multi-step tasks and planning in the future is still very difficult as a problem in AI, in robotics AI. Saying that we have a live combat model that reacts live to what someone else is doing, to the robot's attack and everything, is very hard to believe to me. speaker-1: Yeah, I mean to be fair, I've seen the clips of it and some of the clips, you know, you see a robot basically not hitting the other robot for quite a bit and they always punching in the air and and like tumbling around each other. And then once in a while you see this kind of rehearsed looking action sequence of a kick or of a punch. And then it's almost like kinda the robot just plays that sequence and hopes the other robot is somewhere in the range so that it can hit. It raises the question, okay, how advanced is this autonomous model? If it's just a model where the robot is randomly walking in this octagon and once in a while does this like sequence, is that really a full autonomy or I speaker-0: I think it's kind of in between full autonomy and what Rec and UFB is doing. So basically, like assisted teleop and some kind of algorithm maybe that can output a sequence of ⁓ movements. Because ⁓ at some point during the fight, one of the humanoids lost its head. And they kept fighting. And when it lost its head, the other one started dancing like this. And what was funny is that at this moment, one of the commentators said in Chinese. speaker-1: Yeah, I've seen it. speaker-0: ⁓ okay. The robot lost lost his head, but the operator still has control. ⁓ like, what? There isn't supposed to be any operator, so yeah, still pretty weird. So I don't know. I'm guessing it's just a marketing play to frame it as autonomous, but it's not really the point. I mean, I don't care about autonomy on this, to be honest. I just want to see the robots fight, you know? So what I like about it is that they understood Ngin understood that robotics is inherently a visual business. They got that, you know. Very, very right, in my opinion. Also, something against you know full autonomy is that T Andred is famous for its move where it jumps with the fist in the air, you know, spins around as well. And I've seen Reg does it, I've seen URKL do it. It's still an argument against you know their autonomous gang. But it was a pretty nice event, and I hope this autumn they make another one so I can attempt. The last fighting league is called Battle Bots. And Battle Bots is very different because. It has been around since 1994, so 32 years. And they don't involve humanoids. It was originally televised, I think it still is televised, even though now it's available online, obviously. It was called Robot Wars. And you usually get two robots, sometimes four, in battle royale or team mode, it depends. And they they are in in an enclosed space of forty-eight by forty-eight feet or five by five meters. And The goal is pretty much to destroy the other robots until it cannot move. If the two robots survive the fight, you have three judges to score it. They have 11 points to provide, five on damage, three on aggressions, three on control. These robots are all human remote controlled. Quite often it's a team behind it, and more often than not, it's an engineering team from university, for example. Quite often also you see two operators per robot. One for the control for the movement, and another one for the weapon. Because obviously. These robots they have weapons and you can see so many different ones. Some of them have hammers, some other have flamethrowers, also you can see a lot of ⁓ high RPM saws and stuff like that. The rules are really strict about how controllable the robot is it and that the weapons are predictable, fail safe, and that you can always have a kill switch if it's need. And one of the main strategy I've seen. in these fights is that you usually try to go for the battery of your opponent. Because if you basically kill the battery, the battery will just, you know, ⁓ go on fire and the other robot will stop and from there you're pretty much run. So I've also seen that recently two teams are using AI, but again it's still a little bit difficult to understand how fully autonomous they are. One of them is a robot called Orbitrum. So it has two counter rotating spinners which acts as source pretty much, always enabled. So basically there is no control on the weapons. The weapons are always enabled. The ⁓ the algorithm mostly does pathfinding to go around the other robot or to reach it and to basically damage it. But even when looking at what the team is saying about it, I'm not exactly sure what part is autonomous and what part is not. And another one, which is even more recent, is called Nemesis. And it's a robot that in front of it has two kinds of wedges to kind of put the opponents on it, get under it. And then it has some kind of a saw on the back, you know, that swings over the the other romance. And it's a team from the University of Southern California and they claim full autonomy on both weapons and movements, which is much more complicated when you think about it. Like I said, no humanoids, the most dominant designs here are usually pretty low. They have a wide base and they are pretty heavy and they have some fast-moving weapons. Because if you think about it. You don't necessarily get that many chances at tackling the other robots. So you want to make as much damage in one go as you can. So you need a heavy hitting weapon once rather than a soft hitting weapons and hoping you will hit the opponent ten times. To me it makes me think that the mature dangerous and solved part of robotics is the body. The body is pretty much sold. We know how to make robots that can, you know, attack each other and ⁓ you know, be kind of violent and powerful against one another. And I also like that they have some several decades worth of knowledge in making robots resistant, face-safe, and also how to destroy one another. And I find it pretty sad that they did not set up a data acquisition pipeline around it to gather all the data so that everybody could, you know, profit of it and you know make more resistant robots in the future. Also for humanity is typically data you can use for anything else. speaker-1: Well, who knows? Maybe they will announce their own model soon. speaker-0: Maybe BattleBot Autonomous Robot. And yeah, I think it's interesting to see that the design is completely anti-humanoid. Anything that has to work would be, you know, much more complicated to get to work in such a scenario. speaker-1: Yeah. Now I've seen this show when I was a kid and I used to love it and I still love it. It's it's very entertaining. And even though a lot of companies talk about okay, the applications for humanity and we help everybody, but at the same time you see the the human brain instantly defaults to Okay, let's let's make these things fight So before before everything else is over. speaker-0: I'm surprised there isn't a drone fighting ⁓ tournament. Yeah, races but not fighting. Drone against drone. It would be crazy, man. speaker-1: Yes. They're races, yeah. Fighting will be crazy. Yeah. So many ideas for this. And almost like small battlefields, you know, that whereas like a capture the flag ⁓ scenario and then you have two autonomous drones and a team of two humanoids and then maybe like a couple of robot dogs and then they fight against each other and it's basically like a war game. speaker-0: Yeah, exactly. In the real world. Advance warfare. Robotics Advanced Warfare as a game. Yeah. Cool. So I think that's wraps it up for today. Thank you guys for watching. ⁓ this was the Robotics DAC episode number twelve. Please subscribe to our channel. Don't forget our presence on X as well. It really helps us if you do. And see you in the next one. My name is Leo. speaker-1: My name is Oscar. See ya.