Basil Chatha: Yeah, thanks for coming. Today we're gonna have a talk on computer use agents. So yeah, we have some pretty amazing guests here. it would be awesome if you guys could introduce yourselves. Cool. Hi everybody. Thanks for ⁓ being here. My name is Lucas. I'm a member of technical staff at Anthropic and I work on computer use. Hi. My name is Eric. I was ⁓ early century, like first century GTM, and now I'm f I was first or founding customer engineer. Is that what people say? At kernel. ⁓ hi, I'm Reagan and I'm a founding engineer at Browser Use and I do like a bunch of stuff there. Engineer. Well, could you guys explain your companies a little bit as well? ⁓ sure. Yeah. Well Anthropic. Yeah. Anthropic is a startup that ⁓ creates like large language models. No, it doesn't create the machine god. ⁓ kernel does cloud infrastructure, specifically browsers for agents. So we do all parts of the browser ⁓ browser agent stack. So browsers as well as web agents and LLMs. Awesome. So yeah, could we start with maybe like what's one interesting way that you guys have used computer use agents yourselves in just like day-to-day life? Or like what's an interesting use case you've seen your customers use computer use agents for? ⁓ one fun thing that we have at our office, we call it fridge use. So we have basically a list of groceries that we put the fridge and we take a picture of it and then that picture gets sent to a browser use agent that buys everything on Istacart. That's kind of interesting. For customers. I saw this really cool use case where people somebody scraped, you know, those like job lists, like, ⁓ this list of 10,000 internships you should apply to. And I was like, how do they ever acquire 10,000 internships? And then I saw somebody actually do that with browser use and then compile them all into like a huge Google sheet. And so that was ⁓ relatively interesting. Yeah, for us, it's mostly our more sensitive or more paranoid customers as ⁓ they are using some websites. ⁓ they prefer to ⁓ you know, you know use the computer use controls and be able to interact with the webpage in ways that are a little more human behaviorally ⁓ detectable. Yeah, I think ⁓ cool uses. I mean I one that I think that I see a lot that's interesting is obviously using like just legacy systems because it's kind of interesting to see like a very state of the art model like trying to manipulate some very, very old software that would never have any kind of API or CLI or MCP or anything. So that's a cool one. ⁓ another cool one I think is a lot of people use these to play games. And so it's fun to see how the model will like play a game, try to like you know, figure out some way to speed run it to the end or something like that. So I think that's a more quirky and kind of fun way to use computers agents. Interesting. Like so I used it just yesterday when I was accepting people to ⁓ this event. I think Luma has an API, but you know, I was just too lazy to have it like write a script for that. So I was just like, okay, I like log into my Luma and go through everyone's profile and figure out like who would get the most value out of being here. And then I like accept them. So it like found like sixty six people and Yeah, I just like accepted you guys. So yeah. Awesome. ⁓ no, it was like people put your like a lot of people have titles on their like Luma profiles. They would like look through like, ⁓ what company do you work at? You know, what do you say in your title basically? And just based on that. ⁓ if you didn't put anything, I'm assuming that what it would have done is sorry, this person's not gonna get a lot of value out of this. So yeah. So computer use agents, can one of you guys define them and maybe we can just talk about like the full landscape? Like what does the entire like computer use agent landscape look like. So yeah, who who wants to start? I think yeah, computer use agent is traditionally defined as like yeah, some agent that's primarily interacting with other systems through computer use. I think like in in today's paradigm though, like computer use agent is almost too specific and computer use is more like a feature or a tool that a general agent has amongst other tools. So saying like computer use agent is almost like too restrictive to one way that that agent will interact. you know, with software or whatever it may be. And it's almost like an agent that can also do computer use as well and browser use and call CLIs, et cetera. So that's my take on computer use agents at least. Yeah, I actually very much agree with that. In our customer use cases, ⁓ you mostly see a blend. You rarely will have somebody that is purely using a computer a computer use agent for browsers. For example, you probably want to Have a lot of fallback on the DOM and CDP and use computer use on in selective occasions. Yeah. Computer use you can just imagine is like anything a human does on a computer. So that means like creating a file, also interacting with the browser. That's kind of generally what computer use is. Browse use is kind of like a subset of that. So you can imagine in a typical browser, browser use or computer use flow, I say you, I don't know, went to Wikipedia, scraped all that information and put it into a file that does the computers part, the computer part and the browser part. of a computer use agent. And then the agent part just means you run it in a loop with LM. Okay. If you had to describe like via an example, like how does computer use work, could you walk through that? So like one example that I saw Catherine talk about in like some of the like podcasts she's done. She talks about like how she worked at Cash App and how they do like QA. Could you talk about that? Yeah, of course. Yeah. So this was the sort of ⁓ seminal sort of moment of Catherine when she was working at Block. ⁓ Catherine's the founder of Kernel, by the way. Yes. And ⁓ she was working our block or Cash App, whatever you want to call it. And ⁓ they basically had to do verifications ⁓ that the sort of pay with ⁓ block was working with across all these partner sites. And these are cases where you need to go to the website actually and navigate ⁓ capcha and like navigate ⁓ pro like the actual network layer and you have to make sure things render properly. And ⁓ that's really painful. Run running browsers. basically is a terrible, terrible thing. And that's most of our my my life. ⁓ and so that's ⁓ that was the first use case for Catherine. And it was one of those cases where she was like, this needs to be a thing. This is too much effort and too much work. Sorry. I I just that. ⁓ similar on the QA pipeline, QA is a big thing. QA just means quality assurance, which means you know, let's say there's like a login flow, you want to make sure that user can actually use it and it's functioning properly. So one thing that we've ⁓ done at browse use is well a lot of people use us like labs use us for QA For website generation. So let's say your cloud, your Gemini models, they create this super cool website, but half the time it doesn't really work, right? Like sometimes a button's broken, but it'll look really cool. So they need to test how test if these things actually work. So they use browser use as a judge in RL environments. Yeah. So like how does a computer use agent or computer use actually work? I can tell you for us, basically it's largely it's usually through visual. It's like you take a screenshot of what's going on in the page. And ⁓ you pass it off to the agent along with the original instructions, of course. And then the agent w has a a series of controls that actually is largely defined by Anthropic, ⁓ that the agents know how to send back to ⁓ the browser basically. It's usually like click, it's type, th those kinds of like human ⁓ focused controls. But y maybe you wanna talk to more about it. No, that's pretty accurate. I would say maybe like if you're doing strictly computer use on your computer, something that's maybe different from other tools is like Let's say you wanted to use web search with an API. Most model providers will provide that on the API side. So you just enable it, you ask the model to make a web search. That happens on the provider side. The model gets the web search results. With computer use, the model actually sends a tool call back to your computer, which might be like, you know, click on this pixel coordinate. And then part of your like part of your job is basically to ingest that instruction and actually make the click happen. So That's maybe a small way in which it's different than some other tools that API providers ⁓ provide. But if you're doing it on some hosted infrastructure or browser, then in that case that tool call would go to you and then you would handle the clicking on that browser. Yeah. Anything to add? Yeah. And specifically ⁓ also for computers, some people might have heard of things like playwright and CDP. ⁓ for browsers, specifically for for kernel kernel, for example, a lot of those instructions go down to the Chrome DevTools protocol, which is basically ⁓ as as humans interact with the web, that click or the you typing gets actually translated into some code, ⁓ which is the Chrome DevSo protocol. And so browser agents tap into this protocol and are able to click and do everything that a human can. And so yeah, and Playwright also does this. It's built on top of CDP. So how much of it is driven by like straight up the screenshot and how much of it is driven by like the DOM or things like that? And how has that evolved since like maybe when computer use agents first came out in like twenty twenty four? I think you guys came out with it first. So yeah, how How has that evolved over the last couple of years? I think there's like two different schools of thought on this. So some people prefer some people think like there's pixel ⁓ corner clicking is the best way to go about it. And that's only screenshots. ⁓ some people also bet on the DOM plus ⁓ screenshots. And so, for example, for us, we use both a DOM and a screenshot. And what that allows us to do is and and in the very beginning, ⁓ to answer your to your answer your question, DOMs are really massive. DOM is like all the HTML and CSS you have on a page. Those are really massive. And LMs used to not be very good at ⁓ managing the context windows. And so o over time the usage of the DOM has gotten much, much better. And so yeah, we we are subscribed to the belief that DOM plus screenshots is the the way to go. Meaning like you basically put everything into the context window. Yes. Yes. And screenshots aren't as heavy as you might think. But yeah, it's mostly ⁓ managing the DOM and we manage ⁓ we strip down a lot of it and make sure that we're only using what we need. So I'll say a large part of it's a DO I I I don't know if I could put into like percentages, but using DOM without the screenshots often gets the same results that you would. Yeah. The models are quite good at the code part. And DOM is effectively it's it's code. ⁓ and that is usually the most reliable way of going about it. However, like it's just like when you think about like human in the loop, ⁓ you still have these humans are in the loop. As things change, as people deploy new websites, you need to be able to ha have some layer of like, when things break, why did they break? If the DOM is not guaranteed to be stable, which it's not, right? It's It's not a API. ⁓ you need to have some kind of way of doing fallback. And inevitably the websites, they are still designed for people. They they're meant to be used by he people. And computer use can be a great fallback for that case where you sort of are guaranteed that the screenshot will show something that is intelligible to a computer and therefore intelligible to computer use. Yeah. Nothing to add there. I think plus one to both. ⁓ I think to your point, I think hybrid is probably the way forward. And When you have like QA use cases like what you mentioned earlier, you might want the model to do screenshots that it catches maybe anything that looks or renders weird. But if DOM is more efficient, then it should certainly use that. So great summary. So do you guys use just like the standard benchmarks like OS World when you guys are trying to improve on them? We report on OS World, yes. Internally, is that kind of like what guides your development or are do you guys have like your own internal benchmarks as well? It's hard to comment on that, but we definitely use OS World a lot. Okay, okay. All right. Guys, ⁓ we have things. Yeah, we we we have we have many, many benchmarks. We're a big open source company. We have many, many benchmarks things for stealth and for web agents. So we have things like BU bench. There's also a bunch of other benchmarks outside of OS World, like the Odyssey's benchmark and all those things. So yeah, we we we use benchmarks a lot. Yeah. So I think recently, like within the last month, ⁓ OS World First of all, maybe we should explain like what OS World is actually doing and then we can talk about OS World ⁓ two point ⁓ So yeah, does do any of you guys want to explain it? Maybe like what are some of the canonical examples? Yeah, OS World basically puts like a computer use model or computer use agent in a container and it's then asked to do different tasks on that container. I think it's a Linux machine usually and it's stuff like Edit this photo in this like free software, create this file, like open this Excel file, et cetera. ⁓ and so it's basically a range of computer tasks done in isolated environments. So I think the best models at this point are at like eighty percent. Like what does eighty percent mean for OS World? Yeah, it's likely near saturation. Like with any eval set, there's a real ceiling, which is like eval sets will have some things that are incorrect. And so even if your model is perfect, it'll hit some ceiling, like it can't always hit a hundred percent. So I think basically what it means is OS World V one has been saturated. And so is a really cool journey that models, I think the first OS World score was our score on like 16% or something like that, which then a year and a half later the evolve's completely saturated. So I would say it's cool to see like, you know, this set of tasks, which is like I would refer to them as intermediate tasks. Like they're not extremely simple, but they're also definitely not the hardest tasks that humans do on computers. And so seeing that models have already achieved at least that level is like very, very impressive. Like the model can definitely use a computer better than like my grandma can. Sorry, grandma, for calling you out here right now. But my grandma sucks as a computer use agent versus Claude is now way better than her. So yeah. I think OS World to some degree kind of models that, right? It's almost like maybe like the average person or like a little worse than average person's ability to use computer. And so I think what you're leading up to is OS World V2 is coming now, which are kind of like shifting that difficulty of tasks into more like intermediate to advanced tasks that humans might do on computers and making the evolve much harder so that we have more room to hilk on. Yeah. So like what is harder about OS World Two trajectories than maybe like OS World V1? Yeah. So it's typically more like requires you to manipulate multiple apps, especially sometimes they might rely on each other, like changing the data in one app might change what happens in another app, et cetera. It's also longer trajectories. It's also more tasks. Like I think the original OS World had four or five categories of tasks, like different software platforms that tasks would be done would be done on and OS World V two expands on that as well as making the trajectories longer, as well as making you need to use multiple at once. And so ⁓ yeah, in general, just much richer and I would also say more realistic to how like advanced users that are using computers, maybe like knowledge workers, et cetera, use computers. Yeah. I think ⁓ the median task on OS world v1 takes what like twenty minutes or something for a human and OS world v2 is like an hour and a half or something. And like the what is it? Like the the best model on OS world v2 is I think Opus, right? Opus 4.8. Yeah. It's at like 20%. So like why is it so why is it so low? And like what is the ⁓ path to increasing that? Yeah. So I think like as you make these trajectories very like long and difficult, like the model can maybe do things in different ways or get very close to that end state, ⁓ but not get that exact answer just yet because they are quite long convoluted, like basically trajectories that the model needs to take. And so if you look at like the partial grading score, it's closer to like fifty or s fifty something percent. But I think that that's generally a good thing. Like you want an eval that is like so difficult that your current models right now are like struggling to get good scores on because it means like There really is headroom to improve and it helps you improve. And so I think ultimately it's a pretty good thing, yeah. Because if any takes. Yeah. I mean you guys are training your own model, right? Yes. We have we also have one open weights model. But those are those are all optimized for browser, browser use, but not not for general computer use. ⁓ as OS world ⁓ yeah. Yeah. So like for a browser model, I guess. Why are you training your own your own model? This is largely for of course the reason to train train any model for yourself is to make it tuned to your use case and to make it cheaper and faster. And so for example, to ⁓ the open weight model that we have that's based off of A quen model is it's it's super optimized to the browser's library, which has just around 100k stars. Awesome. ⁓ and so that means that it knows the tool call structure much better and can just do things much, much quicker. And it's it's super, super cheap. And that's that's the number one, yeah, that's why we train our own models, just be more accurate and know know our stuff better. Yeah. One other thing I wanted to ask you, I guess like that reminded me. So yeah, I think Opus it performs the best in OS world ⁓ two point z, but it's also like very expensive per task. ⁓ when you compare it to like another lab that shall not be named, ⁓ their model is not as good, but the like cost per task is like it's just much more token efficient. So like how do you make it much more token efficient going forward? Like what are the levers that you pull? You you don't have to talk about like specific things, but yeah, yeah. I think levers to pull are ⁓ primarily also like effort value is like the new thing, and that's kind of across all labs is like really dialing in your effort value. Feel like most people like to use extra higher max at all times, but that's almost the equivalent of like jumping in your car and just either flooring it or not at all times. It which is like in theory, you get to places faster, but in practice, like that doesn't really work and may lead to accidents. ⁓ and so I think dialing down the effort level definitely helps. Of course, specific prompting for different models also helps, but in general, I think it's a good thing that different models kind of have different sweet spots in terms of like if you're somebody who is like prioritizing cost and efficiency, maybe pri like provider A is good for you. And if you're somebody who's prioritizing like net quality and like tail end reliability, then provider B is good for you. So I think ultimately it's a good thing. And you know, each model family you can again kind of tweak the prompting, the effort value, et cetera, to kind of land somewhere on that quality cost curve. I see. So the levers are basically just a reasoning effort. Is there anything else or not really? ⁓ yeah, I mean you could get really fancy with it. Like you could try like Downscaling your images before sending them so that it's less pixels that the model is looking at. But of course, there will be some trade-off with quality there as well. You could define your own computer use tools to make them either more narrow or more efficient in some way. We see that a lot of customers like to define their own tools and that may help as well. So you could get really creative with it, especially since a lot of that like tool interaction is happening on your computer as well. There's also like strategies with like how you cache the conversation as well and cache breakpoints. So you might not want to send every single screenshot in that history back to the model at all times. So you might want to have like a rolling cache of like the last three images and you break hash every N turns. So we have a ⁓ like best practices actually that we ⁓ wrote up and released like a month ago that covers some of the like kind of more advanced levers you have for ⁓ tuning cost to quality there. But nice. ⁓ so I guess to Kernola, why is fast infrastructure is so important for these like computer use agents. So I've a lot of companies that I've worked for in the past, like we've they've focused a lot on speed in an early stage because it's often a signature for code quality and especially in the infrastructure layer. And so like we focused heavily on making sure our infrastructure is very solid and a performance is like a good benchmark for measuring infrastructure quality. And so like we recently released a blog post talking about all the things that we've done at the sort of Both at the orchestration layer, but also at the actual like per machine, like how browsers are snapshotted and then loaded from memory and ⁓ w the gains that that gets us. Because when you start to spin up large amounts of VMs sort of on demand, ⁓ you actually hit some actual machine level restrictions. And so that's I think and so like but like why it matters for like these agents, I think a lot of it is like we probably all have noticed agents are kind of slow compared to Code basically. And so like we have spent a lot of time thinking about how to make it so that when a user comes up with a task for us, that the browser is as ready as possible. And we can shave, we can't shave any more milliseconds off of our startup time. We're already already in the two digits of milliseconds. We spend a lot of time with like managed auth so that your your authenticated profile is ready. So you don't have to go through the login if necessary. And you can also imagine with like with these like benchmarks, like, You really want like with like context and the way that things fill up. I'm actually I wanted to ask you a quick follow-up question on that. Like you want to make sure that the agent is able to accomplish the goals as simply as possible. So you wanna prepare the browser as quickly as possible. And I wanted to ask actually, I'm just curious. ⁓ with these like benchmarks that like is it one serial agent conversation going through this entire thing or does is it like is it better to have like an agent break down a task into multiple parts and then task this task them to sub agents? is like I I don't know, like I'm just curious how these things are done. Yeah, I think the standard is typically like one agent going off and doing one trajectory. But you can imagine as they get better at computer use, you might end up with like such a big task that you have a coordinator agent breaking it down into sub agents doing the tasks, certainly. Cool. I wanted to maybe push back on this speed point. One thing that I like one thing that all browser companies optimize for is like startup time speed. And but one thing that we've noticed, at least at browser use, is that This number isn't also important. ⁓ because you can imagine, like I just said, agents are quite slow. And so there's no real benefit to starting your browser in, you know, 100 milliseconds versus 50 milliseconds versus 200, because you know, you need to attach the agent to it anyway. But one thing the the more important metric for browsers at least is stealth, ⁓ which basically means how often do you solve captures and how often can you get past ⁓ the antibots of websites that may detect you. And so yeah, the I would say that the speed and the startup time isn't actually also. also important. So on the anti bot stuff, it seems like, you know, every every company has anti bot captcha stuff for a reason, because they don't want bots on their website. So like doesn't that seem like a headwind for your business? Why what what are you guys doing to stop that or g get ahead of it? I guess part of the fact that we ex I I'm standing here today means that I believe that the bots are valuable. I think what matters is that like you represent economic value to the end customer and into these websites. And so this is why I think we've heavily invested in managed auth, because like it's the idea that's like, ⁓ there's a human out there attempted attempting to get something done. Hopefully that thing that they're trying to get done is of some economic value to the world, which means that like the more that we can be aligned with an end user, a human out there, the more that our traffic should I ideally be aligned with the goals of these websites and that we shouldn't be bot ⁓ bot blocked. But in the meantime, it is like ⁓ it's almost a a a difficult path to to go down. And when you look at anti-bot, ⁓ well, though the approaches, there's obviously the simplest, which is like recapture. There is sort of like, you know, they're solving these little puzzles. But over the next ⁓ you can start to see it happening already. There starts to be more identification signals around like, ⁓ like for this particular user, like where are they coming from? What is their patterns of behavior? What is the fingerprinting of their browser? And there's been a a lot of research out there, and it's something that like is an evolving s I think an evolving conversation between like the websites that we work with and ⁓ sort of how we ⁓ sort of work around the existing controls. Yeah. One thing to note, recapture is not the simplest recapture is actually the hardest capture. Recapture V3 like business is actually the hardest capture to get past. ⁓ and but ⁓ side of that and to answer your question, ⁓ Basil, for us at least we also really want to align with the objectives of these websites. For example, Amazon we we also want, you know, somebody to purchase something. And ⁓ a lot of a big part of that is of course being an actual real human, which means logging in with your cookies. So for example, yeah, when when we do stuff, for example, for me personally, the way I interact browse use is I have my cookies uploaded to the cloud. And so I pur when I purchase things, I purchase them on my real Amazon account. And so it's Amazon, it doesn't feel like there's anything different. One thing I also will note is that ⁓ there's a lot of websites that really need automation. So you can imagine those really, really old healthcare websites and those insurance websites that you really that are kind of made to be bad for humans. You don't wanna, you know, interact with these things. So I I would like to make the case that sometimes bots, yeah, are are actually good. And you know, there's there's the other case where, you know, they're not so good, but we we try to avoid those use cases. Yeah. one company that also shall not be named. ⁓ they told me that they are thinking about working with some of these bigger websites to do KYC. Have you guys thought about that? Yeah, I think like Cloudflare obviously has done a lot of work in the sort of web bot auth world, and that's something that we're working with them on too. Like I I think I'm sure sure you guys are as well. And ⁓ I think like KYC is something that we're actively investing in. Yeah. K KYC is like a big new topic. It's certainly quite difficult to to do KYC give like maybe for like the top twenty domains of like, you know, LinkedIn and ⁓ these websites. I'm not so sure that this is going to happen. I there's like maybe like, you know, I I I can't really predict it because for example, websites like LinkedIn really hate when you automate them. So it's hard to imagine that they'll change their mind. But Honestly, there's also a really significant chance that because of how many bots there are on the internet, they kind of have to conform. So yeah, definitely a big it takes a lot of mental space for too and thinking about this KYC stuff. Are there metrics on that? Like how much of the internet right now is like bots and how has that changed over the last couple of years? Yes, Cloudflow actually reported recently, I think last month, the amount of bots on the internet has ⁓ now surpassed humans. When the old trajectory cited by SEMrush was they were predicted by twenty twenty seven. that the that bots or twenty twenty summer twenty twenty eight that bots would overtake the traffic, but it's ⁓ it's it increasing exponentially. It's kinda crazy. Yeah. Yeah. Seems like inevitable kind of that they have to do this KYC stuff because I mean that number's only gonna increase, right? Yes, yes. ⁓ I'm yeah, I'm not so sure. I think that humans, at least for the next few years, will still be interacting with websites. It's hard to for for me to imagine a world where I wouldn't be interacting with a website because For me at least, I like the way my Amazon looks and the way my Walmart, my target looks. And you know want to be able to click around and do all these things there instead of having to interact with, you know, Claude or ChatGPT. Sometimes the interface is really nice. So yeah, I I honestly I I couldn't tell you if if KYC is gonna be a ⁓ humongous thing in future. Interesting. I think it's a migrations ⁓ question, which is like a question of timing. Like you see, like you can say like, ⁓ like all everything should be I I should be able to access all my bank accounts on my phone. Which is true. ⁓ I should be able to see that kind of stuff. ⁓ has ⁓ Bank of America done a great job of building APIs? No. But in the meantime, you have people like Plaid and before them another company called Yodoli that were that were building, they were scraping the websites to bring the APIs forward. And they start doing it and it's somewhat gray at first, but eventually ⁓ the banks become come into agreement with these these API platforms that are meant to actually bring them forward. And so I think that's how I think about browsers and I think there's ⁓ like with KYC or the anti-bot stuff, they're sort of both sides of a sort of a dialogue as to how these more legacy sites, ⁓ like you said, like the insurance or healthcare websites, are slowly brought forward into like the new way of interacting with our own information. So I wanna shift the conversation a little bit. One thing that I've heard a lot of people say is like you can just reverse engineer the the back end of a website and you can automate it a lot better than by using this like take a screenshot, figure out what pixel to click on and then go to the next step. What do you guys think about that? Is it a better way to do computer use or browser use, I guess? I I expect computers. We had a product that did exactly this called Skills, where we went to a website, re reverse engineered the APIs and then did things deterministically. One thing that of course the main counterargument against this is that websites change quite a lot. The API changes quite a lot. So you know you don't want it to break. And Yeah, that's the that's the main argument against it. Also, you know, not everything is accessible accessible via ⁓ API. So for example, when you click on something that creates a pop-up, that pop-up usually isn't registered via API, for example. And yeah, so I would say that the the main thing is you want to use an API when it's available. Same thing, for example, if there's a Google Calendar API or a Gmail API, we should be using that instead of going to Gmail you know clicking around ourselves. But specifically for reverse engineering the traffic, we actually sit ⁓ shut down skills quite recently. ⁓ just because agents have not only become good enough, but Yeah, there's there's so many cases where the the the reverse engineered API doesn't actually fulfill other requirements. I think going back to the bot detection question too, the more that you are reverse engineering website APIs by just watching the cur ⁓ the the fetches, you still need to probably behave somewhat like like a human for some of these websites. And we actually have like ⁓ like there's an API we have like browser curls, so you can make API requests via the browser. And most of the the reason why Our customers have found that valuable is because the biggest question right now is like authentication, which is tied in with again anti bot and sort of like identity. So I think that's I think that's sort of rooted in that question, which is like this reverse engineering, is it is it being done in a way that is like sort of like if you're not authenticated, is that going to be stable for you? And then but if you are authenticated, usually the authentication authentication still is happening through the browser. God knows that there are quite a few MCP servers that have me reauthenticating once a day. And so the authentication that that layer still happens through the browser. I think the browser can still be quite useful for for some time. Yeah, thoughts? No. I think that was that was good. How much of your work is like computer use in general versus specifically browser use? It it's kind of hard to break it down because they're so tied together these days. I would say if anything, it maybe leans a little more browser use than computer use. So yeah. Yeah. What are the unique challenges of just computer use in general compared to specifically browser use? I think the main challenge is like for computer use, you need to like have a computer. Meaning like if it is the user's computer, then like no no, but it's it sounds dumb, but I'm I mean like if it is the user's computer, then your agent is acting on their computer versus if you're doing browser use, you can do something like spin up a ⁓ browser and then do stuff on that browser. Granted, you could also like spin up a VM and do stuff on a VM as a computer, but typically when the user wants to do some computer used task, it's for stuff that's on their computer. And so you need very hacky ways to either be able to do that in the background or the user just needs to like sit there and watch their computer be used by a model. Versus for a lot of web stuff, you are able to just like background it either through C D P or by using a provider like kernel and just spinning up a browser. So I would say that's maybe one difference that makes it challenging. And then the other one is like the DOM piece of it, which is yeah, with browser use, like you get to actually do hybrid like DOM as well, versus computer tends to lean a lot more on just like screenshots and pixels. Yeah. So back to the so back to the evals a little bit. So there was obviously like a huge increase in the performance on OS world. Yeah. Was it just more data? Was that basically it? ⁓ I would say like It was ⁓ more compute, more data, ⁓ you know, scaling laws. Yeah. Yeah. I think we're big believers in scaling laws and so that tends to lead to better performance. Okay. So but also harness improvements as well. Yeah. Which I I I would agree. Yeah. I mean harness improvements have been a huge thing. And the harnesses of course are optimized to do specific functions, have like an optimal number of abstractions for the agents. And yeah, d having those two options for one another is is super good. And we also see a huge dump in performance ⁓ in, for example, our new API, which is built off of open code, a a very popular harness. So yeah. So what makes a good harness for browser use? One thing for for any general LLM, you want to give it as little restriction as possible, as little rule following as possible. So what I by like optimizing a harness, for example, is you you want as few tools as possible and you want tools that are very familiar to the agent. So let's say ⁓ the way that we have our open code fork is that the browser tool is actually just a mirror of the bash tool. So that means that the LLM already knows the shape of bash tool super, super well because it's super well trained. And so using the the browser tool is is really, really easy. And aside of that, yeah, just giving it as much information on CDP as possible, which mostly goes, you know, into pre-training. So yeah. ⁓ Yeah, the the I would say that in terms of optimizing for the har ⁓ the the agent and the harness, that that gap is quite small between competitors. Yeah. So is there any need for like Selenium or Playwright anymore? I would make the argument that yes. ⁓ of there's always going to be cases where these things, you know, make sense. A lot of what ⁓ what a lot of people do is, you know, use both hand in hand. So they'll make a browser use script and then we can cache it and then rerun as You know, some C ⁓ like a list of CDP commands or playwright commands. Yes, there's there's still a case you know, playwright and Selenium. Mostly playwright because LMs are so well trained on playwright. Whatever LMs are good at, guys. Yeah. Yeah, we still see a lot of playwright out there. So it's a question of like migration and how people like how quickly things move. And right now, playwright still works pretty damn well. Have you seen any? I think maybe you were trying to get at this earlier, but like what are the types of use cases that people are spinning up like a thousand parallel kernel VMs for. Yeah, in parallel it tends to be information gathering. They're trying to ⁓ figure out like yeah, I think probably the easiest one to talk about is like shopping basically, where there's a lot of products to pull in and they're trying to build an their own index of what exactly some vendor offers. Usually if you're looking at thousands in parallel in like for when they're representing like humans, right? A user, ⁓ they'll be, you know, like spinning up an agent for, you know, doing some kind of automation and Typically I see like healthcare and insurance like you guys do is ⁓ the same thing. Yeah. For us, if you have a thousand browsers, dude, what do you what could you even do with a thousand? And not to ⁓ I don't know if I can go too much into it. I would say each agent is typically to get give you guys a better idea, each agent is typically dedicated to a single workflow. So let's say that you like QA is a big one. So let's say like you have a list of clients, like one business has a list of clients, and those clients are like, I want you to make sure that our website's always functioning properly. So if a thousand clients, then you can have thousand browsers and a thousand browser agents all running on each of those each of those clients. Yeah. For each of those clients. Interesting. I want to shift the conversation a little bit again. So like long long run agents, like what unique challenges does that pose when you want to run an agent for like an hour and a half versus the way that they were working before? ⁓ yeah. ⁓ we we were we we've thought a lot about this. Actually, random side note, we released something called like game mode today. And ⁓ we games take a lot of effort to produce for LMs and we burn through so much money today ⁓ for ⁓ anyway. For I I will say that for long running agents, there's optimizing for it is is of course largely part of the model, which you guys have done an excellent job at. And making sure that the the agent can use its context window quite well. And that's yeah, it's really hard to optimize for long running horizon agents, like horizon long horizon tasks outside of the the model. Yeah. How does like What different parts of your harness do you have to update to make them better at long horizon tasks? Actually, nothing. I as I noted earlier, you know, the less restriction that you provide to your LM, the less rule following that you give to your to your agent, the better. And so we noticed that, yeah, if you if you actually give the agent agent less tools, less instructions, it actually becomes much better at those long tasks. Yeah. What about like how do you have to think about l memory when you're thinking about these long horizon tasks? Like similar to like the the cache script idea of breaking the cache s every N turns. Things that are heavy for the agent, for example, the DOM layer and the the the the screenshots can be, you know, removed from context, for example. That's one optimization optimization idea. And really good compaction is a big thing for super long tasks that run for hours. And yeah, those those Is that something that you guys can control or is that just up to the model? ⁓ way that it compacts. The way that it compacts is, for example, for for us it's open codes compaction, which is really, really, really good. And you could of course add additional instructions on top of that. But for long running tasks, ⁓ yeah, we've we've we've had a lot of success with it. What makes compaction good? I mean, compaction y you can imagine sorry, so what what makes better compaction versus like worse compaction? I think Lucas, you probably have a better understanding of this actually. ⁓ unfortunately I don't work on long context or hardness, so I I don't, but let's theorize. Let's theorize today. The c why compaction is what makes good compaction. Wow, guys, I don't know. I would theorize that the compaction is of course like just setting the context of another model and getting getting a smaller a s smaller document and I would imagine the ability to pick out, of course, important pieces of memory that are useful. For example, like let's say you put in like an API key somewhere or a password somewhere, the model should definitely remember those things and transfer them over the compaction. And also remembering most recent things that happened, because of course people are running humongo like have humongous conversations. But what happened at the very beginning is often not all as important. So but also of course at the same time, you need to be able to judge everything fairly, where LLMs have had a problem of, you know. Having a bias towards information at the very beginning and the very end. So having a no, ⁓ I was actually just curious with this. ⁓ yeah. Yeah. I was just curious with like compaction when you're dealing with like sort of like computer use screenshots, is it able to compact things very well? I mean for well, screenshot well, screenshots specifically we don't like compact the but yeah, com compaction works. Just fine, thankfully. Thank goodness. Yeah. It works just fine. Yeah. It works. It works. It works. Yeah. So it is, it is it is able to compact things quite well. Cool. Like one of one of the examples that the agents are not so good at right now is specifically on OS World two point ⁓ where you're working on this long horizon task. It's doing pretty well. Like like you were saying that the partial credit is is ⁓ like you're pretty good at partial credit. They're like, okay, like fifty percent. ⁓ but like one of the things that they get tripped up on a lot is let's say you are like an insurance claims adjuster or something and you are maybe like I don't know, like you're updating an invoice. Now you get an email from your boss saying like, Hey, sorry, I gave you the wrong insurance claim. Can you update it with this one? The agents are not really good at hey, like my entire state just changed. I need to update that. Do you guys have thoughts on like how agents should get better at that? I think probably more data would be good. Yeah. More data is that's the always the answer. How about some more compute? Yeah, let's not forget that. Potentially. ⁓ but no, in in all seriousness, I think, yeah, it it's a mix of like, yeah, identifying kind of where these models go wrong and then also to back to the harness, like there there's kind of like the right balance of harness. So like Reagan mentioned, like if you have too much harness also, you could end up tripping up the model too, because maybe you're like writing all these memories and what's your goal and et cetera. And the model doesn't realize it needs to kind of change its terminal goal mid conversation. And so I think a lot of it also comes down to like where the model's being used, like the loop that the model's on to. Those are all the main questions that I had. I think I would like to open it up to the audience for any questions that you guys have. So when it comes to computer use, like which operating systems and browsers are like being focused on, like between Mac OS and Windows and Android and iOS, like what what's the what's being focused on most like percentage wise? I would say it's all of the above. I think developers tend to use Mac and Linux and so those are sometimes treated as first class citizens, but I would say all of the above. So iOS and Android is all also like being worked on, okay? Yeah, if you if you use like a computer use agent and you have like an iPhone like simulator on your screen, the agent can certainly manipulate it. ⁓ I don't know if you're curious on the browser side. For the browser side things, ⁓ that's a little more complicated. the for browsers, the all it's all like fingerprinting and How your browser looks. And so for example, there's actually very few Linux users in the world in relative to the Windows and the Mac users. And so if you overload a website with, you know, a thousand Linux browsers, that website's gonna block you immediately. ⁓ so what we what we do is yeah, we use a mixture of like yeah, Mac ⁓ we use like Mac fingerprints and yeah, so so we we get detected less often. Any other questions? Yeah. Back to the ⁓ context window with like the DOM. I've I've actually seen like the opposite where navigating websites and using the DOM has been pretty inefficient. For example, utilizing Salesforce and going through maybe five to six screens, you've pretty much filled up your contacts window very easily. I was curious if you guys are seeing any sort of adoptions of things like Web MCP to make that more efficient, or do you have any commentary around that? Web MCP is a I'm surprised it hasn't been adopted more. I guess it's like a very new new concept. About the DOM question, it really depends of course on the website. Salesforce is one of those humongous websites that's ⁓ really difficult for any computer use browser agent to use. See we Salesforce actually Salesforce actually wrote paper on it w using us like about how ⁓ browser agents use Salesforce. Yeah, I I'll say that's one of those kind of sort of edge cases where yeah, Salesforce is just one of those websites that's like so tricky. Same thing with like yeah, Microsoft Azure or like any of those annoying websites you might have used too. I I'm I'm curious about how you might deal with a Rexus because the Rexus is designed for humans and how Rexus might actually evolve over time as agents proliferate. A A Rexis? A recommendation system. ⁓ a recommendation system. Yes. And and what was the exact question of how a recommendation system i yeah. So how how do you think about interacting with a Rexis? Because you know, basically a Rexis is supposed to be highly personalized. You go on Amazon, you buy a blue shirt, they they send you all all kinds of blue shirts. Yeah. You know, you know that that sort of thing or you you like Mexican food, you know, and you know, and so forth. So it's it's highly tailored, it's highly personalized. And I'm just wondering what the trajectory is for when it's an agent versus a human. I'll say that for browser agents specifically, we of course like we we proxy for humans. So for example, as a the case I mentioned earlier of the the agent using my Amazon, the agents clicking on my Amazon and so it would find what I would find too. If that so agents actually work with recommendation systems quite well, I would say, for that specific case of like working on behalf of users. Because it's it's it's taking on your persona. Yeah, yeah. It's certainly quite difficult to ⁓ do systems for without that. I I think like where the Rexis system lives just moved in the stack. Like now the model knows you very well. And so when it is going to search on Amazon, it might not need Amazon's Rexis, which previously was like your eyeballs on it. Now it's the model that knows like Victor prefers long sleeve shirts over short sleeve shirts. We've spoken about it. And so when it's searching for shirts on Amazon, it'll, you know, lean towards long sleeve shirts instead. So I think the Rexis where the recommendation system lives maybe moved up the stack. And or maybe you can imagine a world where when websites provide front doors to agents to interact with them, there will be a way where like maybe Amazon could be like, hey Reagan's agent. Here's like a list of like things he searched in the past or like what we believe he likes in case if you don't know about these things or something like that. So I think it'll just shift in the stack a little. Yeah. Because your LLM may know you better than the Rexos does. Correct, potentially. Or or maybe not. And yeah, or maybe they're both interfacing. I d I don't know. But yeah, I think it just shifts in the stack a bit. Right. So you just mentioned about ⁓ the websites and if agents are going to take over more of navigating websites, then if you're building new sites or new versions of the sites. Would you really build it so that you treat bots as first class users as well? And is that a trend you see and how do you see this evolving? Yeah, ⁓ I think I certainly would. And I think like I use this analogy that's like if I were to ask you like, does PizzaHu have a website? Like you don't need to check, right? Like you know for a fact that Pizza has a website. ⁓ If I would have asked you that maybe in like nineteen ninety-seven or something, you would have maybe had to be like, well, let me check. Like there's pretty new tech. I don't know. They're a pizza shop. I could just go and order, call and order, right? I think with front doors for agents on websites, it'll be very similar to that. Where it'll be like, today, if I want to order a pizza using an agent, I kind of need to check. I don't know off the top of my head if Pizza Hut allows like agent orders. But I think within a year or two years, it'll become as obvious of a question as does pizza have a website? It's like, yeah, of course you can order a pizza like with your agent, like, because every platform is gonna have an agent front door. That's my prediction. I could be wrong. I think and I think there to some degree the world will bifurcate into like anti agent and pro agent sites and capitalism will do its thing and we'll see which of those will hold. Maybe being anti agent is right and you can maintain more of that user value and those firms will win out. Or maybe being pro agent if everybody is making their purchases through agents will be advantageous and those firms will win out. But I think it'll play out kind of similarly to other technological shifts that change the quote unquote front door of a business. Going back to this question of we want the agent to act like a human, how do you though balance that with wanting it to have the advantages of being an agent as well? So I think like you you want it to behave like an agent in in ways that like humans so like When you go to a website, like one of the things, because I work with customers quite a bit, but you'll find out that a lot of websites out there are broken. Like a lot. And ⁓ inevitably the ones that our our customers are looking to automate are most often broken. And so when they are you want the agent to be able to interact with the website like it normally would, but you also need to be able to t pipe in other types of telemetry to that agent. Like you want that not only the agent to have access to the DOM and to be able to peek into the network logs and to see what's going on in there. You also want the agent to be able to see the JavaScript errors or the you know CSR errors or whatever type of errors are originating from that browser. So you wanna like you can have the screenshots, you can have the you know C UA level stuff, but ⁓ you also need to be able to give the agent the ability to look into the programmatic interface, DOM, and also the logs of the of the applic the web app, because that typically that is not always gonna be working. Yeah. ⁓ to ask a clarifying question. So what do you mean by like all the things like an An agent that you want from an agent, for example, like maintaining those capabilities. What what you mean by that? Yeah. So going back to some of the examples that you just went through of say you you have a thousand VMs spin up and you want an agent to collect data on sales or on what a company is selling. I think you'd want an agent to do that really efficiently and say, I want to know about every kind of like backpack that's sold in the world. I would love if an agent could do that. in a day instead of like a month, even if it had a thousand VMs spun up and and they were all acting sort of on a human scale of timing click like all the graphics to use. But but yeah, that that's basically it I mean to me there is a trade-off between this is identifiably human behavior that would be recognized as such and this is an agent acting as an ultra efficient human or maybe as you know, that much more humanly efficient than possible. I'm also originally a bit of a StarCraft fan, and so I'm thinking of this in terms of in slightly in terms of APM or something like that. So essentially you want agents, if I understand your question correctly, like agents to be like super superhuman. Or yeah, I feel like there's a tension there between wanting them to act like a human, but also wanting them to have yeah. So for browser agents specifically, you we as I mentioned earlier, both agents and humans tap into the Chrome Devs Tools protocol. And so you can imagine actually those are kind of interacting with the same thing, if that kind of answers your question. And so that like human tension is kind of dissipated because we both use the same thing. So it's kind of everybody said that not to exhaust context video to the end, otherwise it's will start gluc hallucinating. And rather to start new chat. New branch or whatever. What is the trade between losing half of this context and starting new branch and or keep going because of like few steps till the end? And I always ask agent. I have a few more questions should be keep going here or should I start new one? And agent always say no, keep going here. ⁓ ⁓ yeah, I it's just it's really hard to give you a blanket rule because it's very like situationally dependent. What I'll say is like as you reach the end of the context window, there's just a lot more things in the context that the agent has to attend to. Right. And so it's very similar to like I I don't know, if you're reading a very, very long book versus reading a short story, you might be able to keep all the details of that short story in your working memory while you're reading it. Versus if you're w reading a very long book. You know, you might be writing notes or you might have bookmarks the earlier in the book when certain things happened, et cetera. And so it really depends on if the task, if it's, you know, the task has such rich and specific history that, you know, if you were to compact it, that rich history would be lost. Or if it's the kind of task where, you know, the model maybe made like 20 tool calls looking to find something. Nineteen of those output huge outputs that were irrelevant to the prompt. In that case, the model is still at like You know, it's trying to not attend to all that context, but it's filling up its context. And so in a case like that, you might be like, better to wipe this or to compact it. You know, it'll keep the original instruction and it'll free up attention for the model to focus on the important thing. Normally there is no extraction, just full just flo Yeah. So again, yeah, so it really depends on the exact task. But yeah, I I would lean on letting the harness compact as it goes because harnesses typically have good heuristics for that. But I think one of the cool things about all of the coding harnesses and like being able to run longer sessions and agents building whole things and whatnot is you can build a lot of like steering infrastructure and guardrails, like you can diff the code, you can read commits, everything is segregated by files. You can like ⁓ write tests, run tests, look at all the logs and like ingest that pretty pretty well and it's all very like discreet. How do you think about that for computer use and like temporally comparing browser state, for example, and how that differs from how you would like diff code. Yeah, I I think that's a great question. And I think like I tell people a lot that I think there's a lot of headroom in harnesses, particularly for computer use. Like maybe even more than in code right now. And so I think a lot of that should be handled at the harness level as opposed to at the model level, especially for a lot of the temporal encoding. Otherwise you'd be forced to do things like, you know, almost like in just like the video. And make it like, yeah, one frame per second or something. But then you know, if the pop-up is really quick or it's a sliding banner the model might still miss it. And so unfortunately, I don't have a great answer to your question. I think your question is effectively one of the big challenges like going forward. And it'll probably be handled by a mix of the mom is to get okay. ⁓ so ⁓ even as of today, like computer use especially has been like particularly slow and ⁓ browser use is like more or less a little bit quicker, but compared to like a human, it's still fairly slow. ⁓ right. So I'm sure in a couple of years it'll be there. Maybe across both model and the harness, if you could speak to a little bit of like over the past like few months, what kind of optimizations have sort of like led to like quicker browser use or like computer use? Yeah. I will say that once more this is very, very dependent on the model. So we've seen in super small models like Minimax ⁓ and like ⁓ three point two point seven, three point seven ⁓ working super, super well. And and really high, like I said, really, really quickly. And you have to put into context that you as a human, it looks really slow to you, but for a browser agent that has doesn't have all the human memories you do, it's more like imagine if you bunch a website the first time and you're trying to discover everything that's on that website. It's it's actually relatively quick in in in terms of that. And so yeah, but once more, I think it's super dependent on the the model. Yeah, I think another benefit is just like being able to parallelize, like that's effectively a way to increase your speed without needing to actually increase your speed. And I think with computer use and browser use, there's definitely a limit to like if you're going too fast, you're almost breaking the software sometimes because it's not designed to be used in such a fast way. And so, especially for certain websites where like the model like clicks something, immediately gets a screenshot back and it's like a blank loading screen and the model may get confused. So there's also edge cases. Actually, yeah, I hit ⁓ this problem with one of our customers in order to try and make things happen real quick, ⁓ they proceeded to open up like dozens of tabs. on our browser and then to do everything via you know via CDP on those tabs. And that was great. ⁓ our browsers are, you know, we provision them with eight gigabytes of memory. It seems like a normal amount. But the problem is is that Chrome, the beast that it is, caches aggressively and it actually overwrote all the hard disk space we allocated to this VM. So we had to resize it and we had to figure that out. But I was like, that was like a one of those things where we learned from our customers realizing, ⁓ that's the genius way because You wanna make sure that like for a user, you wanna make sure that the profile and the cookies are all updated ⁓ correctly. When so when when they finish the session it all saves correctly. So the best way to paralyze is Chrome built it in, tabs. And it still kind of looks human because like I don't know how many tabs do you have open today? I probably ended the day with like 30-ish. Yeah, so it seems kind of okay. No one's blocked me yet. Thank you for sharing. So I have noticed like a interesting thing that when I try to help computer use agents by clicking or typing when it is working. it actually ended up slowing down because it needs to pause and replense. So I wonder do you think that is the that is due to like model limitation or it is because infrastructure synchronization issue? And also more broadly, do you think the br ⁓ computer use agent in the future is more fully autonomy or ⁓ more like a collaboration with human and computer use together? And yeah. And then if Like whichever the future case you supported, do you think is there any ref research or infrastructure breakthrough is missing for any of the future computer use agents you in you imagined? There's three questions. Yeah. Yeah, there's three questions. So your first question was when a computer use agent is controlling your computer, you try to type in and then the agent gets confused? Yeah, you you actually will slow it down. So it needs to For sure, do yeah, don't don't don't do that. Yeah, don't I would recommend strongly against doing that. The agent is not used to that and so for it it's almost like a phantom like thing that's happening. Yeah. Potentially maybe even thinking it's a prompt injection, so it might like block the flow altogether. So you ⁓ you don't sub you don't recommend like a human and browser to work together. At least not now. It's a cool idea though, for sure. And yeah, maybe for some like educational sort of use case or something like that, it it's really cool. But I think the way the harnesses are designed today, especially, like it's not meant to be used hand in hand. Whether or not that'll be the case in the future is speculation. I I don't think so. I think we'll probably get to a world where like it's not even computer use agents, just agents generally that you are managing a fleet of. Computer use is one of the many tools that they have at their disposal, but You're not really working with the Yeah, I can ⁓ put more context why I try to click or type because I think some tasks it's very obvious for humans to do that. For example, like click some buttons, etcetera. But that takes long time for the agent to find the actual button. So that's why I try to helping it. And then by actually like making it like slower. So then I imagine the second question you imagine like the future of browser agent is fully autonomy instead of human and browser working together. Yes, certainly I'd be human. So yeah, the what what what I see with our customers, it's mostly like, okay, the the browser ⁓ browser agents are always like a little bit s a very well-versed human can do things very quickly. ⁓ but it's mostly about like fanning out. It's like, can you fan this task out across mobile tabs or multiple browsers? And most of the time human intervention comes in. ⁓ which I deal lots with because they come screaming at me, It's broken, it doesn't work. And then I like look through the logs and be like, It worked ninety nine percent of the time coming at me because you want it to work a hundred percent of the time and it it's not quite there. But that's where I see like ⁓ sort of like how humans intervene and making sure that when you have these agents running, that you have good telemetry, good replays, like you have the recordings to be able to review and then intervene in a pro more programmatic way versus actually trying to compete ⁓ with the ⁓ with the agent. Yeah. And into going back your first question, of course, as Luca said, you know, doing things ⁓ like typing in a box or whatever while the browser agent's working is of course gonna de degrade its performance a little bit. And yeah, we're we're moving towards a more autonomous world where once more these things tap into the same protocol, they have the same capabilities. So the main thing, the main like unanswered question is okay, when should I be doing things and not? And for example, Anthropic has ⁓ done a lot of work on this, you know, auto mode, dangerous skip permissions mode, you know, things like this. And so permissions is I think the the final blocker there. But in terms of the capability, I think everything's set up well, especially for browser agents. Thank you. My question is around how much raw capability do you think the models have in terms of being able to, let's say, do tasks on a computer and a browser? As an example, could we use browser use to make a full nuanced mock up on Figma? Or let's say when you have slides, right? And you have text boxes that overlap. I think all companies and all hardnesses currently handle that through, you know, code. You do deterministic checks and you fix it through code. But I'm curious, does model has the right have the right capabilities? What's Figma? I only use cloud design. I I would say in terms of raw capability, I I'm not sure like I know exactly what you mean by that, but I would say, you know, the raw capability of any alum is like a hundred X a human, if I was gonna quantify in some way. For example, could do you think a model could essentially drag A box and make it, let's say, seven pixels larger, right? Using the drag tool which is available on Figma rather than doing that through code. ⁓ so in terms of like a browser cable. Yeah. Yeah, I think it it like it might be possible today, but it's probably not the preferred interface for the model to do said thing. And so to give you another analogy that is humanoid robot-based, imagine you buy a humanoid robot and you want to take a ride somewhere. You can maybe ask your robot, hey, get in the car in the driver's seat and drive me to that place. Or you can call a Waymo. Right. And like the Waymo will probably get you there safer, more efficiently, et cetera, et cetera. But in theory, a humanoid robot could maybe drive a car because it's shaped like a human has eyes, et cetera. And so there are certain tasks like that where like I I don't even know if the question matters because we already know of a quicker, better way to do it. And it's more so the gaps where, you know, computer use is the only way to still do things where we want that raw intelligence to really like play a role. Does that make sense? Yeah. Yeah. I would venture to say probably the agents would rather watch the network requests and then reverse engineer Figma's C D RT protocol and just send the C D R T events through down the pipe. ⁓ because that would probably ⁓ it might actually just be faster. I I feel like that's always the like yeah. And also like it's a question of like when you say ⁓ which is b w what are what is possible, do you mean like today or do you mean like I guess like in this indeterminate future? Because like usually like most of our Most I think that's that's the premise of the browser use question today is like, is it eventually obsolete? And like I guess that depends. Are we as humans just obsolete meat bags? ⁓ in which case we're just like talking amongst ourselves. It's like, yeah, it could work, but like there are probably better ways and like let's just get work done as quickly as possible today. One agent could directly write the binary itself. Yeah, exactly. There's always that. Yeah. Yeah. Yeah, just curious, how much do you leverage imitation learning maybe for specific tasks? Vers or is everything just reinforcement learned? Yeah, unfortunately I I can't comment on that. Maybe maybe Reagan can. How are you guys training yours? I I can't comment on it either. I will say, like in terms of how we're used, we use more f mostly for RL. You guys are post training Quen, right? ⁓ we we already did. Yeah, ⁓ yeah. ⁓ we fine tune twin fine tune quen, yeah. ⁓ it's open points model and how many face. My question is like what's the biggest bottleneck or problem that's being worked on right now? Like I know you mentioned like better models yield like significantly better results. Like what's the biggest bottleneck or problem right now in the in this? Honestly, it's it's about the stealth and the capture solving. That's like the the last, not last frontier, but it's the there's so much work to be done there. More than there is to be done on the browser agent side, I would say. I have one more question. Yeah. What do you guys see like the final state of this? Like to me, it's kind of like one unified app where like just the text input and then you just, you know Say what you want, like buy this and like what do you guys see as the end state of computer use agents? Yeah. We ⁓ Browse use have this model say like our end goal is to you tell a computer to do something and it does it. So no matter what that command is, it will always it's like a slash goal for everything that you do. ⁓ and yeah, that I think that's kind of the end state. You you tell you tell you give a really high level instruction, ⁓ make me a million dollars and somehow it will try every single way possible to make you a million dollars. ⁓ have you heard of T T E O Thoughts to Economic Output? So basically you can think what you want to do and if it is something valuable, then the model will know and just do it. So I'm I'm just messing with the But no, I I think that the end state of this is like still very much in flux and like we like models started with this like chat interface really, where even before that it was like playground mode with pre trained models and then that's gone to like chat interfaces. Then that has moved on to like T UIs and stuff like cloud code where the models like in an active loop and then you know That has moved on to like desktop apps like many that exist that are similar to something like a Cloud Code, but that also handle your files for you. And then now I would argue that's moving towards something like Cloud Tag, which is like, you know, you're a collaborator in SAC that has its own tools. You don't even need to be like in an active session with it. It could just run and do stuff. And so I think we're gonna keep seeing this like up-leveling of like what you are doing until you are more and more like a manager than anything else. And it just comes down to like, You know, what is the best interface to manage all these like super intelligent like things that you have doing work for you? And it's probably not a chat bo like a chat box, I think. I don't know. You definitely see a lot that like the voice stuff is becoming popular. And outside of the joke, there are people trying to do like thought to like or like whispering in your mouth to like text as well, which then becomes actions for the model. And so yeah, I think we're gonna keep getting abstracted away. Yeah. I'd like to ask you a question. ⁓ at least my bet for this interface, this is like fun question to think about, is I think that, you know, glasses is gonna are gonna be a big one. And like I think glasses and like maybe like a neural link type of thing. I think those are gonna be like the two main ones. Do you have do you think what do you think about that for interfaces? We're all gonna be like clawed bots. Yeah. I think yeah, I mean th I mean I think we've all had this experience before, or maybe just me, where like you're working with an LM and you're asking it like, you know, you might be brainstorming with the and It tells you something you're like, whoa, that's a really good idea. Like, what's the next step? And it's like tells you next step and you're like, okay, I'm gonna go do that. And then later I think, ⁓ damn, that was human use. Like that was like reverse. Like it told me what to do and I just did it. So yes, I like that idea and it aligns well with ⁓ T T E yo. Yeah. Basically like we are we become a tool use. We become a tool. Yeah. Is that AGI? Like would you say that? I yeah, I I don't know. Yeah. Hard to define. Yes. Yes, humans as tools is interesting too. I still like what you said earlier about like, you know, ⁓ you eventually someday you'll order your model to get you a like a car and it'll drive you somewhere. But like it does still feel like the car will be another model driving me. Opus is not gonna drive me to my destination. It's not gonna we're not gonna hand it car controls, car use, and then it's gonna start driving me. That would be ⁓ I don't know if you guys are training for that, but it seems a little terrifying. That's a delegated model. Yeah. All right. I have one final question. What predictions do you guys have about computer use agents over the next twelve months? So I think I would say the same one that I ⁓ said when I answered this ⁓ gentleman's question, which was yeah, I I suspect that they're that like in the next six to twelve months, most like companies, websites, et cetera, will self bifurcate into very anti agent and putting up even more blocks or more pro agent and putting up native ways to interact with agents, up to the point where, you know, a year and beyond it'll be obvious that Most places either support agents or no places support agents. Yeah. I mean for the w in terms of the web, I 100% agree. You know, it's definitely going to there's definitely gonna be that split. In terms of the capabilities of the agents themselves, I would say that for for long horizon tasks as LMs get better, I think that's then a a huge place for growth as, you know, agents become better at managing their context, compacting ⁓ these long horizon tasks like, ⁓ fetch me. two hundred and fifty Wikipedia articles, generate a research paper or work on this math problem for like, you know, two days in a row. ⁓ these problems become much, much more feasible as the models get better. Yeah. All right. Awesome. Well that's all we had for today. Let's give it up for our panelists. Thank you guys.