Nik: Welcome to AI and Design, where we explore how artificial intelligence is reshaping the world of design. I'm Nick Martillero. Dan: And I'm Dan Saffer, and we're faculty at Carnegie Mellon's Human Computer Interaction Institute, the HCII. Each week, we break down the latest AI developments, dive deep into topics that matter to designers, and talk with fascinating guests who are right at the intersection of these fields. Nik: Whether you're a designer working on AI or an AI practitioner interested in design, we're glad you're here. Dan: This week we've got some news from and about OpenAI and no, it's not Apple's lawsuit against them. Nik: And we also welcome some guests, Jason Hong and Shaw Vic Das, faculty at the HCI and co-founders of the startup Fugu UX. Dan: But first, let's start with OpenAI. So OpenAI launched their latest model, GPT 5.6 SAL, last week. Now, we don't usually cover model launches here, but this one had two interesting things about it. And the first was that Even though it's a new model, it jumped to the top of the Design Arena leaderboards that track how good models are at design. Now, if you've never heard of Design Arena, and I frankly hadn't, according to their website, this is what they do. So Design Arena is the world's first crowdsource benchmark. For AI generated design. We give the same creative prompt to the top AI models, show you the results side by side, and let you vote on which one is best. I had never heard of them before. Nick, had you? Nik: Yeah, I've actually heard of Design Arena and I thought it was a cool idea, Being able to benchmark how well these models are doing a design, I think is something that is missing from most of the standard benchmarks that the companies use to evaluate when they're developing their models. Typically they're doing problem-solving things. A lot of them are doing code writing, software development challenges. I imagine now Design Arena is actually think probably getting big enough, and actually With this pretty big article showing that GPT 5.6 soul is jumping to the top, I would imagine that now the companies are starting to leverage these as benchmarks and they're working harder to try to get visual design and interaction design actually part of the training regimens for the models. Dan: Yeah, I'm wondering what design are they really tracking here? Is it is it really just visual? is it also a lot of interaction design and UX design? it's hard to tell. Nik: Yeah, I think most of it is pretty visual because that's what you see represented in a lot of the structure of say the code when you're looking statically at it, an interaction emerges from the user actually well interacting with something. And of course the functions are there to define that, but I don't know if it's as easy for an AI model unless it's actually working in an interactive mode, using a visual language model to actually click through the code and test things. It's hard for it to actually understand that. in regards to what I feel when I've seen Design Arena, a lot of it does feel visual design and even the things that the article here that we link to in the show notes is talking about is that there are a lot of AI anti-patterns that GPT 5.6 soul seems to suppress. Dan, you'll be happy to see that it does not seem to put in purple. Dan: Yes, I did see that. Amazing. five point five did a lot of purple gradients and apparently that's gone now. Which I'm like, yes, thank God. Nik: Yeah, one of the other things that's really interesting here about how five point six seems to be working is it seems to be leveraging good templates and then personalizing those. Basically customizing a known template, a known design pattern. And that is I think sensible. I mean, that's how a lot of us work, from what we're doing. I mean, either we have a design system within the product we're working on or we're going to follow common design patterns, within the say sector that we're in, right? If I'm creating an e-commerce site, I'm usually not trying to create completely different, checkout experiences from all of my other competitors. I actually want that to be as seamless as possible so that people actually buy the stuff on my e-commerce site. Dan: Right, you don't you're not reinventing the wheel every single time, but there are times where you want customized bits. That's just part of the job. And so it seems like from what they're saying with their evaluation is that it is combining these templated bits with customized outputs. it's putting things together in interesting ways. And that's pretty cool so that we won't be stuck entirely with the designs of twenty twenty four for the rest of our lives. Nik: Design Arena has some really cool speculation that that OpenAI had worked pretty hard to improve GPT 5.5 with much better design abilities that we now see in GPT 5.6 soul. Some of the stuff that they show on the website, and people should go and check it out, is they have these really cool visualizations of basically the design space, it's images of say websites and things projected into an embedding space, and there are really clear holes that basically they're these just big blank spots. That GPT 5.6 soul just doesn't generate things in. And this is what they're speculating as these are where the AI anti-patterns are, where the rather these are where the AI patterns lie in the holes that now they've cut out. And basically, people at OpenAI had sat down and they've basically developed their systems to say, stop doing this. Dan: they're the purple gradient black hole. Anything there, toss it. Nik: Another thing that's actually really interesting to see here and is making I think this a potentially competitive option for designers out there to look into is the cost differences. GPT five point six costs less. than Fable in terms of how much it costs to say input tokens than the output tokens you get. and then in addition, it's also faster than a lot of other models. So they compared it to GLM 5.2, that's another model that's gotten a lot of hype in the last couple of days that's quite cost effective at generating code, but that basically OpenAI is a lot faster with this. And so I think that this is really showing that open AI is taking seriously at least front-end visual design and trying to improve upon what's being generated. And it leads us to another small story, which is the fact that Figma has now integrated a connection with OpenAI, not just with Claude. Dan: Yeah, I mean part two of this and I don't know, this may seem a small story, but maybe it's actually a small story hidden inside of a bigger story, which is that now you can use five point six in Figma Make, I mean five point six is open AI's most capable model. And with Figma Make Prior to this, their most capable model was Claude Opus 4.8. You can't use Fable in Figma make right now, and maybe never, because now Claude Design and Figma are in a little bit of a of a tussle where they're both trying to go after the same thing. So it'll be interesting to see if Fable ever makes it into five point six. So I think it's an interesting thing because we don't know that OpenAI is looking to make the open AI design. it maybe there's a new partnership here between five point six and Figma make that they're weaning off of Claude. So I think that could be a pretty interesting development, big if true, as people say these days. well as people said Fifteen years ago. Nik: Yeah, and I think this also speaks to the fact that just the market right now, in regards to what models we use, is just becoming way more diversified. I know that on the show we often talk a lot about anthropic models, partly because we're users of them. But honestly, I'm using open AI models. I'm using local models, some of the distilled models out of China, Quen. these are becoming more and more sorry, do we take? Dan: Well the big news this week was Kimmy, right? That we didn't even cover. We're not even talking about here. Nik: Right, which you can't even sign up for on certain sites, right? Because they're they've had so much demand. And a lot of that is being driven by cost. actually there was a really cool chart that I saw earlier today in Scott Galloway's newsletter. and it was showing the number of tokens that are being generated and consumed across US and Chinese models. And the US models have actually they had a rise and they're actually staying pretty stable. But the Chinese models have really just accelerated in how much people are using them. Now it could mean that just more people are getting access to them. Of course, the chart actually shows token uses through open router which is a piece of software basically helps to route you to all kinds of different models and then you can pick and choose what model you use. But it's interesting to see that. I mean it seems people are really starting to think much more about how much things are costing. And right now these distilled models do cost a lot less. And so it raises the question, with Figma now having open AI as a partner and having access, at what point do they also start opening it up to all kinds of models? That you might want to use. And are they trying to, develop their own models internally, which is actually what different companies are doing? Microsoft, for example, they developed for GitHub Copilot their own coding agent, which is faster cursor, developing their own coding agent, actually a derivative of Kimi, but then with some extra stuff which makes it ultra fast and I think maybe works better for their users' workflows. So, I'm curious to see, is Dan: Yeah. Nik: Are you just going to see that if you have a software which is a great front end for working, do you basically just need to have all of the AIs be able to talk to it? And the benefit of all of the AIs is none of them actually work that differently, right? They're all language models Dan: the meta discourse around this is what is the future for your open AI is for your anthropics? Is it that they become AT and T where intelligence is a commodity and they're basically competing on price for intelligence because all the intelligence is more less pretty good some a little bit better than others some a little bit worse than others but some practically free especially if you're talking about an on-device model that one could be entirely free versus are they going to be something like an Apple or a or a meta or something like that is a that is a brand that has products that you are using, are you gonna keep using clawed code, clawed design, open AI work? are you going to be using those or are you just going to be using things that are built on top of those? And I don't think we know the answer to that yet. everybody's talking about it right now, and so it's interesting to mention in light of this announcement by Figma and OpenAI. Nik: Yeah, what I imagine we'll see at least in the near term and that I that really could be beneficial to a lot of software is figuring out how to find the most cost effective Language model, visual language model based technology that you can embed into your product to give a much better user experience, but you're still providing the core product that your users know and love. Right. I think at the end of the day, switching out to a completely different tool, there's a lot of cost in doing that if you have lots of workflows. And so just making the workflow better. keeps you in and potentially gets more people invested in the software, but you have to do that a basically the most cost effective way you can instead of just burning tokens. Or you're saying, right, some someday I'm gonna be, using OpenAI's work as my OS for doing my work. But that basically requires them to recreate a ton of software and a ton of ways of doing things. I'm not sure where Things make sense. I think there's a place in the world for both completely AI native new workflows, and then there's a place for, bolting on. And I don't say that in a in a negative way. I mean literally in this really, this is an amazing addition that you can add to your product that just enhances what it does. That, you might have even had these ideas for a long time and you just were never able to do them just because the technology wasn't there yet. Dan: Right, and there's the old adage that every new technology is treated like a feature and bolted on, and then a couple years later we end up with the native versions of that technology. and we saw this happen most recently with mobile. we had all of all of our desktop stuff somehow got ported to mobile, but it took a couple of years before we started to get things like Uber that only made sense as a mobile product or Tinder is another great example made makes great sense as a mobile product and less so as a as a desktop product and we're not there yet we don't know what those AI native products are going to be yet because I don't think we've seen any of them I mean I guess arguably you could say things like Claude Code or Gemini workspace or Claude Co work or now OpenAI's work. Boy, there needs to be some differentiation between these names. that these are the closest thing that we have to something that is truly AI native and could not be done years ago. That just feel extremely different. Other than the actual AI models themselves, which are something completely different that we hadn't really seen, although we definitely have had things like chatbots for decades. but speaking of something completely new and novel, or maybe not so new and novel, our last open AI story, this is a an all open AI day, is that they announced a new piece of hardware that's already sold out that I thought was pretty cool. It is a piece of hardware that how would you even dis it is a piece of hardware that Nik: It's a macro keyboard. It's a it's a it's an auxiliary macro key keyboard specifically designed for working with codecs. Dan: it's called Codex Micro is the name Nik: Yeah. And it is a little, side module that has let's see how many buttons here. It's got one, two, three, four, five, six buttons. it's got a knob, and then it's got some one, two, three, four, five. And it's got six other buttons, but they're RGB LEDs lit that tell you what agent you're working on. And the whole idea behind this is that you can sit there and have a piece of hardware that's monitoring I guess up to six and probably more, but at least six agents that are working in codecs, and then your main commands that you might need to do. So accepting, rejecting, push to talk so you can talk to it, starting a new chat are just simple. key presses. Dan: And that the keys will light up and change color based on what the agent is doing. So one agent, if it's idle, or if it's thinking, or if you have an unread chat or it needs your approval on something, the buttons light up based on the state of that agent. So you can immediately jump into that particular agent without doing all this tab flipping that you might do Right now. Nik: Yeah, so this is a cool piece of hardware. it is two hundred and thirty dollars. which for an external, macro pad feels a little bit steep. I mean it is it is cool looking, I will say. I think that Dan: Yeah. It it it's super cool looking. So it was designed by Design Studio Worklauder with OpenAI. I would buy it and I don't even use this. I would well actually I wouldn't buy this. But I mean i if I had a similar thing, yeah, I would I would definitely be interested in it for Nik: it. Dan: cowork, which is what I use, my version of this, I would definitely do that. Nik: Would I buy it for two hundred and thirty dollars? Probably not, because you can get awesome amazing keyboards. And for those who haven't gone down the rabbit hole to mechanical keyboards, you can spend a lot of money, but also you can get some awesome keyboards for like, a hundred dollars. Now I imagine that this was clearly a small run product. it's already sold out. So by the time you look at it, sorry everyone, it's sold out. but the thing that I love about this is that it really harkens back to the older Unix Keyboards. I don't know if everyone here knows this, but there used to be, when you bought a terminal interface for an old Unix system, there were all these other keys around it that were for specific Unix-based commands. And so you could actually have all these macro keys to do stuff for you. And I really like the idea of like, cool, I can have that again and add that, to my laptop, my current Word computer, which only just has QRD and then command function keys and arrows and stuff. But People have had these kinds of things before, so it's cool to see, what's old is new again. Dan: I was thinking back to the Doug Engelbart corded keyboard way back in the sixties, stuff like this. But I love this kind of intelligent hardware. I think it's a really interesting and fun space that I wish we saw a lot more of to make more of our digital lives more physical and more ambient than just another screen I think having these kinds of tools and accessories that are hardware is just a really interesting space and I hope we see a lot more of it. Nik: Yeah, I'll say, for folks who look this up and you think I don't know, that seems goofy, I'll say a l one of the things that is really awesome to me on here is the push to talk button and then the hotkeys for say accept, reject, branch, do something else. I've been utilizing a speech to text program and I just have a hotkey on my keyboard. And so when I'm working with a code agent now, actually I mostly press and hold talk and then let go. And actually it's having to go in and click the user interface for these accept or reject or do all this stuff that is a little annoying. So I can actually imagine exactly you're saying Dan, the Doug Engel Bart, I can have both my hands down on the keyboards. I don't have to use my mouse or my mouse pad and I can sit very comfortably and mostly talk to my computer and then have just a couple quick action buttons. And I'll say I don't know, I've been having some funny wrist problems. I hope I'm not developing RSI. but it's actually been super nice to be able to do this because I could let my hands rest. So for as goofy as some might think about this, I actually think there's a really nice physical, almost accessibility use case here. and then also just a general just being able to be, focused in on the tasks at hand, as opposed to having to cue in on things with my mouse. Dan: Right. But I think it would be great to see more of this hardware. Our next guests are not working in hardware, they're working in software on an app called Fugu UX that is all about detecting usability and accessibility issues in products. And we will have our conversation with Jason Hong and Shavik Das coming right up. Nik: Now we're sitting down with Jason Hong and Shavik Das from Fugu UX, an AI user testing and automated usability audit company. Jason, Shavik, thank you so much for being on. This is actually super cool to talk to a company and people who are thinking about user testing. We talk a lot on the show about creating designs, but actually there's a lot where you could imagine using AI to evaluate and audit the work that you're doing, especially all the work that's out there. So Great to have you on. curious, maybe kick us off. Tell us a little bit about what you've been working on and what you're excited about. Jason Hong: Yeah, thanks for having us, Nick and Dan. so the basic idea behind Fuga UX is that we wanted to make it dramatically easier for companies to find usability problems on websites. I sort of joke that you know in the movie Sixth Sense, the kid sees dead people everywhere. And the curse that I have is I see usability problems everywhere and it just drives me bonkers. And so I felt like you know it was time for AI to That it was ready enough that it could actually find a lot of these kinds of usability problems. So my co-founder and I, Shavik Das and I, we decided to try to create a company that could help automate this whole process, making it dramatically faster and cheaper and easier for people to find these usability problems through AI user testing, and then also prioritize which problems to find as well, too. Dan: what can we do now with AI that we couldn't do 'cause there have been kind of usability testing kind of services in the past, but what can we do now with AI that is different, unique and cool? Sauvik: I like to think of it this way. So I think, you know, historically, if you you could you could go back to like the late eighties with Jonathan Gruden's paper about the case against interface inconsistency. I might be getting that title a little bit wrong, but you know, UX And has always required some level of kind of like subjective judgment. And that's been really difficult to codify into fixed rules. I'll give you an example of this. So information sent is a very simple kind of concept conceptually in HCI theory. It's basically, is there enough context around a link on a web page such that the user would expect to go where that link takes them? Very simple concept, very easy for us to understand, but there are trillions of links on the web, and it's impossible to come up with a standard if-then set of rules to assess information synth. So until very recently, it's been pretty difficult to codify a lot of these kind of like usability heuristics into reasonable rules that give you reasonable results a lot of the time. we're kind of there now, because you can provide a lot of context. and many kind of like modern, kind of like frontier language models have enough of this kind of context for you to be able to sort of construct simple enough rules to be able to assess things like information sent. These like very simple concepts that have historically been very hard to codify. So that's one thing that AI can do now. it can sort of like scale up judgment in that sense. two user testing is incredibly important and nothing that AI does right now should replace user testing. But user testing has also been historically quite slow and quite expensive. And so when you do the user test with the real people, what you really want to figure out is is the product meeting user needs? You don't want to be, you know, have your user stumbling into like very simple, sort of for lack of a better word, dumb problems. Like they can't see the link because it's, you know, not accessible enough. and so with our synthetic users now with agents, we can sort of rapidly identify these like simple issues that hopefully you fix before you go to your real users and try to figure out is your product really solving a problem that is addressing human needs? and so that's also not possible in a way that was never really before possible. and so the combination of these two things sort of like gives us this kind of like new frontier of what AI can do for UI and UX. Jason Hong: Yeah, and following up on what Shavik was just saying, AI r just changes the whole economics of user testing. So if you look at how a lot of corporations do testing today, there's something called the testing pyramid, where there's a lot of unit tests that are done and a fair number of service kinds of tests, so these are end-to-end kinds of tests, and very few kinds of user interface tests. And that's not even user usability tests, that's just making sure the user interface works. And so all this is sort of done internally, and then if you're lucky and if you have enough time and money, you can also do some usability testing as well. But with AI, the economics of that really changes dramatically where it's relatively cheap and easy now to do these kinds of usability tests with AI agents. Again, we still advocate that you should have real users, but now with AI user testing, you can iterate much faster and cheaper. And at the same time, while you're doing these AI kinds of user tests, you can also check for a whole bunch of other kinds of things as well, too. So you can shift up the pyramid so that you can do a lot more of the kinds of user tests and also at the same time get the benefit of doing service tests and unit tests that the AI can also just build in for you at the same time. Nik: I'm curious, when you're talking about doing audits on websites, are you talking about products that are primarily built already? Or are you talking about doing them on products that are being built? Because one of the questions I have is if I'm using, say, an AI code agent, couldn't it also, if I tell it, hey, don't make dumb usability errors, and I prompted it in a certain way, you would just get usable code from the outset. Jason Hong: the current way that we're targeting things is primarily for sites that are already built. though we're also considering how it could address the use case you just mentioned, which is you know very early stages of design where you might want to know like, well, what are the competitors doing or who are the potential competitors, what kinds of features do they have, what kinds of user interfaces do they have, and then identifying the key user journeys of those, and then perhaps, you know, mocking up some potential user interfaces on what likely key user journeys you might have. That's definitely something that's is on our radar, but we decided that we wanted to focus on existing websites first because we felt that was easier and there's already a really large market for that too. Dan: Yeah, so who are the target customers for this? Is it kind of small and medium businesses, large business like who like who's the service for? Jason Hong: Yeah, that that's a great question. So we've been primarily talking to a lot of product managers, the people who lead a lot of these kinds of teams, because one of the things that we found is that there's always a lack of usability in these teams. It might be that the team might not have An actual usability person, they might not have a UX designer on their team that can help run these user studies. But even if they do have a usability person, another challenge that we've also seen too is that that person tends to be overburdened. They might be split across multiple teams, and they might not have enough time to actually devote to all the different projects that they might be assigned on. And we've also seen that there's a lot of these internal intranet kinds of projects that just have pretty much no usability at all. And so this is also something that we felt could also help those kinds of teams. Dan: Mm, mm-hmm. Nik: Wait, so are you gonna be able to like f fix, you know, some Oracle system that I have to use internally? Jason Hong: Yeah. Dan: Ha ha ha. Jason Hong: So some things are beyond hope. this is actually something funny. Shavik and I were discussing this on our internal Slack channel, a while ago, which is why is it that B2B software is just so dreadful? And the the economics of it are just that, you know, People will buy the software for B2B software before really evaluating the usability of it. compare that to B2C software, business to consumer software, where the usability of the site itself is part of the whole experience. And if it's hard to use, you're probably not gonna convert. But that's not really the case for B2B software. And on top of it, the other challenge too, is that there's a high kind of threshold for pain that your problems are so high and so difficult for managing things like you know, enterprise resource planning or your employee workforce management, that you're willing to accept training and other kinds of things because the software costs so much and the training is relatively small, so you're willing to accept that. And then lastly, there's just not that many competitors for a space. So the kinds of things that Oracle offers, there's not that many competitors there. Dan: And once you're embedded in like a workforce too, it's very hard to untangle that. Like so if you're you know, we use workday and we've probably had a contract with them for you know, two decades and so it's hard to untangle that when all of your systems are now all based on a B to B software. And there's no upside for them to like going around fixing usability problems. Like we're already there unless they were gonna lose us as a client. Sauvik: It's funny that you should mention it because Jason and I actually I won't mention which system, but CMU has asked us to provide, some usability audits on internal systems that we may or may not use in the future. And they're pretty bad. and, you know, we were just talking about, well, we can definitely show them how all of these things are bad, but where is the incentive for them to fix it? And so that's the other thing that we're hoping, longer term. So going back to Nick's question a little bit earlier about, well, can't I just tell Claude code, hey, make this usable? I you you can, and they're pretty good right now at like getting some of the basics out of the way. But I've actually done A-B tests, you know, so like you just ask Claude, hey, make this a beautiful design, make it usable. And, you know, certainly there are prompt engineering techniques you can use to get the most out of them as far as usability goes. But You know, having like kind of like pointed insights into like, hey, the information sent of this link is off for these reasons, that gives Claude Code and these other coding agents so much more kind of like context and specificity to go on. And they can make, you know, both macro fixes and microfixes at much higher levels of fidelity than you could by just kind of like prompting them generically. Jason Hong: Yeah, I I actually have some conjectures about this as well too, that we can use a lot of these AI design tools to create websites and other user interfaces for us, but I suspect that the weak points are when you start scaling up. So once your website becomes really, really big, trying to make sure that every page is sort of consistent and usable and makes sense, that's gonna be really hard. Especially if you have multiple team members who are working on different parts of the site. so it's basically sort of like a Conway's law kind of issue where, you know, your software architecture matches your team your team architecture or you know how your team is divided. And so, you know, it's going to be the parts where one team's user interface interfaces with another one and that's where the inconsistencies are going to happen. Sauvik: Yeah, I think there's like this long tail of usability issues that never get addressed because they just snut high enough up on the totem pole. And I think this is where a lot of value can come in as well. Nik: Another question that I had was how primarily are you approaching this? 'Cause when we're doing something, say Creating code and working with Cloud Code, we're usually sitting there as a designer, we're looking at stuff, and then we're translating a usability issue into language and then working with the a code agent to fix something. And actually, you even said, like, you might say, hey, this link doesn't have a good information sent, let's improve that. That's your translation into language. But if you're working on sites that are already developed, how are you, how are you approaching this? Are you using like visual language models? Are you doing something else to do this? Jason Hong: Do you mean how we're finding problems or do you mean what we're recommending as fixes? Nik: How are you finding problems? Like, you know, if I actually like give you something, like does it interact with my site? Do I set up tests? Does it just start navigating it? Does it figure out tests to do? And then is it working from a visual perspective? Because this is also something that so many things can be seen, but they might not be as easy to determine from even the page source, especially if you're using technologies where most of your page source is actually gobbledygook React code that that's all hidden. Jason Hong: Shavik, you wanna talk about the pains you've been experiencing? Sauvik: Yeah, well, we don't have an hour, so I won't get into all the details. But we use a a variety of approaches, let's say that. So for our AI user testers, it's sort of like what you're saying, Nick. All you give us is your website, and then you give us, for example, a task that you would want your users to be able to complete. So in your case it might be something like, you know, Hey, make sure that it's easy enough for a PhD student to find out whether or not I'm recruiting, or something like that. And then, you know, one of the one one of our offerings is AI user testing, where we will instantiate a series of AI users that have different personas, different digital literacy literacies, etc. and they will essentially start from your homepage, the homepage that you give us, and try to Do that task. And they'll try to do it in a diversity of different ways to kind of like emulate what different sorts of users might do. now I want to be clear here that this isn't emulating anyhow any individual user might do it, but this is emulating you know statistical patterns of here are some like broad ways that users might approach accomplishing this task on your website. And then it'll identify so it'll actually be like clicking on links, clicking on buttons, opening up accordions, and it'll think out loud. It'll tell you, you know, quote unquote think. It'll tell you it's like, hey, this is why I'm doing this next. and at the end of that, you get kind of like a report of like where the synthetic users got tripped up while they were trying to accomplish this task. Now we also have is an an audit-based offering where you just give us the URL. And then the audit agent essentially tries to understand, hey, what's the design intent of this website? What type of website is it? So hey, it's nickmardo.com. here are some competitors, right? Like so so so so Nick Nick Nick's website is a lot like, I don't know, like Michael Bernstein's or Jason Hong's and whoever else. and then it'll say it's like, okay, so this is what Nick's website does, and this is what those other websites do. Here are some example user tests we might want to run on. on Nick's website and then we instantiate the synthetic users to see whether or not they can accomplish that or not. And then it also uses sort of like, you know, visuals, it uses front end source code. and you know, there's certainly been a lot of crap we've had to do for lack of a better term again, to make sense out of that, you know, front end gobbledygook, as you say. but you know, we're we're we're getting closer and closer to something that's like really sort of like actionable and useful. The the the longer we work in this space. Jason Hong: Yeah, I actually have a a funny story related to this. So as we've all experienced, CAPTCHA's are one of the things that are designed to block you know, AI bots from using websites. And so sometimes when we have our AI agents go through websites, they'll see the CAPTCHAS. And sometimes they'll do really funny things. So one of the the agents are actually narrating what they're thinking as they're going along. I'm using the word thinking very loosely, of course. So one time one of the agents said, I really hate CAPTCHAS and clicks on it and it's like, good, I can keep going on now. Or another time we had a CAPTCHA that said, you know, I really hate CAPTCHA, but I'm really glad that I can get through them so fast. it'll help me save time and I can spend more time with my kids. So the the AI agents really say some really bizarre things sometimes. Sauvik: Yeah, yeah. if I can actually add a little bit more color to that. So the agents are smart enough that, they have these backstories, essentially. And so they'll they'll know for example what to enter into different forms. So we had an agent that was reviewing kind of like this school website and instead of putting in, you know, its own name, it put in its quote unquote daughter's name. Because it knew it was an elementary school. So it's like I said, I'm put in a daughter's name here. And so w we actually have a lot of like internal agents do the darndest things kind of memes, going back and forth between us. Dan: So so how does that work by do you have like different agents for different domains then? Or do they come up with different backstories depending on what they're being thrown at? Like how does that work? Sauvik: Yeah, so we actually have a stable set of twenty agents right now because the backstories are quite rich. And so essentially what I did was I looked at kind of like US Census microdata and I sampled from that space in order to get like a diversity of agents with different digital literacies and different you know walks of life, etcetera. we're careful not to try to, you know communicate this in a way to suggest that, hey, this is like a middle-aged man or something like that. but instead we just want a diversity of quote unquote backstories so that the agents will do different things. so depending on your task, for example, if you tell us it's like, hey, we want agents of a certain demographic or something the honest is on us to say it's like, hey, these aren't people from this demographic, but here are here's like the the the closest facsimile to that that we can give you. but also you could just say it's like, hey, I just want like a diversity of agents or something like that. And then we can, you know, sample that space and give you five agents who are quite different from each other, who will go about doing a task in different ways. Jason Hong: Yeah, it it's also worth pointing out that this is a really interesting issue with user testing in general, which is It's okay to have people try the exact same paths because you're really trying to understand the nuance and the range of behaviors. but in some of the AI user tests that we do, we actually explicitly force diversity. So what I mean by that is agent number one, we want you to primarily do search. agent number two, we want you to primarily do browsing. agent number three, do something different from agent numbers one and two. because we're trying to increase the range and coverage of things. And so that's something that We built in as a feature because we felt that a lot of the people just wanted to know sort of the range of issues and they can also qualitative look at things and see does it make sense for them to make the change or not based off the results. Nik: One of the things that you mentioned earlier is right, there are a lot of heuristics that we often use as designers for usability testing and usability issues. Like we do a lot of use you know, heuristic audits on our own. are you building off sort of our traditional heuristics, do you did you have to develop a new set of heuristics to work that actually like these AI systems can sort of actually an enact testing? Jason Hong: Yeah, so we started it out by compiling a really gigantic list of heuristics, ranging from small low level details like for example HTML metadata, search engine optimization, and so on. And then there's some flow level things all the way up to like, you know, site-wide kinds of issues as well too. And then there's also just generally cross-cutting kinds of heuristics, so the most famous of which would be Nielsen's heuristics. And based off of that, we also gathered a lot of examples of those, and we also got a whole bunch of websites where we manually labeled that to help train the AI systems as well. So in some ways, the AI heuristics that we have, the AI bots that we have, also are somewhat trained on Shavik and my own brain, where we were labeling things that we said, this is good, this is bad, this is actually a very good rating, this is a very bad rating, and so on. on, you know, a few dozen kinds of different websites. And so in some ways I guess the AI bots are sort of like our own children where we've been trying to teach them different ways of doing things and trying to teach them what's good and what's bad. Dan: but it sounds like they each have taken on their own personalities so it sounds like there's like this combination of like your heuristics and like the world's heuristics plus the origin story and so they now each have their own personalities and motivations seemingly. or am I over characterizing it? Jason Hong: Yeah, this is o of course one of the dangers of AI is that it's really easy to anthropomorphize things. especially when you're watching one of our AI user tests and it's narrating and it's also telling you things and sometimes is making jokes. sometimes they'll even start speaking in Spanish or in Japanese, which is really surprised us. so I I think there are elements of that. Dan: Ha ha. Jason Hong: i i it is hard not to think of them as people. I whenever I'm trying to comment on a bug that we found, inside one of our agents, I always have to be careful to say it or he or she or whatever, because I'm trying not to anthropomorphize them. Part of this is, you know, us as computer scientists and also researchers that we don't want to mislead people to think that it can do far more than it really can. but also just try to make sure that people understand, you know, there are limitations to what we can do today. But the argument we also make is that most usability issues are actually fairly basic. They don't require that much domain knowledge. So, you know, you can grab any of us right here in this podcast and we could probably point out on any website, even if we didn't have domain knowledge, we could probably point out, I would say, probably sixty or seventy percent of the usability issues because you they don't require domain knowledge. Things like this navigation bar doesn't make sense or it's too small to click and so on, is almost context-free. Though there are a lot of things that do require some domain knowledge. So for example, one thing that some customers have asked us about is Is things like, well, can we have domain knowledge inside of it? So for example, I'm working on construction, and I want to make sure that the website, we do use a lot of construction terminology, but we want to make sure that it's something that someone who has already worked for a few years would be familiar with. So those would be things that would be outside of sort of a generalist kind of expertise. so we are looking at ways of how we can incorporate that kind of terminology, that kind of expertise, but right now it's still sort of experimental. Yeah. Sauvik: So I I kinda like to characterize them as my little Tamagotchi pets. You know. they are they they certainly do their own things. and that's my design. And that's actually helpful when trying to identify usability issues as well. Because, users also sometimes do the darndest things on your website. And so it's helpful to have broad coverage of like here are different ways that your website could be stress tested by a user trying to accomplish this kind of goal. So it's not any one user who would do exactly this, but here is a likely path that some set of users might take. Is your website ready for it? know, that's the sort of thing that we hope to be able to identify and flag for people so that they can prioritize accordingly when they're making their next update. Jason Hong: Yeah, if I could sort of go back to one of the earlier points about, you know, what's the value proposition and how would you use this? You know, the current state of user testing might be it can take days, weeks, or even months to do the recruiting, to run the user tests, and to summarize the results. And here we're saying, well, you know, we might not be as nuanced and as in depth as a lot of these full fledged kinds of user tests, but you can actually get really good results in a matter of hours, and that makes it possible for teams to iterate far more quickly than they can today. And so that is by itself a special kind value proposition and especially for teams that just have very little usability support right now. Dan: Is i is the output then going to be something that you could pass off to a coding agent and just be like, Okay, well, the Fugu s found all this all these problems. Here you go, Claude Code or Seoul or Codex here, take these and see what you can do with them. Jason Hong: Yeah, th that's a a great point. that's something that we've been discussing internally. So right now the output is a report of usability issues along with the evidence of those issues. though we have been discussing things like, well, you know, export to JIRA or export to linear or any other kinds of ticketing systems, seems like that would be a natural match. And then maybe exporting in a format that some kind of markdown file that these coding agents could use in the next iteration of things. In fact, you You know, sort of a pipe dream that we have, a really long-term kind of vision, is what if we had auto-updating websites where it's going through the whole loop for you? So as people are using the site, then it comes up with redesigns, it runs simulations on those redesigns to see which one is most likely to work best, deploys the site, and then gets real data on it and then iterates. So now you almost have like a constantly updating website that is being tuned to what users are trying to do. as well as actual user data along with simulations. But that that's still like, you know, far past research idea stage. That's you know sort of beyond the state of the art of what anybody can do today. But I think that would be a really great kind of long-term vision. Sauvik: I think also worth mentioning here is that that is not what Fubu does right now. So you know, going back to your original question, Dan, I think was, you know, is the output just something you can hand to Cloud Code? I think right now the model that we have for our customer is really kind of like the product manager. and so we still think very much so that Jason Hong: Yeah. Dan: Yeah. Sauvik: based on kind of like what we're finding right now, not all of these fixes absolutely need to be implemented, but rather we're flagging them so that somebody who really owns the product really understands it, can work with their team to figure out what's what what's the appropriate thing to do next. Now not all websites and companies are gonna be so well resourced to have like an entire dedicated design and edge team for every single component. And so longer term we do also wanna provide an offering where you know we can sync with your Claude code. And as you're doing this, for the websites that, you know, hey, like a nonprofit, let's say, hey, we really don't have the design or the engineering staff to be able to do any of this. maybe this could be a cheaper alternative to get like quick wins. It's never gonna replace good design judgment from a real designer or good code from a real engineer, but it could be good enough for a lot of these kind of like long tail use cases. Dan: Right. When I've spoken to engineers, one of the things that they love so much about the era that we're in is that They don't have to deal with the P threes and P fours. That that is something that designers and or AI can just handle. And so I'm imagining that that is the world where these kinds of like, this doesn't have enough contrast. Well, okay. Go go do that. Here it is. Here's the problem. You figure out how to fix it. Sauvik: Exactly. You don't have to go through like an entire PR review cycle to increase the contrast of this button, right? like my model for all of this is I would love to live in this world, where this minor version change doesn't necessarily require the full loop of the entire team all the time, right? Like so when you go from version one point one point two to version one point one point three. You know, Fugu can be in that process that allows you to like iteratively improve like these small things on your website based on evidence. but you're always gonna need kind of like the people who are really like using their judgment to figure out are you fixing a real pro a problem here for a real person to go from version zero to version one, right? Like the zero to one transition, I think, is always i is the part that's most interesting to us, and it's the part that I think You know, if you just give it to an AI, it's not gonna do that well at. But going from one to one point one or one point one to one point one point one, you know, that's the sort of thing that I think we can automate. Jason Hong: Yeah, what I was gonna say too is that respect to Dan's question, different organizations organize their teams in very different ways. So for example, it might be the case that things are divided up by a specific pages. So this page is owned by this team, this page is owned by that team, and so on. Or it also might be the case that teams own specific components. So this team owns the login, and this team owns this content module, this other team owns this other content module. And I think the challenge here for a lot of AI in general, not just our stuff, but in general, is how to align the capabilities of the AI such that it can maximize the effort of those teams. Because you can imagine that if it's something like contrast, then it's like, well, contrast cuts across the entire page, it might cut across all these different teams. And then there might be endless debates about it and so on, which is not really great for efficiency. But that's something that the AI could also do very quickly if everybody agreed on it. So I suspect that the long term issue is that AI can really help out with productivity in a lot of places, but it also requires a lot of changes in how these teams are organized. It's sort of like the productivity paradox that we saw in the 80s and 90s with computers. You know, the joke was that you could see the computer revolution everywhere except in that you can. economic numbers. And I suspect that, for a lot of teams it's going to be the same thing that until we can start building all these amazing kinds of capabilities, but we have to figure out how to align the humans and the AI together so that we get the best results. Dan: end. Nik: Yeah, I think that's a great place to end. Jason Shabik, thank you so much for being on the show. And for those who might want to reach out, do you guys have contact information if people wanted to learn more about Fugu? Jason Hong: Yeah, come to our website at fugux.com, F U G U U X dot com, and you can get a free trial and see if you like it. And yeah. Also appreciate any kind of feedback you have. Dan: Thanks for joining us here at AI and Design Podcast. We'll be back next week with more AI, more design. We'll see you then.