Basil Chatha: 90% is being generous. It's like 99.9%. I don't think I've written a line of code in like eight months. But when I actually did my onside appinch, like I like came in with Claude Code you Sonnet 3.7 in my own API key, and I was like, we're gonna cook today. Expectations for what like an individual engineer can accomplish have like shot through the ceiling because now you're no longer bottlenecked on the volume of code that you can produce, like the cost of producing code trending towards zero. I do think that there's a certain type of engineer. Who's losing big right now? And that's like this sort of ivory tower, like esoteric syntax type of person. Meta is now measuring how effectively do these engineers adopt AI as part of their performance reviews. I think like Microsoft is actually tying those to manager performance reviews. So how well does your whole team adopt AI? So today's episode is with Siddharth Nanda. He went from writing every line of code by hand at Microsoft and Atlassian to shipping twenty five or thirty thousand lines of code a week and almost none of it's written by him. He's currently working at Finch, which has raised more than $20 million to build an AI-powered platform to help automate admin tasks at personal injury law firms. So it's really interesting to see his perspective on working in big tech pre-AI. And now at a startup where more than 90% of the code that he writes is written by agents. We get into why he hasn't written a line of code in eight months, what that means for engineers, PMs, and designers, whether the classic software engineering career path still exists. We also talk about what happens to managers who have stopped coding. Why the cost of producing code is going to zero and what that means for how teams get built and run at startups and at enterprises. You don't want to miss this. Let's get into it. So you were a software engineer at Microsoft for a couple of years. I would say like before ChatGPT came out, right? That's right. Yeah. So can you talk through the engineering process at Microsoft maybe before ChatGPT came out? And I'm assuming that engineering processes didn't really change for maybe like a year or two, even after ChatGPT came out, just because the coding agents weren't good enough. And most recently they they've gotten good enough to where maybe you should be using them. So yeah, maybe you could talk about how was engineering before then and how have you guys been changing your engineering process like recently because now you had a new company also. So yeah, absolutely. So I started at Microsoft in early 2021 as a new grad. So I think like two things there, like my programming abilities were fairly ⁓ novice level. So a lot of hand holding, a lot of guidance was needed onboarding to like an enterprise code base and like being a full time IC for the first time. Back then, I think like it's a very relatable experience to anybody who is in the industry. So a lot of focus on syntax, learning a new language. ⁓ all the code I wrote at Microsoft was written in C sharp. So it was really about writing syntactically correct functional code and putting a lot of effort into testing it. So in those days, ⁓ if you have a new graduate on your team, one of the first things you might do is actually put them to work writing tests for existing code where, you know, there's been reliability issues or the coverage has been weak and That's kind of what my first project was as well. Like just raising the test coverage of that code base and sort of like exploring there. ⁓ I think like two things were different as well. Like one, the iteration process for how you actually went through writing a unit of code was much smaller. So you would actually set out to write a smaller block of code that did maybe like one or two things, and you would test it quite rigorously before you shipped it. I think that has clearly like evolved. ⁓ through time. When I was out of LASIN in, you know, 2022 to I think like late 2024, that was when we had our big like chat GPT moment. And you'd be surprised because adoption was already happening sort of quietly behind the scenes. Some of you guys might remember like the IntelliSense plugin for Python that Microsoft released in VS Code. So There was quite a bit of hey, like let's start experimenting with tab completions, let's start experimenting with autocompletion, even before it got, you know, to the extent of what Cursor got very famous for. Most internal releases at big tech companies at the time were through like a single, a very gated portal. So you had to apply for access. Maybe there were like two models being served there, ⁓ GPT 3.5 turbo and GPT four. And you could use it very sparingly. The rate limits were very, very low. So Kind of working on coding use cases was not the thought for many, many people at the time. It was a lot of, hey, let's put in this sort of unstructured data and see what comes out. Let's see how we can sort of productize this with maybe replacing some of our existing ML solutions. And the coding moment happened maybe like six months after that, where you saw an explosion of software engineers at the start just copying and pasting code into those internal portals to write unit tests at first. Interesting. Wait, so six months after, so when would that be? Like twenty twenty four? I think it would be closer to mid twenty twenty three when all of that was was happening. I believe the initial release of like Chat GPT was early twenty twenty three or even late twenty twenty two. Yeah, I think it was December twenty two. Yeah. ⁓ so okay, that makes sense. Yeah, because like personally I remember when it came out, the first stuff that I used to do with it is like I know Python really well. But I don't know JavaScript super well. So I would literally write something in Python and I would be like convert this or translate this to JavaScript. And that would really help me like learn actually. So I thought that was like really interesting. Yeah. And I I think like the ambitions at the time using those tools were like quite small. Like, hey, like paste this function in and then convert the language for me. Maybe not like take this whole module and go read back. Yeah, yeah. And so now now you're doing more like bigger changes. Like you're Literally telling it, hey, just go refactor this entire code base? Or how do you use it now? For context, if we look at like the last week alone, I think I produced maybe twenty-five thousand to thirty thousand lines of code. And I want to preface this by saying I know that's a horrible metric for judging things. But about like sixty sixty-five percent of that was actually done fully on Devon. So I was just multi-tabling Tons of Debian windows, like knocking out a whole bunch of smaller things, some feature work, some planning. And then the other like sort of 35 to 40% was done entirely with codecs, actually. So zero handwritten lines of code in there and everything from like fixing smaller bugs, launching smaller features to working on really big pieces of like a project that we have ongoing. Yeah. ⁓ what did an engineering team at Microsoft, for example, look like? When you when you joined. And I guess has that changed now? And how do you see that changing maybe like a couple of years from now? I think like the can't speak too much to Microsoft since I was there for about a year and a half. But the two interesting things that I can say about that was they were both pre flattening teams. So in twenty twenty two, there was a big, you know, hey, we're gonna eliminate middle management and like sort of make our engineering teams like broader across the board. ⁓ This was before that happens. Teams tended to be in units of maybe like six to eight engineers, ⁓ working with like a manager, a tech lead, PM, and maybe also a designer. And they tended to own like a narrower scope. So you might own like maybe one microservice or like one unit of work as a team of that size. During my time at Atlassian, we started as one of those teams, maybe like six to ten people. And then Host flattening expanded to a team of almost like 20 people owning multiple microservices. So there was more of a push to, hey, these engineers are working across teams quite frequently in these sort of modules. Let's go ahead and like group them together under single manager and hope that the execution speed goes up because there should be like less decision making friction between these like pools of engineering teams. Yeah. And was that true? Like, was it that? Like the friction decreased and you guys could actually increase your productivity when you increase the team size? I think it was difficult to measure for a long time. But ultimately I do conclude like yes, the sort of initial assumption that like a team of like six to ten engineers was the maximum that you could have was probably an incorrect one. Especially at like big tech companies which have varying levels of process and bureaucracy in Things like approving designs, I think has made a meaningful difference in what you can ship and how you can actually get alignment if you're only needing maybe like one or two like senior engineers or tech leads to sign off on your proposal versus multiple of those across teams and multiple managers as well. Yeah. So also on the manager front, so like basically the number of managers did decrease at Alassian. Okay. Yes. I think uniformly across the industry, just an annihilation. Yeah, yeah. So last week I was talking to Lindsay Simon at Russell. ⁓ and he was like I like I looked at his GitHub commits and it was like, ⁓ back in like twenty seventeen, twenty eighteen when he was like an IC, he was doing like two thousand commits a year. Then like he became a manager and it went down to like three hundred commits a year until you hit like twenty twenty three and then you see his commit volume and it goes back up to like two thousand and he's still a manager. Which is like pretty crazy. The expectations have changed as well. I've seen things at companies like Google actually have, I think, always done this. They've always had some managers continue to operate as tech leads. And that was largely seen as unique and not industry average. But I have friends at like Ramp now who like tell me that h they have like something like six to eight direct reports and the expectations of being a phenomenal IC have not gone away. Like they are pushing just as much code as they did as ICs before they had any direct reports at all. I also think that's part of a push to sort of make like management and performance cycles like happen faster and invest more on like internal technology to sort of report up on performance and spend less time on those cycles as a whole. Yeah, interesting. So have you have you noticed like what's happening to these engineers or I guess like the managers who, you know, they kind of lost it. I feel like there were a lot of managers who They go into management because they don't want to write more code. You know, because they're kinda like, ⁓ they have families now. They're a little older. They they're kind of like, ⁓ I just want to manage teams. You think that's basically gone now? Like you cannot do that anymore? That's not a valuable or a viable career path. I think I think this is a complex answer. There are certainly individuals who, you know, after maybe like some amount of tenure, want to do less work. And I think the pathway for them to get there has changed. Maybe that's now more going towards like a classic like Fortune five hundred company or big bank or somewhere that's still in an earlier phase of like tech modernization, like still investing in, for example, like cloud migration. Overall, what I have seen is a large number of managers actually really embraced being great ICs again. So these were people who were phenomenal ICs before they became a manager because of like need. Hey, you are one of the best communicators on this team. We'd like you to be a manager. ⁓ and they were They pretty happily returned to being ICs. And then the other thing I I've broadly seen from there is the managers who sort of elected not to do that, who continued in management roles, actually like really stepped up to that. So they, for example, like my manager at Alassian, fantastic guy. He suddenly had the number of his direct reports double. And she really stepped up to the plate. So we didn't see like anything change from him, like outwardly in terms of. Hey, like the level of detail or attention he was giving to like the amount of work that the team was doing or what we were working on or even performance reviews. So I think like maybe everyone had some like excess capacity that they could absorb. ⁓ and then sort of like the final point there, I don't know if you've noticed this, but I've increasingly seen a trend of, you know, people in tech in their early 30s like signaling more like willingness to be like, yes, like I'm gonna have family, I'm gonna have kids, and I don't really intend to. dial down my ambitions or my work schedule, I'm gonna find a way to have everything. And I think like more of the people I've seen like fit into that bucket than like the first bucket of like, hey, I want to do like less work and like sort of wind down ⁓ the more ambitious parts of my career. Interesting. How has the role of a manager changed? 'Cause okay, so now now a manager has to start like writing a lot of code or they want to. Maybe it's a combination of both. So yeah, how has their role changed, if at all? I think At big tech companies, I still don't think there is a huge expectation that the remaining pool of managers is significantly involved in the code base or writing code or doing things that aren't like traditional things. I think they now face like more pressure to like facilitate among like a more direct group of like maybe directors or like VPs and are reporting to like higher levels of the organization than previously and on more initiatives, right? So they're simply managing more people. Therefore, running more projects and therefore that like reporting workload has actually increased for them. On the other hand, you have like sort of what I'll like call like gigastage startups all the way down to like maybe seed and series A companies that have managers for the first time who are writing a lot of code. And I think like the extent of their management responsibility really varies. For example, I'm at a startup right now. I report directly to the CT. And our CTO basically empowers every engineer to be like, hey, you own an entire product surface area. I have absolute trust and confidence in you that you're gonna like solve the problem. I'm here for input, not to like be like, you know, hand holding you or like driving technical direction like for like any specific aspect of your work. And he's more working on connecting all the engineers together who are owning like huge scopes of that project and making sure that, hey, we have a cohesive experience in the app and that the app in general is like facilitating success for operations. So that's certainly one model. And then the other model I've seen is like. That sort of combination of like a super IC who does have direct reports. And these individuals are actually largely resorting to one, like AI powered, like workspace management tools for things like performance reviews, for things like collecting data. A friend of mine who's manager at Ramp, like regularly uses like AI to analyze all of his direct reports commits. So you can have like communication of here's what we actually did this week in a set of bullet points. And that really like streamlines the workload of like, hey, I need to know in great detail what each individual is doing. And since you can do it in five minutes with AI versus like 30 minutes in like a one on one conversation. Yeah. Interesting. At some of these like smaller companies, I guess, maybe just like in general, how do you see the role of a product manager, like a designer changing inside of a inside of a team? Difficult question to answer. My company doesn't currently have a product manager, for example. Yeah. So essentially their role is broken up between engineering operations and even our users are like acting as like quasi-product managers and like really taking a hands-on approach to driving the direction of the product. I think this approach works better for specific kinds of startups that maybe are more focused on hey, like operations and outcomes versus like, hey, like we are selling. A a like traditional SaaS-like product that has like seat-based pricing without like maybe having that direct face-to-face with users as often as they'd like. And so having a product manager helps like facilitate that relationship, basically, when your users are external. I do think product managers, from what I've seen, are being asked to do more in terms of both like making prototypes. So hopping in with tools like cursor, tools like Cloud Code, and actually coming back with, hey, like, This is how I think we can accomplish this, or this is how I think this can or should function. And like on the design front, so product managers are also hopping in like products like I think like paper is one that's pretty cool ⁓ as well to to come back to you with like mocks of like here are like real components in your code base. Here's like the sort of mock I came up with. ⁓ can we do X, Y, or Z? Designers' roles are also changing as well in that, again, I think like, Cloud code, like all of these coding agents are basically transformative technology here. Being asked to do something they've never done before and like hop in directly, ship their changes out. And I think like that field is actually really hasn't had its like aha moment yet. Like I think the tools are less mature. Like there isn't like a clear incumbent, there isn't a clear winner. But we certainly have like a lot of high hopes for like, hey, how can our designer become like more involved and ship things themselves more independently versus like having to like come up with something and then balance those priorities among like different engineering pods. Yeah. So you said before that most of your code, like 90% plus of your code, is written by AI? 90% is being generous. It's like 99.99%. I don't think I've written a line of code in like eight months. Okay. Okay. So 99% plus of your code is written by AI. So is the expectation now for basically anyone on a technical team that they also need to become engineers. Or maybe that like anyone on a technical team just needs to be able to be a PM, be an engineer, be a designer. Like you need to be a hybrid basically now. Cause I was talking to like Lindsay Russell last week. And I was like, so like what's happening to these roles? Like, do you need like PMs anymore? Do you need designers anymore? And he's like, I just view everyone as like a builder. He was like, I think the the like the technical role is just like builder now. And it's just everyone's got to do everything basically. Like do you agree with that, disagree with that? I think Certainly I think the expectations for what like an individual engineer can accomplish have like shot through the ceiling because now you're no longer, I guess, like bottlenecked on the volume of code that you can produce. Like the cost of producing code has actually like is trending towards zero in my mind. And so a lot of it becomes more like, yes, like product understanding, problem understanding, what should we build and how we should build it. But there is a real like sort of infrastructure maintenance, like system design burden that is like increasing on engineers as well. So as you sort of mass produce code at scale, like your time starts to go towards how can I make sure that these like features I'm like releasing aren't slop, right? Like that they perform reliably, that all the different like units of the system like work well together and are connected in a way that make sense from like a data modeling perspective, from a performance perspective. And so you actually start to lean a lot more on those sort of skills versus purely like syntactical, like how should this function look? Cause I think AIs generally produce great code with with some supervision, with some, you know, occasional spot checks that that it hasn't gone too off track. And I think that the expectation that everybody is a builder is certainly true, but you have to be cognizant of where people are spiking. Right. So for example, like you might have a team that has like two like ML leading people who have a really deep understanding of like, okay, like here's how we can like look at reasoning traces. Here's how we can build evals for this agent. You might have one person who just has like a great feel for like how UX should be, like making things smooth, easy to understand on like the front end. And you might have like someone who like is very deep in infra, like just generally knows like, hey, like these are going to be bottlenecks for us as we sort of scale the volume of documents we're processing, or maybe like objects that like are saved to our database. And all of those people are like still naturally like have learned that knowledge and I guess like not necessarily like pre-AI, but maybe during like takeoff times, right? Like in a different environment. And I think all of that is like still really valuable. Like you have to know where to like point and shoot for lack of a better word. And you need sort of these different archetypes to work well together to to really accomplish like a big goal. So I'm still like really, I guess I'm like really bullish on like individuality. Like if you are interested in infrastructure, if you're interested in like front end, if you're interested in design or you're like just like really, really like rigorous about your like product principles, by all means you should like continue to go deep there. AI is not going to sort of like diminish like what you can offer there, but rather like enhance it because you won't be as dependent on others for like the things you're weaker at. Like your strengths are going to multiply and your weaknesses are sort of going to, you know, get get fortified or supported in some way. Yeah. What would you say? Okay, let's just talk about how you've been getting up to speed on like agentic engineering. When did you first start? Kind of like going from traditional software engineering, just like writing all of your code all the way through where you are today. And how has that transition happened? Where are you learning all these things? You know, how have you like what are the trials and errors that you've like come across? Like how has that transition happened? Yeah, absolutely. I think when when I was at like a big tech company, again, the rollout was kind of slow, right? So you have limited access to these like internal AI portals. And then as access started to ramp up, that's when we started to really discover, I think, the world at large. Like here's what AI can do, like with coding, and here's how it can fit into your workflow. The problem with that is actually that big tech companies can't build the tools that you have the demand for fast enough. So, like I said, like in the early days, we were like, please write unit tests for this file. And then you would copy paste like maybe an entire Java file into this portal and then copy paste all the output back into a Java file, like put it in there, fix the formatting, run it, and then iterate like that. And it was just like a very clunky workflow. Sorry to cut you off, but were you even allowed to use like the state of the art models, like Claude or like opening IG back then was just worse, right? So at Microsoft, like nothing had launched yet by the time I had left. But at Atlassian, like we had access to GPT four. And then I think we had access to Claude Opus three. think was like the and those two were frontier at the time. Yeah, yeah, yeah. I'm asking 'cause like ⁓ one of my roommates he was at Google and they weren't allowed to use like Opus or any like ⁓ any of the GPT models. So like they had their own internal model and it was really bad and everyone like did not like using it, but they they weren't allowed to Yeah, I think that was the advantage that I had of just like being at like a SaaS company that is a customer potentially of all of those three like major labs and not like working at a company that was like producing or like working on AI like themselves. Yeah, yeah. So we had like much broader access, I think, than employees at Microsoft or at Google would have had at the time. Though I think like Microsoft, I imagine, got like quick access to GPC models thanks to their agreement with OpenAI. Yeah, that makes sense. But yeah, sorry to cut you off. No, all good. ⁓ sort of like getting back to that, like I just think there was a lot of demand. Everybody knew the workflows were clunky. Like everybody was doing stuff that like we knew like wouldn't scale And so when I actually like left Atlassian in like late 2024, like and I finally had some time to like just build things that I wanted to. Like I had already been like experimenting on my own. And so I would just like be on Twitter, like see the announcements for the new tools, like read engineering blogs and like download them and try them. I think that's like a great habit that's stayed with me through this day. Like I'll test like one to two like new tools per week. I actually just tested one called conductor last week that like manages both of your like Cloud code and your codex sessions locally, that I I thought was quite nice. And I think like X and Twitter are just like a very, very powerful tool. Like if you set up your feed right, if you like follow the right people in the space, like I kind of think that these companies and these labs are actually prioritizing X announcements over any other channel. Like simply following the OpenAI like developer account will let you be like, ⁓ they launched something new for file inputs far faster than like you would have normally noticed. Or like before like their email blast for like the community even reaches you. So I think it's just like some pretty incredible alpha ⁓ there, like just staying being like sort of terminally online and like looking at those and like downloading and trying things. Have you noticed that some engineers are way more productive than others? And what's the difference between between the more productive ones versus the less productive ones? Now, like specifically right now? I think at least at my company. Every engineer is like an absolute like superstar. Like I think their abilities are beyond my own. And like it's a like an amazing experience to learn from them. I think we are all roughly within like the same order of magnitude in terms of like the code we're producing, the number of features we're launching each week, the set sort of outcomes we're able to accomplish. More broadly, I do think that sort of stellar engineers at big tech companies are going to start gapping their like less. talented peers simply because they're going to be willing to adopt anything that will make them like more productive or like achieve an outcome versus like having an ideological stand. There's plenty of like big tech teams I'm aware of that have like varying levels of AI adoption. So you might have like one user who's like really, really into it and they everybody else views it as a sort of curiosity or is like skeptical to adopt. Even at companies where you can see like, hey, like they're actually producing and training a frontier model. there's still that reluctance. I think most recently you've seen companies like Meta and Microsoft actually tie AI adoption into performance. So I think Meta is now measuring like on individual engineer basis, like how effectively do these engineers adopt AI as part of their performance reviews. And I think like Microsoft is actually tying those to manager performance reviews. So how well does your whole team adopt AI? And I think both of those are like sort of directionally aligned incentives and like hey, like you need to adopt this tool if you're gonna like keep up and making sure that that, you know, last quartile of of engineers on your team like are still like gaining these sort of like performance benefits and advantages introduced by these coding agents. How do you make sure that like your team is staying up to date on everything rather than leaving it on the individual to be like, you know, the chronically online, chronically like on Twitter and just like testing things out? Like, do you guys have a method of kind of sharing your best practices? You build a culture around it, really. So I think like especially at Finch, like there's a very high level of enthusiasm to share share new things. So if you pick up a new tool and it's working really well, you'll be like, This is really good. Everybody hop on now. Like this is you gotta try this. That happened when GPT five point three codex was released. Like I hopped on, tried it, like we were all together, like at this big table, and I was like, This is it. Like this model is just so good. Everybody switch what you're doing, get on this immediately. ⁓ And occasionally you'll like receive pushback, which is really interesting. Like, no, I think like Opus four six with like agent teams is like working better for me. And then you'll like have a little sort of like standoff and you'll be like, okay, let me see like how hard this other person's approach gets me. Let me actually put these two things to work in parallel, one in a work tree, one not, and like kind of see which gets an outcome that like I'm more happy with. And I think like building that sort of culture in your organization of like just having like extreme openness about like, here's my stack, like, here's my workflow, like. guys like adopt it, being that sort of having that sort of enthusiasm to share is is really, really important. Yeah. ⁓ okay. So it's it's a little more like informal. It's just kind of like something comes out, someone says, Hey guys, I've been trying this out. This works really well. And they just kind of like share it naturally. It's not like every week we have a meeting where some we're like, hey, just share whatever's ⁓ you know, been I honestly think a weekly cadence would be like too slow. Like there's something every day that people are posting on Slack, like here's a little tweak here, like Here's something new I introduced to Devin so it can do this. Like, yeah, if we like change this like config in our code base, like I've noticed that like the AI understands this type of data model much better. Let's do this. So I think there's the appetite to squeeze every bit of performance you can out of this thing is just insane. Like you need it to work at the highest level of performance because as soon as you like sort of break through a wall there, you know there's even more work you can do on a weekly basis and more you can achieve. So there was like a study where I think it was a Stanford study and they were saying that, yeah, everyone thinks that they're becoming more productive with these coding agents, but when you actually measure it at like, you know, a meta level at the company level, like productivity seems like it's decreased. Do you agree with that? Disagree with that? Seems like you probably disagree with that, but why do you think maybe they're Like I can certainly see like, let's say you take a Fortune five hundred company and you like broadly roll out AI tools. Depending on the culture of the company, you might see people like basically bank those those time savings. Okay, I'm gonna have the same level of output every week, but I'm gonna get my work done faster. And then I'm gonna go like take an take a long lunch, ⁓ maybe go home two hours early. Like those kind of things I think like are going to be like common. I think it's like human nature when you're like not as invested in your work or like don't have like a I guess like a deep like personal feeling about it that you're like, all right, let me just take the time and like head out. But I think like especially the sort of like batch of AI native startups that's like emerged in the last two years, like I think we all view it as like, hey, this timing right now is like existential. I mean, a startup is always like you always feel like you're having an existential crisis. And AI, I think, kinda amplifies that. So you're hoping like if I get this performance gain, if like I improve my productivity, then there's like that little bit of extra like thing I can squeeze out this week that'll like make this thing true or like change the direction of like what I'm working on or even the business as as a whole. So I think this is largely cultural and not really attributed to the tools themselves. Yeah, interesting. So back to the point that you were making earlier where you were saying seems like everyone on a technical team will spike in a different area, but they can overlap much more than they overlapped previously. How do you interview people now? So I'm assuming like yeah interviewing before was, you know, yeah, it was like leak code problems, that sort of thing. But yeah, how do you interview now? Has it changed? So we still like emphasize like technical fundamentals like product sense system design in earlier rounds. I think like in in on-site, we typically like give access to any coding tools of a person's choice. So when I actually did my on-site of finch, like I like came in with with Claude Code, with with you know Sonnet three point seven and my own API key and I was like, we're gonna cook today. I think like it's been interesting to see our different like candidates' preferences in that regard as well. And it also kind of acts as a as a screening factor. Like if you have candidates who are more reluctant to like adopt AI coding tools, but are like seeking to work at like an AI native company, like you really get the chance to evaluate, like, okay, like then how does this person like approach problems? Like, are they being like deeply intentional? Are they like recovering on these different elements when we have them actually produce code during their on-site? And then you get like you know, some superstars who come in like have absolutely like up to date like knowledge of these things and they have all of those sort of like fundamentals that just makes them great, great candidates. And then you're just like desperate to hire them. Yeah, yeah. So like fundamentals meaning can they solve like lead code problems like traditional software engineering? I don't I don't think so this is gonna sound funny to my friends, I think coming from me, 'cause I've done hundreds of lead code problems, if not like yeah, more than a thousand or something like that. And I just fundamentally don't think like lead code itself or like the DSA itself is like what you're being like tested on. It's just like a useful abstraction to like think about how you walk through a problem that has like some complexity, how you like explain your solutions, the alternatives, and then arrive at some conclusion. Now it's mostly a convenience because data structures and algorithms are taught at like every university. Like there's some like baseline knowledge that you can expect all people to have. So theoretically, if someone is like a good problem solver, they don't need to do like any lead code questions and they can just like walk themselves through it with like some book knowledge and like a first principles approach to like problem solving. I don't think like such abstractions are going to be like as needed. Like I think it's going to be easier to evaluate someone's product solving, like or problem solving process. through like how they actually build a thing, because you can just say, here's an AI coding agent, like you're free to use like whatever resources you'd like, but you maybe have to like walk us through it or do like a presentation on it afterwards that like explains all the key choices that shows how you arrived at a conclusion. Maybe even put it in front of a real user or a real operator and have them evaluate the usefulness of what they built. And so you can actually just be a little bit more rigorous about like the same factors that you're evaluating. versus like a lead code style interview, which was designed as a sort of like compact, like convenience based like framework. I'm rather harsh as well in that like I just don't think there's like any anything to like things like standardized tests or like lead code and stuff like that. Like learn a good problem solving approach and like apply it. Like that's like what we are supposed to do as engineers. Yeah. Are there like certain skills that you're starting to shift in the way that you're interviewing or like What skills are you looking for in interviews now? Then that was different from before. Obviously, like, yeah, do do these guys use coding agents and like do they are they up to date on all that? But specifically on the technical fundamentals, has that shifted? Like, are you more on the system design rather than the actual syntactical knowledge? Do you care more about that? I don't think it was ever really possible to drill that deeply into like syntactical knowledge because someone could like still solve the leap code question in like whatever language of their choice and then just like muck about in an enterprise code base and like refuse to use interfaces. Like yeah. So like that was never a thing I felt like we like felt that we had a good ability to evaluate in the first place. I think like the good engineers like pre-AI and good engineers post AI look remarkably similar. Like it's the same group of people. Yeah. It's not really being driven by like AI adoption so much as again, like just great, great like problem solving fundamentals. And I think that still comes through on like things like system design, on things like product analysis. Yeah. I I don't think like the core qualities of like what you look for in an engineer has changed. I do think that there's a certain type of engineer who's losing big right now. And that's like this sort of ivory tower, like esoteric syntax type of person who like made it their whole like sort of archetype of ⁓ I gatekeep because I'm like really good at this like niche thing. And now that's been like commoditized and that, you know, that's been devalued a lot. Do you think it's expanded the pool of potential engineers? Like people who maybe used to be PMs or designers? Like has that those have some of those people started to make their way more more towards like the engineering side? Or is it basically the same pool of people still? I kind of think it's the same pool of people. I think generally what I've observed is that core product people, core design people want to continue in core product and design roles. I also think that. the AI products that are sort of like enabling their workflow are not as mature as like coding agents. And like the types of productivity gains that they're about to receive will look different in like something like three years or six months. I have seen them be more have more time on their hands to be able to talk directly with engineers as well. So they are already like receiving those improvements. They can be more embedded in your teams and they can be more embedded with users as well. So but I don't think like I'm seeing a lot of like step over of like Hey, I'm gonna like convert back or back into being an engineer or convert into being an engineer from one of those like more specialized roles yet. Yeah. Interesting. Yeah, actually back to what you were saying before. Like Codex is better than Opus now, or like in in your opinion? I I have a little bit of a bias. Work on a lot of like at the moment I'm focused on a lot of like very core like back end logical things. And so yeah, Codex feels very ergonomic to me. But Opus four point six is still like a fantastic model. I just kind of think like the cloud code. Harness is going through some ups and downs. Like I have noticed it performs a lot better inside of like factories, droid versus cloud code at the moment. Yeah. ⁓ actually on the harness engineering side of things, there was a there was an article I'm sure you read, like harness engineering from OpenAI. Can you explain what harness engineering is? And then also, yeah, like how how is harness engineering sort of becoming the thing that everyone should be good at? Yeah, I'll give you both the more like consumer oriented and then like the builder oriented like perspective of it. But like an agent is just calling this like LLM API and like loop, right? Like and you're giving it a set of tools, like giving it a set of things it can do. And that's what we like conventionally refer to as like a harness, maybe like a piece of software or like how you've sort of crafted the loop and the tools that the agent can go run with. What we have seen, I think like as consumers is that harnesses make like a huge, huge difference. Like I could tell you probably like which versions of like clawed code like really st stood out to me as ⁓ my God, I loved it with like this model version, only to like feel like the next week, like ⁓ it's not working as well for me or like maybe I've like fallen to a prompting pattern that doesn't work as well for this thing. This is also starting to like show up in like benchmarks. So like all of us like open up like terminal bench and like look at the scores, like on every new model release. ⁓ to see like, okay, like, is it good? And is it good in like their harness? I think a great example of that is like Google's Gemini models, which I love, love to use in production. Like fantastic for like one-off API calls, things of that nature, and just like not very useful in Gemini CLI, which is their like sort of coding agent. So you see like some variance there. I think like from a building perspective, like anytime you're like building an agent, you're like inherently like sort of crafting some sort of environment or some set of tools for it to use. And You kind of want to be like conscientious about like, hey, like what I provide and what I don't provide, like is going to make like a very meaningful difference in like its performance on this task. And especially if I decide to overload it and give it the ability to perform on like multiple tasks. So like finding that balance is challenging. I think like running evals on like your individual harness for like your individual corporate use case is difficult. I'm excited about a number of new like eval companies and eval products that are launching. But like what you're able to evaluate feels very vibe-based at the moment on our side. Yeah, yeah, yeah. What what kind of evals are you running on the agents? We do a lot of like classic, like LLM as a judge type thing. So like I'll set up like a quick, like config driven, like A and B profile or something like that with like a t third judge. And my like sort of approach to this is like to bulk process it by putting that whole thing in its own evaluator loop or like having like setting up like a coding agent to actually read all of the like LLM as a judge outputs and like sort of label those for me into like a CSV or something. It's I would not call it a particularly scientific approach to like evaluation where like, okay, like it's very rigorous. It's being driven by like like industry standards or anything like that. I think like what you're looking for, especially at product companies is something relatively time bound to like discover like that sort of like tree of options that you might be like building for and then just have like some ability to like tune that based off of like vibes or based off of like the output of like a really, you know, almost like simple kind of LM as a judge loop that you've set up yourself. I haven't adopted any like eval products or like eval services into my workflow yet. ⁓ and we are far more serious about like Hey, like this a long running agent. Let's go actually like do a rigorous evaluation of it versus like here's like every little thing that we might build that this is AI and we have like evaluate it very deeply before it launches. Yeah, yeah. What so you were you were talking about the Codex Harness versus Cloud Code. Like which one do you like better right now? I I know you mentioned codex model you like better, but do you like the harness better? I've been writing a lot of code in Codex Desktop, actually, ever since the desktop app was announced. ⁓ interesting. I think like Especially when I I have maybe like ten coding agents running as a time, but like ability to like visually see all the sessions, like in them, get the notifications when they're complete, like see those visual indicators and just like jump around has been like super helpful to me. I know that Cloud Code's desktop app has gotten a lot better over time. I was like an early user there and it was it was kind of not very performant. So I like abandoned it for a bit. But I have been using plot a lot in the last week just because I love how easy their like plug in system is. And so you configure your like sensory plugin, you configure your like Notion plugin. ⁓ and you you just have like a loop of like, hey, like fixing and documenting like different reliability issues and exceptions that are happening in your code base. So those things are kind of sick. I wouldn't say like I'm ever like one hundred percent on one and zero percent on the other. Like you will find like 10 tools open on a machine at any time. Do most engineers at Finch kind of prefer the desktop apps versus the CLIs? I think desktop app adoption is about like fifty-fifty at the moment. Like the real, the real like sort of spike for us is like we're all in love with like Cognitions Devin. We put like a lot of effort. We put a lot of effort into like making it fit into our workflows and making it good. And because we're like a sort of like operations driven company, like the ability to like launch it directly out of Slack conversations with like non-technical like legal associates is just like a godsend. Interesting. You have the ability to like triage maybe like 10, 15 things at once and like expect like really good results because it's able to like very verify things within its own sandbox, which is awesome. Yeah. So I I think like there's a lot the CLI desktop app makes this still like maybe fifty fifty, but the adoption of like Devin is like outrageous. Like it has been the like single largest contributor to our code base, like in the last like three weeks. Why did you guys decide to use that versus like building your I mean, just everyone using the Cloud Code CLI or the codec CLI? Yeah, I think like CLIs are really awesome tools because you get to run them on your local machine. You get to give them access to your local setup. They are very like comparatively very fast. But what you will notice is like depending on how you've like set up your like local environment or the like characteristics of the services you're running, You will quickly become like memory limited. So on my like MacBook, I can run maybe like two different work trees at like a time. And I am like working on like a side project with like another engineer to like spin up some like real sandboxes and like get us access to like just more memory. So we can run more of those, run more coding agents locally, ⁓ quasi locally, I should say. But I think the async use case is different. Like if you go into like Cloud Code web or if you go into like Codex Web, your ability to like configure the container and like give it access to like unique tools or like unique capabilities to like verify its work are still more limited. Now, I don't doubt that like within the labs, they probably have some like pretty incredible internal tooling that will probably be available to us as a product sometime this year is my hope. But just having the ability to like launch like really great async agents, which no matter how much time they take, they have access to like a unique set of tools and can verify their own work, provides like a different level of like convenience to you, especially like I think in the context of our business, like having like in house legal operations. Yeah. Have you tried the slash command, like slash remote ENV on cloud code? Cause I assumed that's running on the cloud, right? Yeah. So if you use like in ampersand, I believe you can like push out your session to a close. like a cloud instance. ⁓ I think like this will work better for like companies that have like monorepos. So like if you have your back end and front end in a single repo, then you can like push that entire session out onto the web. If you're like us and you actually have to like simulate a monorepo in some way because you have a back end repository and a front end repository, that gets a little bit more challenging because you can't actually do like instruct claw to use like, hey, go into the front end. start that and like test this backhand change to confirm that behavior appears. Right. That I think is like really just a killer use case. And we're like working on making bad food so we can use some of these tools as well. I think like the more tools that we can use, the more options we can give people, the more that like I think everybody everybody will find something that fits like ergonomically into their workflow. Yeah. So on that topic, I guess I I want to get a little more detailed on like some of your workflow that you really like. So let's let's just talk, I guess, like the CLIs, because that's what I'm most familiar with. I'm not super familiar with Devin to be honest. But like do you use skills? And if so, like what skills have been actually useful for you? Yeah. So we write a few custom skills actually. So we write some for like code cleanup, some for code review. Those have like our own unique guidelines. We've also like experimented with like custom skills for like browser testing. So like you might write a skill that like explains how your part of the app works and how like a coding agent that has like another like browser use skill can also pick up that knowledge and be like, Great, all of this is loaded into memory. It doesn't need to spin up a pool of like 10 X4 sub agents to like go learn something only to like immediately compact after the fact. Yeah. So that's been a real convenience. I think in terms of like publicly available like plugins and tools that I really liked, Sentry is just fantastic. Like The level of like thought put into like their tools is really good. And so I've gotten a lot of utility out of that. I also really like the Notion one. Just being able to like pull and push from Notion, like as you wish, from like your coding sessions has been like really useful to me because sometimes I'll be able to like report up like findings on incidents to like non-technical people through Notion or even like document like product concerns that I'm seeing, like, hey, this looks a little awkward. Like I need these notes. I know people who like are really big fans of the like linear plugin as well. So like some of the again, like all of the classic, like workplace, like SaaS stuff that you would find, like I I see a lot of adoption there. Yeah. What about hooks? Do you guys use like pre and post tool use hooks in Cloud Code? We haven't invested as much into hooks. I know one coworker who like has like certain hooks that he likes for like running automatically running like linting, I think like after like every edit or something similar, if I remember that right. But a lot of the times like I just haven't found the value of like writing a hook versus like just putting like some set of instructions into like your main like Claude MD or agents MD and then having that done like on a per session basis versus like on a per tool call basis. Yeah. And what are your thoughts on the Claude.md and agents.md files? Cause I was just I think I I saw a tweet from like Theo where he was like, Yeah, they basically are useless. I kind of don't think they're like that useful, to be honest. Like I think that I think that there's like certainly areas where they work. So like code style is like one where like if you put it into your cloud MD, you're like generally going to get like Cloud following your directions on like, okay, don't use like inline imports or something similar like that, which Cloud loves doing. It's still a useful directive to put in your Cloud MD file. I think where people sort of lose the plot on the Clot ⁓ D is just putting like a ton of like business context or like application context into it. I think the natural behavior of agents in these harnesses is to like explore really to like read code, whether that's in line in the like main session or through sub agents or through an agent team, like these things like want to research and like want to touch code files. So like trying to like limit them or like try to provide more of that context up front like hasn't really proven to be a successful strategy with like Cloud agent MD files. Yeah. And what are your thoughts on like Git work trees and just like running parallel cloud code sessions? Do you guys do that a lot? Like how how do you keep all of that in context? Like it's working on like five different features at the same time? I really like so I really like agent teams. Yeah. For like, hey, like you can work on multiple aspects of a single feature. And if you have like Tmux, you like see all those individual sessions and like keep track of that. For like using work trees, I really, really like conductor. I hate like running all those like git commands to like go manage them. And like the advantage of this is it's like a desktop app that will just like spin up the work tree with like a cleanup, like with a startup and cleanup command that you give it and then just like run your cloud code or like codec session in that window. So that's been really useful for me. Again, I think I'm like more memory limited right now than I wanna be. So like the number of work trees I can run on my machine is like limited to I think like two full ones with like back ends and front ends. And so like I lean really heavily on like full async like agents to like really get the advantage of like parallelism there. But I'm optimistic that like we will solve that problem by like just like building some solution for it ourselves. Yeah. So are you now like basically not coding at your computer? Are you just like sending off messages in Slack like on your phone as you're out and about? I definitely do that. So if you like open my phone, there's like five tabs of like different like coding things, like different sessions that I'm keeping an eye on. But when I I love my focus time. Like I love the ability to sit there and like read through, like keep an eye on each of these sessions, like read through their like reasoning traces, like make sure that they're like going the direction I want. I think it's way too easy if you like kick off a bunch of like random coding agents without like strong enough directives for you to end up with nothing useful. And that that cannot happen. Like that is worse. The like little bit of convenience you gain by like not having to pay attention is like immediately lost by the output being like not understanding like what you wanted to do initially. And so I tend to be like a really active manager ⁓ of my like sort of swarms of agents and like keep an eye on each one, align them and like use like lots of plan files and just generally make sure that I know that like each of that those things that I set up like is going to accomplish the task. And sort of the time I recover, like while it is just like patching and making edits, I'm usually doing like more discovery. So I'm either like with users, like you know, sort of seeing like the issues that they might have in their current workflows, like sort of planning like the sort of next phases of my implementation, looking at strategic things. Like I feel like it just gives you time to go hustle, but like you definitely want to be very, very involved, like how they are actually set up to go like make those code edits. Yeah. Also, before you were talking about agent teams, like have you noticed agent teams work better than just like sub agents on their own? And also because you mentioned you use Codex over Cloud Code, I don't think those features are available in Codex yet, right? So Codex does have like an implementation of like subagents and you can like change the number of available like agents and like the number amount of like depth that they can have as well. I've played with that somewhat, but like my codex at the very least has access to a sort of quasi Agent like setup. What's funny in codecs, they'll like get like names for each sub agent too, which I think is cute. Interesting. I think the biggest thing for me is the reason I like agent teams more than sub agents is that I get more control over like what those individual things are like working on. So like if I'm like planning like a big like front end refactor, like I might say like use an agent team and I will lay out and like name each agent in like cloud code, like specified like the model that needs to be used, like for each of those agents. like give them directives and then like also just like be able to like have a little bit more control over how like each session is going versus like the earlier like subagents implementation, which is just like, hey, we have these like boilerplate things. Here you go. Bye. Yeah. I think that there are teams out there that do this with like fantastic taste. Like if you've seen AMP code, which I think like used to be part of like Source Graph, like Each of their sub agents like uses a different model with like a different set of instructions. And I think those guys have some have some really good ideas. Like I love like how they've set up like their Oracle sub agent. That's what like actually made me more bullish on like GFT 5.2 in general, is like they were like, they picked this for Oracle. That's a that's a good sense that that's gonna be a a good model for planning. And like, for example, like I think I referenced them before, but Factory's droid, like they just launched like mission control, which is their like implementation of this. And I've been playing with that to some like interesting results. Interesting. So that basically goes back to the point you were saying that you really like to actively manage the agent. So like agent teams basically allows you to do that. Like as they're thinking, you can kind of like guide them one way or another. You don't have to wait for the output of a subagent to start guiding the I guess like the orchestrator. Yeah, exactly. Like if you see like maybe like it's confusing like two modals that are very like similarly named, you can like address that like as it's happening versus like At the end, you're like, wait, these changes aren't where that I wanted them to be. Yeah. Interesting. So I want to shift the conversation a little bit. I know we're like running close on time, but I want to shift the conversation towards maybe the non-technical teams. I think that's really interesting, especially in your company's case, because that is a big part of your company. Like, what percentage would you say of your company as the engineering team versus the operations, legal, that sort of thing? I don't even think we mentioned like, What Finch does actually. Yeah, I think I think I can give like a very brief intro to what Finch does. Like we're basically a startup like focused on enabling Americans to get broader access to justice. So a big like part of like kind of our thesis is that there's more work to go around than can be like actually solved in any like unit of time that you like mentioned. And so we're starting with personal injury, working our way out from there. Yeah, that's very interesting. Cool. So I didn't have any other questions unless there's like something you think I missed that you wanted to go over. No, I think this was a fun chat. Thanks for having me on. Yeah. Yeah, for sure. Awesome.