Isaac: Welcome to Never Rewrite. I'm Isaac Askew. Jeffrey Sherman: And I'm Jeffrey Sherman, and today we're going to explore whether AI and AI development changes software fundamentals around work in progress and PRMR merge size. Isaac: Okay. Yeah, it sounded like before the show you had kind of a story about this related to testing at your current company. Jeffrey Sherman: Right. Without going into specifics, or let's go into specifics, but also not. Isaac: Mm-hmm. Jeffrey Sherman: I'm working with a code base that is organically grown PHP from 2008, 2009. So Isaac: Mm-hmm. Jeffrey Sherman: if you can imagine, or you don't have to imagine. PHP has early PHP had a well-deserved reputation. It's a horrible language that lets you do horrible things. Isaac: Mm-hmm. yeah. Jeffrey Sherman: Th there's a wonderful meme, if you care to look it up, about if you had to ham you need to hammer something and you have a shoe and a coke bott a glass coke bottle, and what should you use? And it this is an analogy for early PHP. if you had to use some of these two things to hammer a nail. And This code base has grown organically. So it you know it's a real code base, it's really been running. and the developers Isaac: Mm-hmm. Jeffrey Sherman: have been writing tet unit tests in this thing for like 17 years. And Isaac: Ooh. Jeffrey Sherman: because of the way the early system worked, and the because of the way PHP worked, especially the early PHP. where you had globals were an extremely big thing. Like it would be hard to have an early PHP system that was not full of globals, session globals, Isaac: Mm-hmm. Jeffrey Sherman: global globals. and also, you know, back in the prehistory of why they did it this way, they did it this way, it also has a persistent database that they use across unit tests. And so Isaac: Yeah. Jeffrey Sherman: Any test that didn't take a lot of effort to make sure it reset the global the persistent state to back to the original starting state was quite likely to get tripped up on something else that some other system needed to set or unset. And so this system it ran in a certain order, and if you ran it in a certain order, everything was safe. And if you ran each test individually, it was safe. But you couldn't run some of the tests and have any confidence that a failure was actually a failure. Isaac: What was their rationale for the the unit test slash end to end test bleed? 'Cause usually with the the unit tests you want to isolate and mock like database calls and you're just testing functionality and not like the response or whatnot. It sounded like they wanted to kinda hack the unit test to be end to end test in this case to make sure it modified state correctly. Jeffrey Sherman: Correct. These were more end-to-end tests. the rationale is that's what they did. there was no rationale. Isaac: That's a good rationale. Jeffrey Sherman: it it just or there may have been, right? Again, this is a 17-year-old code base. I have not been at the company for that long that I it it was there Isaac: Yeah. Jeffrey Sherman: and Things have inertia, so because the code base was set up this way, this is how testing was done. People just wrote more tests. Isaac: Right. Jeffrey Sherman: Right? You could go back and invent an entirely new testing setup, testing harness, which we eventually did. the because of the performance issues, we had what they were called slow tests, which are highly stateful end-to-end, more end-to-end tests than unit, and then we had fast tests, which were true unit tests. Isaac: I love that characterization. Yeah. Jeffrey Sherman: Everything was marked. And y the it ran. Isaac: Okay. Jeffrey Sherman: so yet we did eventually, and by eventually I mean like five years ago, split out the end to the stateful end to end tests from the original tests. But because both exists, there are thousands of these stateful end to end tests that masquerade these unit tests, and developers continue to add to them because it's easy and it's better than not having them, right? Isaac: Fair enough. Jeffrey Sherman: Yeah. And so I asked myself, or my a colleague and I asked ourselves, could we fix this with AI? and the impetus for this was the we have build pipelines and the build pipelines. The the spread between the fastest group of tests and the slowest group of tests was like three and a half minutes. But we couldn't. shrink the the group. We could we couldn't reorder the tests because reordering Isaac: Mm-hmm. Jeffrey Sherman: the test wasn't safe. It would everything would blow up. So as an example, I I I naively not naively. I had it just take I I had a AI say, hey, re redo the sharding so that the tests are packed for for time. So limit the time spread between any of the shards, any of the groups. Isaac: I can imagine it didn't do well for that first pass, depending on how you prompted it. Yeah. Jeffrey Sherman: Right. Well there were thirty-two shards, then they all passed, right? 'Cause all the tests pass if you run in the right order, and they all pass if Isaac: Mm-hmm. Jeffrey Sherman: you run them individually. Thirty-two shards, I had it repack it for time, and all thirty-two failed. And so we had AI where it was like, okay, well, here's all these failures. Look at these failures. These are all due to leaked state. and so it was okay, have all of these tests reset the state that they need on setup, and then try to have it clean up the test the state at the end. Right? So don't leak more. And so 30 or 40 merges later, we Got them all, it was more like 50. Got them all working, Isaac: Mm. Jeffrey Sherman: and I said, Okay, cool. Are we really done though? Shuffle it up again. And that time, like 18 of the suites failed. Isaac: Nice. Yeah, okay. Jeffrey Sherman: And so we w we continually did this. And after about a hundred round or after about a hundred MRs, each with one or two test files in each MR. I said, okay, maybe maybe we're going about this the wrong way. Instead of trying to fix each test individually, let's have it do a blanket thing. Right. And I should say I had it trying to do fix all these things individually because I was going at it in a breakfix manner of this is broken, fix it. and do lots of small things and then merge it in. Isaac: Yeah. Jeffrey Sherman: I'm like, okay, well, what if I tried to have the AI be proactive? I said, okay, look. Analyze these things. Like here, you know, this is PHP. Look at the session object, look at the global object, look at these things, look at these database things, the the dependency injector. These are the stateful things that exist in this test suite. Look at all the tests, find things that get changed and don't get changed back, and then come up with a list, and then let's change all the you know, set up all the teardowns to fix. All of those things, and then it'll be great. Isaac: Yeah. Jeffrey Sherman: Mm. Isaac: I was gonna say I'm just kinda so the main the main problem here is just the the fact that the the suite takes a long time to to run and Jeffrey Sherman: Mm-hmm. Isaac: changing the order of these tests to optimize for the time is risky to do. So the process here is essentially just trying to train the AI to most effectively order these tests for you. So your your initial prompt was This some something along the lines of we know these things are stateful based tests and Jeffrey Sherman: Mm-hmm. Isaac: that when we change them, they'll break. Organize them. And so it did a decent first pass after well, not first, but fifty passes, MR passes. After that, it was able to make the test suite faster, essentially. It was able to to make them all every test pass and the suite was faster. But you were like, Is this a hardy solution? Are we gonna run into this problem again? Is that the main Re reason why you kept going. Jeffrey Sherman: Correct, you're right. We basically we had a system where yes, we had this this test, they could run them in this order, and then if we shuffled them, we could make the the build process like twenty-five-thirty percent faster on every build simply by changing the order the tests were run. and so we changed the order and then we fixed all the things that were broken. But as soon as we started adding new tests, this problem was likely to crop up. And this has been an ongoing Isaac: Yeah. Jeffrey Sherman: problem with this code base for as long as I've worked there of you added some tests, some of these integration tests, and now suddenly this other test that is totally unrelated doesn't work anymore because you've now because your new test shifted the ordering and now this test doesn't work. And when you fix that test, it can cause Isaac: Mm-hmm. Jeffrey Sherman: a new cascade of other failures. And so this is a perennial developer problem. Right. Not only is this Yeah. Isaac: So I got a I got a question then for for your as you started talking to the AI and trying to come up with solution some solutions for this. Were you did you get curious about like maybe I can just randomize every every time the test suite runs, like tell it I want you to run the test suite five times before you say it's good. And every time you run it, I want you to run it in a random order. Like like order all the tests and then just shuffle them every time or something like that to kind of give it a a test or something to play with. Jeffrey Sherman: Yes, so we set up a second branch, right? So we set up the first branch was we want it this to be as fast as possible. We want the build to be as fast as possible. Isaac: Okay. Jeffrey Sherman: We set up a second branch that was we want you to run the tests in as random of an order as possible. Isaac: Nice. Okay. Jeffrey Sherman: We never had to run that second branch twice to get a new bunch of tests that were failing. Every time we ran Isaac: Oof. Okay, yeah. Jeffrey Sherman: it ran it randomly, we got failures, and so we would fix those failures. We didn't need to run it five times and like just. Isaac: Did you have like AI run it and like in like expand like its thinking process and like did you see that it was going, I've run into this failure again after only the second pass and like see how it like it how it tried to to change its own behavior to run the the to change how it's running the tests? 'Cause sometimes I've noticed like I can see w I can see what it's thinking if you expand the window depending on the model Jeffrey Sherman: Yeah. Isaac: that you're using. And then so if you if you only had to run it twice, the the chore is almost to give it to it and be like, Run it five times and come back to me if all five ran successfully. And if not, find out why. And then just kinda like watch it and see what it's thinking and why like I why it like usually usually you'll see one area where it's like, again I failed the task. Let me see if I can change this so that, you know, something else happens and you just kinda sit there and watch it. Have you tried that? Jeffrey Sherman: No. So that's where I was going with for this thing and the best practices, right? Because we could have done that. Isaac: Mm-hmm. Jeffrey Sherman: We could have said, hey, AI, right? Here's a it's it's totally random. It's gonna shake the, you know, it's gonna shake the snow globe, and things are gonna fail. And I want you to shake the snow globe. Then fix all the failures, then shake the snow globe, then fix all the failures, and keep going until you can shake the snow globe like five, like you said, five times. but what I realized, or and again, this is my decision as a human developer, is when I had it come up with a list of all the things that we would need to change, after we were at least a hundred merges in. And and I l it I said, okay, you know, come up with a list of all the things that you would need you would need to change to be fairly confident that you were not leaking state. And it was more than four hundred things. Isaac: More than four hundred Jeffrey Sherman: Four hundred t test classes. Isaac: Test classes, okay. Is that a bad thing? Jeffrey Sherman: Right, so they no, but the it's a big thing. But if I did the snow globe thing, but I said shake this thing, Isaac: It's a big thing. Jeffrey Sherman: fix the tests, I would be left at the end, you know, ignore how whether or not that was an efficient use of my tokens, I would be left at the end with 400 changed files and a green build. And I would have no Isaac: Okay. Jeffrey Sherman: idea, like I I guess I could merge it. No. I would have no idea about what had changed. Isaac: And now we're back to the other problem from a past episode, which is, you know, when you really let AI build everything for you, and and review everything for you, at some point you might not know how any of this is even working anymore. And so if you want to keep understanding how things are working, in your case here, you have to go in those this is what you're getting at, I guess, is this the smaller pieces. So you can't approve four hundred tests at once. Even if it looks green because you don't know if some of those are false positives or if they've changed some kind of thing to make it pass 'cause it gets clever 'cause it wants it to pass for you to, you know. Jeffrey Sherman: Right, yeah, it it could pass. things that we've noticed when we were working with the you know, doing it one off at a time. Again, this is an organic PHP code base. There Isaac: Mm-hmm. Jeffrey Sherman: were four flavors of database operations. Data like, you know, there's the very low level, here's PDO, do the database Isaac: All level. Jeffrey Sherman: thing, and then there was a homegrown ORM and there's Laravel. Isaac: man. Jeffrey Sherman: Sorry, there's two, three, three variants of homegrown ORMs plus Laravel. And it would get confused. It would start writing PDO, or it would start using Laravel in sections Isaac: Yeah. Jeffrey Sherman: of the code that don't use Laravel. And that would work, right? There's nothing fundamentally wrong with okay, I'm gonna use PDO Isaac: Mm-hmm. Jeffrey Sherman: and have it write raw SQL to reset state. Isaac: Yeah. Jeffrey Sherman: But like that's not idiomatically correct, and you're just making the whole thing slop. Isaac: That honestly is another good call out too, is that that that confuses human and AI. When you when you when you do unconventional things, which tend to happen if you built a library before another library got famous and big, or like Jeffrey Sherman: Mm-hmm. Isaac: I one company I worked at, they had built their own test framework. And it was after PHP Unit, this is PHP, it was after PHP unit existed, but they made their own Jeffrey Sherman: Hmm. Isaac: assert true, assert false, assert equals Their whole library. They had I c have no idea why they built their own custom one. But anybody looking at that thinks, this is just using it it was the exact same format as PHP unit test suites. And I'm just like, this is not even using that. This is and I had to fix a bug in their test suite for like equals operators. Right? Jeffrey Sherman: Bye. Isaac: So when you when you let AI loose on that too, it also gets confused because it makes the same assumption that humans do that, this is using a very popular framework and it might not be. And then separately. Another issue we saw was like i if there's if you end up using Jeffrey Sherman: Well I did see Isaac: go ahead. Jeffrey Sherman: yeah, like it it did it's did for one of them it made its own comparison thing. I'm like I looked at the MR, I'm like why why is you not just using assert? Isaac: Right. It'll do anything it can to make it pass. So you have to be careful. but yeah, using conventions I think is is even more important these days for the sake of AI because you'll confuse it. And one other thing before we get back to to the test topic, I just I happen to know this notice this recently is there was this really weird flow at a company I'm working at where we would mark something as like a bill as paid, and we would say paid is true, like a bullion. And we had this other flow where we wanted something to be not showing in our main portal, but showing for partners. And so we'd mark it as paid to hide it from our portal, but then have a not paid, another not paid column, and then another like I forget this other column we used to show it in the partners panel so they could continue operating on a bill that looked like it was paid from our end. So whenever AI hits it and it was like, no, this bill's already paid, we're like, no, that's just not. It's actually elevated to the partner technically, but it's still not actually paid. It's just paid true does not mean it's truly paid. Yeah. Yes. And AI was confused by this. I mean, Jeffrey Sherman: Right. This this is very enterprisey. You have a Boolean and now you need another one. So you add a second Boolean, now you've got four states, but two of them are invalid. Yes. Isaac: it because we are too. Anybody any junior engineer coming in and looking at how you've set it up goes, wait, what? 'Cause it doesn't really make sense. If paid true doesn't mean it's actually truly paid, that's confusing. And so AI will run with too. So any Back to the the testing conversation, th those conventions are super important. And I think maybe, and this is what I've been trying to do recently too, is clean up around language that we're be we're using at our companies and especially Jeffrey Sherman: Mm-hmm. Isaac: cleanup on dead files that provide context to AI that it does not need. That might confuse it, is becoming more important for us too. Jeffrey Sherman: Yes. so yeah, to bring it back to the question that we were talking about to begin, my partner, my colleague and I started this project and we're now we've merged two hundred changes to the test suites. Isaac: Mm-hmm. Jeffrey Sherman: And we think we're getting close to being done. But the main thing th one of the things that's happened is we have merged two hundred times into the main branch. Like we don't have this giant 200, 300 long file whip out there, you know, with all this theoretical value. Isaac: Work in progress, yeah. Jeffrey Sherman: Yeah, work in progress. Everything we've done immediately got merged back in. And so we know Isaac: Mm-hmm. Good. Jeffrey Sherman: that none of these things broke anything. we also, because we kept the humans in the loop, we know that the changes are reasonable. But also importantly, like one thing that we noticed, because when we went back to look, like, okay, we're getting close to the end, let's let's make sure we have numbers to talk about, right? Because like, okay, there's a three and a half minute spread between the fastest and the slowest. One thing that we learned when we went back to look at the numbers is, that spread has shrunk to ninety seconds. It's only a minute and a half now because People developers have continuing to add tests. We've been re doing things to cause this re So even on the main branch where we have not act intentionally caused the reordering, reordering has happened, but there was no cascade of side effects, which normally happens when that happens. Isaac: Yeah. Jeffrey Sherman: And so even though the project is not done, and even though we haven't gotten the goal that we hit weren't going for, which is to make the build faster, we've already, because we kept merging everything as we went along, we've had this benefit, this silent benefit of, and Isaac: Yeah. Jeffrey Sherman: we've quieted down the the ink the random disasters of flaky tests showing up. Isaac: Nice. So what's the what's the end goal at this point? Like at what point are would you call your project finished? Because imagine like there's only so much you're gonna get diminishing returns at some point. Jeffrey Sherman: Right, yeah, we We have not done the snow globe, although maybe at this point we should do the snow globe because now it's small enough. But our current thinking is when we get the current round done so that the build is green again, we'll go, we'll get the 90 seconds, we'll declare victory, and maybe Isaac: Mission accomplished. Jeffrey Sherman: mission accomplished, maybe maybe we will set up something in the background to snow globe. and see if we can get it to snow globe but also you know fix this test. I guess there's no reason it still wouldn't have to because I would still set it up. I would have it snow globe, right? Here's shake it, fix these things, and then when you fix the test, cherry pick, you know, create a new branch, cherry pick that and build and tag me with a new thing so we can merge that change back in. Isaac: So I'm gonna ask an an evasive question. Is is there a possibility that the tests are just written in a dumb way? Or that the unit tests that we've made into end to end tests may be because you you said people kind of took that and ran with it and then just made their own tests. I assume, maybe wrongly, that somebody decided to do this database state test on something that's actually quite simple. And didn't even need a database. They just kinda did it because that was the pattern everyone was doing it for. Is there any way to clean up just the tests styles themselves so you don't run into the state problem as often? And maybe extract some of the happy path tests to make sure you got decent coverage, but you don't need it for everything? Or has that already been considered? Jeffrey Sherman: I I already considered it and I already did that. So a few months ago, an earlier version, before I realized that there was such a large spread between the slow test suites themselves, I said, Well what if I could extract some of the slow tests that are not actually slow, that aren't, you know, stateful, and move them over to the fast. Because the fast test suite runs in like thirty seconds. It's fine. Isaac: Mm-hmm. Jeffrey Sherman: and so I did. I got one or two percent of the tests and again I used AI. I had just had to move it over, boop, from slow to fast, and then I had to fix some stateful cascades because I removed tests, which caused the orders Isaac: Mm-hmm. Jeffrey Sherman: to shift. so yes I did that. I I move anything that didn't actually need state has been moved from one side to the other, which also got you got a small prof it it It took a few seconds off the build time because it no longer was setting up state and trying to reset the state and build you know doing all that stateful work for cut tests that didn't need the state at all. Isaac: Okay. Evasive question number two. Are there any tests that you found that you could have just plain out deleted? Like they were actually already encompassed by another test that did the same kind or similar work that already proved what a smaller test was doing that you could have deleted. Jeffrey Sherman: We didn't ask the question. I thought about it, but I didn't want to. I don't have enough confidence that I would trust the AI's judgment on it, and I didn't want to spend the time proving it to myself. And so sort of letting the sleeping Isaac: Fair enough. Yeah, okay. Jeffrey Sherman: dog lie. And and and I feel like I would have more confidence if it was a fresh system, as opposed to one that has grown organically for almost Isaac: Yeah. Jeffrey Sherman: two decades. With you know, Isaac: Got it. Jeffrey Sherman: many styles and many things. Isaac: Yeah, I think the the way you're you approached it is kind of the way I approach it. It's like essentially I call it l the lazy approach. At first you're like, Do we have to do this? and after we have to we we have to do this, okay, well do we have to do all of it? You know, and then just kind of keep breaking apart there and then when you're finally left to I case actually we have to solve this problem and there's not a way to simplify it any further than that. All right, well let's get to work and then from there going through and finding ways to chick the snow globe or improve the process and the bootstrapping, whatever we need until we we get the we get the problem solved. So I dig it. And and the smaller d delivering in smaller chunks too, especially. Because I know at the you know company I'm at today, a a four hundred test change is not something people would just blindly approve yet. People are getting too hasty with AI. but usually that's like this is too Jeffrey Sherman: I I would not approve it. Isaac: much. This is too much to change. We don't know how it works. And then you'll go through like you're talking about and you'll find, you just cheat wrote your own definition of assert equals. And that's how you got it to pass. Come on, like you know, and AI still it'll it'll still cheat and get clever. And so having some kind of double check to make sure it's not doing that or that your test is actually doing what you need it to be doing is still important. And it's still something of AI is not really reliable on yet. It's getting better, but it's still not reliable. Jeffrey Sherman: Right, and there there is something to be said for, hey, it's just a test, it doesn't really matter if it's but if it's not actually testing what you think it is, then it matters. Isaac: Yeah. Yep. Jeffrey Sherman: Awesome. Well we actually ran long. thank you all Isaac: All right. Jeffrey Sherman: for listening. I'm Jeffrey Sherman. Isaac: And I'm Isaac Askew, and this is Never Rewrite.