Isaac: Welcome to Never Rewrite. I'm Isaac Askew. Jeffrey Sherman: And I'm Jeffrey Sherman, and today we're going to do part two of can Jeffrey Finish Off a 10-year-old rewrite using AI in two weeks? So the short answer is yes. Thank you all for listening. see you next Isaac: See you next time. Jeffrey Sherman: time. Yeah. no, so to rehash, the it was sort of a test and sort of something that needed to be done of The company had a V1 version of their API, which had grown organically without much sense. They had a V3 version, which was very restful, and because it was very resty, it was one, it was not very functional because all the intermediate joins and connection pieces relationships, all the intermediate relationships. Ended up being additional calls. So something that you could do in one call in the V1, you wouldn't have to do like five or six calls in the V3. And so human users hated it because it was so much extra work. And then there was a whole bunch of stuff that just wasn't done. The ever we talk about this many times with the half-finished rewrite, where you start a rewrite, you get some of the stuff out, you put it into production, and then there's this long tail. in this case ten years literally of well yes, we'll get to the rest of it when we get to it, but you never get to it, so then you end up supporting both versions in perpetuity. And so the challenge was can I, with a little bit of help and all the AI c I can use, get finish off this the not the f migration, just the rewrite. And I think that's an important caveat there of getting the V1 API is deprecated, but it's not in date they're not turning it off anytime soon. They're Isaac: Mm. Jeffrey Sherman: they're not forcing a migration and a lot of the usage is from from customers, so you couldn't put a you timeline anyway. But just getting it to the point where there is no reason they can't upgrade before it it's a necessary condition before you can go forward. Also the MCP server only works with the V3 API and other additional things that are just better about the V3 API. So you know before you can do anything else You need to finish the rewrite. Isaac: So how did that from where we stopped last time at the end of the episode? that was was it a full two weeks? Jeffrey Sherman: It was a full two weeks, yeah. Isaac: Okay. So how did you how did the from the moment we ended that episode and Jeffrey Sherman: Mm-hmm. Isaac: you went back to it, how did that continue from there? Jeffrey Sherman: So what I did is I had AI. And in this case, my tool of choice was Devon AI. that was the tool that I the first tool I got, we had did not have clawed licenses at the time. And Devin was entirely sufficient for my needs. And so that's the tool I used. And I certainly wasn't going to try like, well let's change your your main tool in the middle of a two week challenge. No, no, no. So ev when I say AI, in this case I was using Devon AI. and in the weeks since I've been playing around with Claude and I will say roughly similar in terms of product. Isaac: Interesting. Okay. I haven't used Devin, but everyone, you know, is head over heels about Claude, so I just assumed it was better. Jeffrey Sherman: Yes. Claude's got some great things, Devin's got some things. They're slightly different, you know, it's VS Code versus say IntelliJ. You might like one better than the other, but there's not fun there's no not much of a fundamental difference. So I had Devin create a list, like do the analysis. And this was, hey, here's the published API V1 docs, here's the published API V3 docs. Tell me the gaps. And it gave me a list of 75 things. And one fundamental mental shift that I made is stop treating this like machine output and start treating it the same that I would treat a list of 75. Items in a burn-down list from a human project manager. Because you know, in my 26 years of professional work, I have gotten many giant lists of burn-down punch list, you know, punch lists and burn-down items. And they are mostly right, but there's gonna be gaps that the things that are missed. There's gonna be stuff that is wrong, just totally wrong. There's gonna be stuff that's subtly wrong, and That's just how it is. And that's how and so the idea that I've been handed the machine has now done this vibe-coded thing. And I'm like, yeah, it's going to be mostly right. There's going to be stuff on there that's wrong. It's gonna miss things. And I will just treat it the same as I would treat input from another human, which is humans make mistakes. And I'm I'm not going to be annoyed at it like I would with a machine output of like, this thing is wrong. It's like, yeah. Isaac: Mm-hmm. Jeffrey Sherman: Alright, I'm gonna go item number one. and I went just because I have experience, I looked at the list, you know, I gave it a quick scan and I said, I'm going to start with what I think are going to be the most difficult things. Because when it gets when t when crunch time comes, I can have it I feel much more confident that I can set up missing filter options or missing sorting options in bulk. Isaac: Yeah. Jeffrey Sherman: than I can for reports with stuff in them. Isaac: This goes back to our other call out for why a lot of rewrites fail too, is they don't start with the hardest problem. They end up deferring it and to the very end and going, oof, this doesn't solve the thing and you like the half finished rewrite, 'cause turns out the hardest thing to do takes a lot more time and scope. So you inverted it and correctly worked on the hardest pieces. Jeffrey Sherman: Right. A a lot of the hardest pieces, it was like step one, do the language is PHP, both for the V one and the V three. And the V one, as you can imagine, being older PHP code, you know, if this V three's been out for ten years, V one's obviously gotta be older than that. It was tightly coupled to the state model, it was tightly coupled to globals, all the things that were normal in PHP 15 years ago that you would not want, it's there. And so I started having the age an agent iterate on the code, I'm like, okay, you know, extract all of the stateful things, the global stateful things. Extract those into variables and then make a a telescope of you know here's the original function and now all this original function is gonna do is extract all the global stateful stuff and then it's gonna call the new version of the function and it's gonna add those as parameters. So now I can have the v3 version. work entirely different in terms of parameter collection and calling, but still call I'm off screen, still call the the new version of the function because I've now extracted all of the stateful stuff. Right. So Dustin Rea: So you're purifying it, like PHP's version of purifying the function, you're removing side effects, removing state, and then you what you're left with is a pure function, same inputs equal same output, which means I can test it and guarantee it. Is that the stra the is that the strategy? Did I get that right? Jeffrey Sherman: Yes. for things that were I am part of the the the goal was not a pure one-to-one rewrite. It was a you know v1 was it's a hodgepodge of all kinds of things. V3, it's restful, it's JSON. And so but at the end of the day, if I'm migrating a report, then w I need that function that's going to generate the data. And what I need to do is extract all of the stateful stuff around it, all of the formatting stuff, so I can get to the data piece, and then I can reformat it. So so I can plug in the inputs as you say, make it pure, and then reformat it to the standard spec and go forward. And so that was the first round, few rounds, because that was the more delicate work, the the Isaac: Mm-hmm. Jeffrey Sherman: slower stuff of like, okay, well, I I can't finish this off until then, but I'm gonna do that. And so I started there and I I just kept working the system working the list of okay, what do I think is next? What do I think is next? And for each item on the list, sort of the first step was always to challenge the agent of like, okay, this is the list, this is what it says. Is this true? And now you know, then it would come up and it would still give me, you know, it it would explain these gaps to me and I would be like, That that doesn't make sense. I'm like which endpoint? 'Cause it would get confused. It'd be like, which end I would ask which endpoint are you fixing it this with this task? It's like, you know what, there isn't actually a V1 endpoint, but this V3 API is missing this logical piece, and so it doesn't make sense without it. And I'm like, Isaac: Yeah. Jeffrey Sherman: you are right. And let's make a new epic and we'll put that Jira on that put the ticket on that new epic and we'll get back to that later. But we're not doing it now. Isaac: Yep. Jeffrey Sherman: You know, the Which again is exactly what would happen if I had a human project manager. Isaac: Yep. Jeffrey Sherman: Like, you know, this isn't really part of the thing, but it it's necessary, so I just put it in. It's like, no, no, no, no, no, no. Right. And this this is one of those things that junior developers don't know. And as you evolve to a senior, this is the kind of thing that you learn is you can't trust any human. Isaac: We also talked about it as like the Boy Scout rule kind of thing. Like don't Jeffrey Sherman: Mm-hmm. Isaac: don't do the Boy Scout rule. Do not leave it cleaner than you found it. Like don't touch it, especially if it's legacy code, it's not like untested. Leave it be. Who knows what will Jeffrey Sherman: Yeah. Mm-hmm. Isaac: happen if you fix it and it suddenly starts working after ten years when it was broken for ten years. Jeffrey Sherman: Hehehe Isaac: You don't know. Dustin Rea: Can't tell you how many bugs I yeah. That's probably an entire episode of bugs I found after fixing a bug. That were way worse than the bug that I fixed, also. Jeffrey Sherman: Yes. Isaac: Yeah. So just let it stay broken. No reason. Mm-hmm. Jeffrey Sherman: So that was an interesting thing. while going through, I made a one endpoint hit and I'm like, all right, let me test it. So one thing that I had the agents do, because it it's a very data-rich environment, it was difficult to do like an end-to-end test, I would have it create a Python shim where it'd be like, all right, create this data and then I also, in addition to starting with anything that seemed hard, I also did the gets first. Right? Verify to me that if I have some data that is created, whatever, I can get the same ver values back in v1 and v3. They're going to be formatted totally different. And Isaac: Mm-hmm. Jeffrey Sherman: I would have the agent write a Python shim, like, all right, call this endpoint in v1, call this endpoint in v3, and compare all the fields, even though the data Isaac: Nice. Jeffrey Sherman: format's different, but compare them all and verify that this thing is the same. And so I would do that and I I ran it with the gets first because that you know once I knew that the gets were the same, then I can move on to just testing post and put of you know, I've created this thing in v1 and I post the same more or less the same data in v three and then I just read it off the gets and you d does it look the same? Because now I don't need to worry that the v1 and the v3 aren't the same. Sorry, that the gets aren't the same. So knock off a verb at a time. Isaac: Mm-hmm. Jeffrey Sherman: I I did that as well. again having it write Python for me, which would do the comparisons to make it make my unit testing in a staging environment, which with data rich environment easier. and that worked really well. That gave me a lot more confidence than simple unit tests and simple end to end tests, which are still extremely useful. They don't not knock them at all, but just for this kind of work where I need it To be much more coherent, it was very useful. The as I was going along, what I what I realized is hey, we've got this v1 API, which we've said is deprecated, but that's all we said about it. And we have the v3 API. And I started having the system add links. I'm like, hey, in this v1 API, and we want users to move off of it. Let's give them a link to the V3 API equivalent so that we can do it. And then this also became a a another way to check the burn down list. There were, I don't know, 50 endpoints in the V1 spec. If I could come up with a this is the way you do it in V3 and add a link, then I knew I was done. And so I would have it do that. And this led me to two discoveries. The first discovery, you'll never guess that the docs were out date. I know that Dustin Rea: Mm-hmm. Isaac: Mm-hmm. Excellent. Jeffrey Sherman: is shocking in this day and age. But the the dock, the v3 docs were wildly out of date. There was a lot of functionality that already existed, but it just wasn't documented for whatever reason. Because humans had to do it. So. Isaac: AI didn't catch that. Like sometimes I'll have AI catch the fact that things are they exist in the code and they weren't in the the list that was mentioned. Jeffrey Sherman: AI caught it when I asked it. Isaac: Mm. Jeffrey Sherman: So with that original list of 75 things, after I realized the first time that, wait, the code already does this, which it caught, I'm like, okay, go through, find all the stuff that already is on this list that you can close out by just simply updating the docs. And so I did a few passes of that of okay, let's let's update the docs and close out stuff. Isaac: How many did it catch? Jeffrey Sherman: I would say about 15 to like 20% of the total Isaac: wow, okay. Jeffrey Sherman: gaps were simply due to out-of-date documentation. So to me, that's as a takeaway, step one would always be true to the docs. Which is an intelligent and rational thing to do anyway. But dealing with humans and the amount of labor involved in chewing the docs, nobody does it. Isaac: Mm-hmm. Jeffrey Sherman: Nobody did it. Isaac: No they do it. Jeffrey Sherman: Well, no, they still don't because it doesn't occur to people. But now you Isaac: Yeah. Jeffrey Sherman: should. Dustin Rea: C can you go into a bit a bit more about how you like validated at the end? Like you kind of talked about using, you know, writing shims and I I imagine you had some like kind of maybe some custom test harness or harnesses and fixtures. But like how do you go about, you know, one, knowing that you didn't break anything or uncover, you know, issues that were laying around. And then how do you like maintain continuity with the teams that, you know, are using those old endpoints? It sounds like it's just an opt in. Update, so it's up to them. Jeffrey Sherman: Right. In this case it was an opt-in update and I specifically framed the project that way where in terms of the challenge, I did not need to get anyone to migrate. So everything well almost everything was new and additive. And so I did not need to get anyone to migrate. Now I do, like now, outside the challenge, I am working on getting people to migrate. But in terms of a two-week challenge, if you have to rely Isaac: Mm-hmm. Jeffrey Sherman: on other people and then getting them to migrate. the out the window. Dustin Rea: Confident are you that there won't be like production incidents or production issues from the migration itself? Like, is that like what's your confidence level? Jeffrey Sherman: That's a hard question because you're saying the migration issues. I am fairly confident that this giant blob of work that I've done, in and of itself, is not going to create production issues. Migrating from a V1 endpoint to a V3 endpoint and All of the weirdness that somebody might do in an integration has a much higher chance of creating an issue. so but going back to some of your other things, I absolutely found bugs. I found bugs that were more than 15 years old. And then I had to fix them. Cause I'm like, well and again, this is another reason Isaac and I have always been so anti rewrites is hey, you did I I'm doing this rewrite but there's this version that we don't care about and there's this new version. But hey, I've now found a bug, so I gotta fix it here so I don't have to recreate it there or else you know, recreate it or or what have have you. Dustin Rea: That's another one of those junior mistakes that you make before you're experienced. You fix the bug along the way and then now you have no verification of like if there's a new bug, you don't know if it's because you've tried to fix that bug or if it's because your implementation. You just, Jeffrey Sherman: Mm-hmm. Dustin Rea: you know, back to just fixing bugs again. Then you loop and loop, you get you know, go down a rabbit hole. Jeffrey Sherman: Right, and this is also a problem with larger MRs and slower cycle time with the more whip you have. and that became a Isaac: Mm-hmm. Jeffrey Sherman: problem that I encountered where I did around a hundred pull requests, MRs, to to knock out these 75 things, but due to the company's release cadence. The initial release got they have a weekly release cadence and this is a two week challenge. And the first release got screwed up and they basically punted and so n ninety of my changes went out in one burst. Which am I which killed a lot of the value. Yeah, it's a lot of surface area. Dustin Rea: A lot of surface area. Isaac: Why don't you fix the bug that you found? Why don't you just leave it be? Mm-hmm. Jeffrey Sherman: Why did I? Isaac: Why don't you just port it one for one to V three and then go through after? Like just defer it in the same way you deferred the other epic stuff. Any other issues you found? Jeffrey Sherman: Because it was so small that it was faster for me, especially with AI, to fix it. Like it was a one liner. Isaac: Yeah, fair enough. Jeffrey Sherman: And and so it's like alright, by the time I put it to the side I could just have the AI fix it. If it was something subtle Isaac: riding the jury ticket to keep track of it was takes longer than just fixing it, yeah. Jeffrey Sherman: Yeah. Like if it was something subtle, like the reports that I ported, which I did a one-for-one, if I had found a bug in there, I probably wouldn't have touched it. Right? For those, I only touched the parameter validation and the output formatting. And if there'd be if I'd found a bug in there, I would have pushed it to the back. I would have said that no no, I'm not fixing it. and and I did have an issue where I found Isaac: Got it. Jeffrey Sherman: an endpoint. That hadn't worked because we had removed a library back in 2023. And so this endpoint had not worked in three years. Like it was impossible. It could not have been working. Isaac: Mm mm. Jeffrey Sherman: It was like an XML serializer and wasn't in the system anymore. And I'm like, well. So I put that on the list of like, I'm just not gonna do it. I'm not gonna fix this thing so that I can port it. But clearly it hasn't worked in three years. So clearly this feature's not in use. And we should think about. what it is and do we actually want it and if we do I'll build it. But not gonna fix it for the sake of fixing it by readding the library when clearly no one's using it. Isaac: I could see a case where like even leaving it broken, like like let's say it returned like a a four a four four or something like that. And then you go, let's just delete it while we're here. Let's just clean it up. And then somebody had some weird code that handled a four four, but when you deleted it it was a five hundred now and now it like kills somebody's integration, even though it was broken the whole time, it was broken in the way they could handle it before. That's just another one of those weird Jeffrey Sherman: Mm-hmm. Isaac: reasons where you just don't touch it. Just leave it be. Jeffrey Sherman: Yes. Isaac: There's always a way to piss off a customer. All right, so what's you're done with it? Jeffrey Sherman: well, I'm done, but I'm not done done. by which I mean Isaac: Mm. Jeffrey Sherman: the last of the changes haven't been released. So I'm still waiting on that long cycle time to get the code out into production. And then when the f last of the code hits production, I can Isaac: Mm-hmm. Jeffrey Sherman: update the last of the V3 docs. and those go pretty much instantly. But the V1 docs have another w cycle time of a week and it's a different cadence from the code release. So the last of the codes should hit production this week, and then I can update the V3 docs, which means that I can update the V1 docs in production next week. Isaac: Got it. For Jeffrey Sherman: And this is you know Isaac: for your V one and V three w I know the response looks correct between the two, like with as far as the fields that you wanted to return from a get or whatnot. are they served by the same underlying functions? Like could there be a bug, for example, that one of them returns the correct fields but the other bug returns the correct fields in a very non-performant way? Such that now that you switch to V three, there's an issue because of a bad runaway query, even if right now at a at a low level, pinging it one by one, they look correct? Or do you have them both using the same, kind of serving the same data? Are they refactored just different Jeffrey Sherman: That is absolutely a potential issue. For some, it's using the same functions, and so it's the same data, and there's no new risk. For others, it's a different code path. It should be a better code path, but there's no guarantee because no one's been using it this exact way. Isaac: And as you roll it out, you'll probably catch it too. Like once you migrate a couple of people, then you see, there's a spike here that you can catch it then too. Yeah, I mean you d optimally you Jeffrey Sherman: Right. Isaac: don't want to catch it live, but you know, there's a way to catch it early. Jeffrey Sherman: Well, this is where if I had you know, if if if there had been more, you know, a continuous continuous release kind of system, then it would be less risky just because there'd be less surface area for any individual endpoint that I'm Isaac: Right, right. Jeffrey Sherman: migrating. But in this case, it's not terrible because it's their additive and people aren't in using them immediately. Because, you know, it launches it you put an endpoint out, but it's dark, nobody knows it's there. Then you update the V3 Isaac: Mm-hmm. Jeffrey Sherman: docs. And then you update the V1 docs. And before at the before times of like last year, it would take a very long time for humans to notice and do anything about it. Now, well like one of the reasons that I wanted to include the links was AI will scan the docs and build integrations. And people do that all the time. And by putting the v3 links in the v1 documentation, now I'm telling people, telling the AIs more than the people to use it. Like, hey, use this instead. And so it should get hit much faster. Or, you know, the the lag should be much lower. Isaac: Cool. You got me thinking a little bit too about like the how fast this was. Like if this was a ten year problem and in two weeks AI can do a pretty pretty good job at getting it mostly there. Imagine another couple of years or even sooner that it does it in two days and then eventually two hours. At some point with that timeline shrinking, just theoretically, I'm wondering if This even even the concept of the whole thesis shi thesia shipping, iterative delivery is just an implied obvious path. And no one has this concept of the big bang rewrite anymore. That was like the scoped of the a year-long thing. Now it's like, yeah, when we say rewrite, of course, we mean iterative delivery a hundred thousand times in, you know, ten minutes to get this thing rewritten. And now we have another kind of super definition of what a rewrite is. And people do rewrites all the time because it just encompasses like the obvious safe flow of this testing and iterative release cycle. I wonder if that's like a future way of Yeah. Jeffrey Sherman: It very well could be. Yeah. Dustin Rea: Just doing it at a much faster speed or like much many more cycles, yeah. Jeffrey Sherman: Right. If if AI you're if you have an agent doing the same steps but it doing them much, much faster, then I think that is safe. It's where AI takes a giant Isaac: Right. Jeffrey Sherman: jump, the same giant jump that humans try to take, that's where it becomes not safe. But it as the harnesses get ever better, I I think it will tend towards the better path. Isaac: Yeah, yeah, the thing that looks like the the risky rewrite is now actually all the things that we worried about has already been considered in that small amount of time. I imagine we will get there. And then maybe maybe rewrites actually maybe we do change to always rewrite because by that point in three years, rewrite always means iterative delivery in a very quick, safe, AI driven way. Who knows? Could be the future. Jeffrey Sherman: don't don't change the name of the show, Isaac. Isaac: We'll have to rewrite the name of the show. Well, Theseus ship it. Jeffrey Sherman: yeah. Dustin Rea: Ha ha ha ha. Jeffrey Sherman: AI rewrites. Yeah. But yeah, the the big things to Isaac: It's sell better for sure. Jeffrey Sherman: me, the big takeaways are if you treat your like a bur well, first of all, this is a scoped, like it was a contained problem of V1 versus V3 published things. Which made because of that it bounded the problem in a way that made it attackable. Like, is this done? Yes or no. Isaac: Mm-hmm. Jeffrey Sherman: I was able to use y Yes AI hallucinates, it hallucinated a bunch, but because I could bound it and check it and just had the innate distrust that if a human had done the work, I I was okay with it. And then finally I I was able to just verify it, you work ways to verify it against itself. Isaac: Mm. Jeffrey Sherman: And so. It worked. Isaac: Congrats. We can let us know after it's done done. And then part Jeffrey Sherman: Yes. Isaac: three. Probably not enough content for part Dustin Rea: Yeah, your migration risk. Isaac: three. Yeah, I mean I mean i yeah, if something crazy happens, Jeffrey Sherman: Yeah, migration risk. Isaac: that'll be a right episode. What was missed? Jeffrey Sherman: Yeah, I I could see i in leaderships we we 'cause we had a meeting and afterward the leadership person was like, this is great, it's so it's done. It's like everyone's migrating. I'm like, No, no, no. It's like if Dustin Rea: Mm-hmm. Jeffrey Sherman: you if you remember the proposal, no migration. And I could just see Isaac: Yeah. Jeffrey Sherman: in his eyes, he's like, this was a necessary but not sufficient thing. And I don't have everything I wanted, I have everything I asked for. Isaac: Yep. Jeffrey Sherman: Yep. Which again, the disappointment of a rewrite. Yeah. Isaac: Well I hope I hope the migrations go well. Right. Yeah. There's so many different ways it can go wrong. and i just expectations versus reality is just one small piece of it. Not even the technological part, but like you could do everything right technologically and have a misunderstanding of what was delivered and then still you didn't do you Jeffrey Sherman: Yes. Isaac: like the product people are mad or this you know, CTO is mad. I thought Jeffrey Sherman: Right. Isaac: you were supposed to deliver this. Jeffrey Sherman: Right. You could get off the V1 API now. That doesn't mean you have. Isaac: Also the customer outreach to get everybody migrated and some teams I'm sure that they you have to have a lot of hand holding. Not every company, as a sp like you serve a lot of small businesses, not all of them have an IT department that can sit there and go, yeah, I can go ahead and correct all my routes to point to the new V three. There's probably a lot of customer support. Like that's like that's the whole next level. Jeffrey Sherman: Mm-hmm. Isaac: So that's a big migration. All right, shall Jeffrey Sherman: All right. Isaac: we leave it there? Jeffrey Sherman: I think we should. Thank you all for listening. I'm Jeffrey Sherman. Dustin Rea: And I am Dustin Ray. Isaac: Isaac Askew, and this is Never Rewrite.