Isaac Askew: Welcome to Never Rewrite. I'm Isaac Askew. Jeffrey Sherman: And I'm Jeffrey Sherman, and today we're going to talk about having AI monitor your release process. So if you're working at a large-ish company or a large distributed system kind of a thing, you're gonna have releases, it's gonna take a some amount of time. And if you're like most people, you don't have the best metrics around what's going on in the release. especially if you've got an older, messier system where There are errors in your error log continuously, and now you're looking to see: well, did the rate of errors in my error log change as I'm releasing? And so one thing that I have started doing, and that we'd like to talk to you about today, is the idea of, hey, look, you can have an agent check the error logs, of simply say, hey, we're doing a release, watch the error logs, and see if anything new starts showing up. Because that would be very quick indication of, there's something bad in the release. or if Isaac Askew: Right. Jeffrey Sherman: you could do the con inverse of, you know, it see if anything disappears that isn't listed as a thing that we fixed. Because that that would also Isaac Askew: Right. Jeffrey Sherman: be strange, right? You what you don't want as surprise, I guess is the where we're going, is and this is historically difficult. if you don't have a very clean system, it's very difficult to go from If you've got say a hundred errors per minute, it's way too many for a human to Isaac Askew: Yeah. Jeffrey Sherman: pay attention and notice trends on. and it's it's just hard to know what you're looking for. But AI does a really good job on this because it's just a purely statistical thing of well, what do you see? And how many of these things should I knock down? And it can grind it out pretty quickly. And You can then just see, okay, well, it ha you know, are there any new errors that we're hitting in production that we missed in our testing? Isaac Askew: Yeah. I like this idea. I think like I've used Sentry before to like check for exceptions, but there could be errors that are not exceptions or n they don't make it to the the try catch kind of area that you'd normally report, like capture exception if you've manually done like your own configured catch block. and there's th th a lot of times i I'll monitor like I'll have sentry and like ground cover, which is logging system. I don't know if you've used that one, like CloudWatch essentially. Jeffrey Sherman: Then not okay. Isaac Askew: so I'll I'll check both whenever I deploy something and I'm like keeping an eye on the sentry channel, make sure there's any there's no new exceptions, and I'll keep an eye on the ground cover logs to see if there's I'll just kind of refresh and see like the last five minutes is there anything odd here and I usually see like an error whenever there's like a deploy and the session dies and reboots or whatever. Jeffrey Sherman: Mm-hmm. Isaac Askew: Something like that and that's in the obvious like this is this is fine. But then I'll just look and to inspect especially like any of the the files that I edited. You know, if they're a very common like this like a web the webhooks repo I was working on recently is a good example of this. If there's a bunch of new webhooks that I've added or a bunch of e webhooks webhook codes webhooks code that I've edited recently that, you know, gets hit frequently. I can check and filter on that particular file Jeffrey Sherman: Mm-hmm. Isaac Askew: to make sure it's still behaving the same. And I've actually manually downloaded those logs and gave it to AI and said, Is there anything, you know different here or can we use this to see what files are not being hit. Like I use this to clean up dead code to show Jeffrey Sherman: Mm-hmm. Isaac Askew: here's the access logs the last ninety days. Tell me which endpoints we never hit 'cause if we haven't hit them in ninety days, it's probably dead code. And it found all of for me. But the idea of what Jeffrey Sherman: Nice. Isaac Askew: you're talking about here is like almost like an N C P connector to CloudWatch or to ground cover that it could feed it right into it. Yeah. Jeffrey Sherman: Yeah, you'd need MCP or some kind of real-time or nearly real-time thing, because if you've got an a a a release and it gets usually what you do is you roll it out in stages of like, okay, five percent, ten percent. And if you've got good coverage, then great. And if not, you're you're really hoping that some customer is gonna squeal. Isaac Askew: Mm-hmm. Jeffrey Sherman: that and tell you i in the five percent range as opposed to not seeing it until you get to the tent. and that happens, but like for that to happen, something has to be terribly wrong. You're not gonna notice, you know, a a low level thing that's gonna happen ten times an hour. But it's still it's there. It's gonna be an error and you know, the dorometrics of just how long does it take for it between when an error gets into production and you notice it and then fix it, the clock has started w with your release. Isaac Askew: Yeah. Yeah, and it could not even just be like small things that they wouldn't notice, but it could be a big thing that takes a long time to notice that actually has been dead on deploy but you didn't know that and the customer doesn't know that and then a couple of days later, then it's it's very apparent this process that might Jeffrey Sherman: Mm-hmm. Isaac Askew: run nightly or whatever didn't run or runs every couple of nights. You know, so those are the cases where having some logging or having something in the background diffing. how your current your your prior process and your rather your processes prior to deploy and after deploy just to see if there's any kind of anomalies. and like you said too the the inverse matters too, like if if suddenly all the air logs go away and you're and Jeffrey Sherman: Yeah. Isaac Askew: you're used to a hundred air logs a day, you're like, that means the reporting's broken or something you know, the logger broke. You know. Yeah. Jeffrey Sherman: Right, well that would be more obvious, but yes, if you if you expect a hundred errors a day and suddenly it drops to ninety and you didn't fix any of the thing that would lead you to believe that there should be ten gone, Isaac Askew: Yeah. Jeffrey Sherman: you're like, well I wonder if the Isaac Askew: Yeah, like an upstream process could have finally fixed its payload such that the downstream Jeffrey Sherman: Mm-hmm. Isaac Askew: one accepted the payload in one particular case. And you know, and then AI can like, you fixed you know, John fixed so and so three days ago with this commit on a different repo. That's why you're seeing the reduction. You know, it can make those connections a lot easier, especially with the complexity of like many systems talking to each other than than Jeffrey Sherman: Mm-hmm. Isaac Askew: we can digging through each. Jeffrey Sherman: Yeah, it's not human doable until you miss out on just the wide range, the the long tail of stuff. You you just miss out. Isaac Askew: Yeah, I think that's a good idea. I think that would be like something I'd like to implement. because like I said, right I I was manually searching for it and then I was downloading the Jeffrey Sherman: Mm-hmm. Isaac Askew: JSON and giving it anyway, that's the perfect perfect automatable thing to have it like a post deploy process to wait it can already health check and wait to see if the deploy happened and like wait as a process. And then after that like Jeffrey Sherman: Mm-hmm. Isaac Askew: start looking at the logs and seeing am I seeing the same rate of two hundreds, four hundreds Whatever. Jeffrey Sherman: Yeah, and if you've got a deploy script, which I hope you do if it's more than you know, just click a button on say GitLab or something, like if you've got a long phased out deploy, then there should be a deploy script with checkpoints. Should be relatively simple to set up an agent to like, okay, we've now released it on ten percent of our machines, ten percent of our instances. Now tell the agent to go look and see, you know. What do you see in this version that you don't see in that version, and vice versa? And tell us, and that report is now a useful guideline of do we go forward? Isaac Askew: Yeah. I think this might be why a lot of companies like Datadog and Snowflake and what is the other one that's like data centric, similar to Datadog. New Relic maybe, or Grafana. Jeffrey Sherman: I don't know. New relic yeah. Isaac Askew: i they've been experiencing a big boom lately, Jeffrey Sherman: Mm-hmm. Isaac Askew: when it comes to AI 'cause I I I suspect it's having all this logging and Telemetry and other things that are really useful for understanding how your system's behaving as extra context to feed that humans can't just quickly re- same thing for like very quickly if if my if I try to get my Docker network to come up and something happens. The error Jeffrey Sherman: Mm-hmm. Isaac Askew: might not be very obvious, but it I can say, Claude, can you check and see why it didn't work? And it'll like look through all system logs faster than I could type it and go, I can't see it's because of this particular issue here. Here's Jeffrey Sherman: Mm-hmm. Isaac Askew: what we can do. You have to you need to remake the network because of so and so or last time I did this it was like a compatibility issue with Mongo. Mongo got upgraded. so yeah, catching things like that, I think giving it more context is I I would be surprised if we didn't see companies go this route. If if the if this doesn't already exist as an idea somewhere out there. Jeffrey Sherman: Right, so that's something that like New Relic and Datadog did very well even before AI, but it was very expensive. And here Isaac Askew: Mm-hmm. Yeah. Jeffrey Sherman: it could be as simple in the setup that I'm using, it's as simple as Claude plus Grafana MCP server. so Grafana runs on top of Loki. So it's all open telemetry, and you just ask it, hey, connect to the MCP server, check it out. Boom, no giant setup. And it it works. And you can set it up it you know in just a few minutes and it will help you with your releases and you will get more confidence, you'll catch bugs faster and everything else. And I suggest you give it a try. Isaac Askew: I think I'll set this up at a company I'm working with, and then report back to you for for a follow up on this one for next time. Jeffrey Sherman: Okay. I I will hope maybe have it more fleshed out. It's a good simple thing. thank you all for listening. I'm Jeffrey Sherman. Isaac Askew: I'm Isaac Askew, and this is Never Rewrite.