Olaf: So it's all about this balancing act. So how much money are we willing to spend on trying this new functionality out? Because you cannot predict any everything in advance, right? You need to get it in the hand of people with the right amount of quality to get the right data back as soon as possible, right? That's the big challenge. Sebastian: Welcome to another episode of our new season of Beyond Vibecoding, partnering with ImpalaSearch. André Neubauer: go to TEC, an executive search agency in Germany. Sebastian: In this podcast, we explore the transformational change in software engineering and knowledge work in general. I am Sebastian Heidemeyer zu Erpen, CTO at Ecosia. André Neubauer: And I'm Andre CTPO at Trusted Jobs. Great to have you back. In the past episodes, I actually have learned that technical experience, including strong CI, CD as an important asset on the way mastering agentic engineering. Sebastian: Yes. And in today's episode we speak to Olaf Moldenvold of Circle CI. And CircleCI obviously knows a thing or two about CI CD pipelines. And Olaf shares a fascinating vision on how signals from production feedback directly into product development, maybe the dark factory, and the CI CD pipeline. Surprisingly obviously plays an important role in that, right? anyway, it was a fascinating discussion and We hope you enjoy the discussion as much as we. Welcome to the show, Olaf. It's great to have you. And as Olaf: Thank you. Sebastian: usual. we'd like you to introduce yourself to our listeners. Olaf: Yeah, thank you Sebastian. Happy to be here. And of of course also André. my name is Olaf Moleveld. as you can hear from my accent, probably I'm from the Netherlands. I live in Utrecht, which is a nice little city in the middle of it. And I work as a technology advisor in the CTO offices of Circle CI. So I work very closely with our CTO Rob Suber. And I tend to kind of work between product strategy, product marketing. thought leadership, try to get the story and the brand out there and also kind of feedback this information into what we're doing at Circle CI. My background is software development. So when I started in the end of the 90s, the World Wide Web started happening. So I was kind of like run surfing the first wave, doing all the stuff that e-commerce, content management systems, very cool. And then at some point, as you also mentioned, Sebastian, I started moving to the dark side, doing less coding, more consultancy, architectural things. and at some point I was very heavily involved as a consultant in the content management and e-commerce platforms, and then Docker containers came along, and we started this re seeing this requirement from the business side of things where they started to do A B testing. canary releasing without doing monolithical kind of releases and then I started a company with a good friend of mine called FAMP.io. actually it was an e commerce platform first and then we started zooming in on the progressive delivery thing. so we did that for about seven or eight years and then we got acquired by CircleCI and that's happened I think about five years ago. And here we are. Sebastian: Awesome. Thanks a lot. André Neubauer: Awesome. Sebastian: So also you've seen a lot and you've seen yeah several like changes in the industry, right? That's also a pattern we're seeing people who've seen a couple of changes it's easier for them to basically also approach the new change that we're seeing right now. and yeah, it's so it's it's really great to have you and I'm curious to learn how you at Circle CI see this change and and what you think about it. But before we get there, as usual, we like to go into the status quo of our interview guests. And you are obviously less on the on the coding side, right? But I'm sure you are using agentic tools, workflows for the non-coding side, right? So how are you utilizing agents? And if you have an example then Please let us know. Olaf: Yeah, so maybe I can kind of slice it up between the different ways of working during the day or the week. So within CircleCI we we use a lot of entropic tooling, like cloth and so I think one of the cool things they were seeing besides obviously engineers using it to create code, to review code, to look at the PRs and basically help them As a kind of like a superpower intern or bunch of interns to do their work faster and at a higher quality level. We also create a lot of skills in-house. And I think that's one of the cool things that I'm seeing where non-engineers, maybe for one cool example for is that we have this design department that does all the branding that we have, and they created a skill in Claude. To create branded slides. It's really like the and you know what a hassle it is to get your slide deck in the right brand because the brand changes, you know, colors change. And now it's basically like, okay, Claude, this is my outlines, these are my slides. can you create a slide deck for me in the in the proper branding? And then it starts to do all this stuff, and it actually works really nicely, which is I think a testament to the fact that Agents are not only for coding or for your code related things, but they are super powerful ways of helping you to do the I would say non-differentiating heavy lifting, but having the right branding when you do a public presentation is obviously important, right? So it really helps me to focus on the content of a slide deck and not so much on if the color is you know, the hex code of the title is André Neubauer: Yeah. Olaf: good or not. So I think that's a that's super nice. From a personal perspective, I use it a lot for brainstorming. So we are a distributed André Neubauer: Okay. Olaf: company. Of course we have offices in in Silicon Valley and and globally all over the world so there's colleagues you can work with. But in my in my life I I work a lot from home and I I'm a person that likes to or needs to talk. And rehash my thinking in a focal way because it helps me to kind of refine my thinking. And typically I would kind of like abuse a colleague and and and put up a whiteboard and then start talking and then rehashing it. But now I do a lot of this thinking and brainstorming with AI agents. Like this is my ID, can you be very critical? Or you are now a hacker news commander. So you're super critical and negative and shoot holes in my in my idea, right? So that's that I think that's really nice. it it it requires a little bit of a tuning to make sure that the AI is not the this is a great idea. I'm not sure if you have the South Park episode, André Neubauer: Yeah, sure. Olaf: this is a great idea, let's do it. I I I want it to be very critical because that's helpful to me. And then the third thing, I don't code for for money, let's say it. anymore. I code for for fun. So I do a lot of like these personal citizen coding projects, right? I have an LG television. how cool would it be to have a slide deck in there that shows all the art of the Rijksmuseum Amsterdam? It's not there, but there's the API. And I don't have a don't have a clue how to build this thing for an LG television. Figure it out. And then boom, a few hours later I have my own personal app running on the television. That shows all the artwork from the Rijksmuseum. And it's only for me. It's not public. So yeah, th these are kind of like different ways I use it. Sebastian: Absolutely. And so I I do also the the third thing. And sometimes for such a project, vibe coding is perfectly fine, right? So you don't need to have the Olaf: Yeah. Sebastian: best setup around this. Just you vibe it and if it works, great. If it doesn't, yeah. So what you tried it and yeah. And and maybe Exactly. Exactly. Olaf: Yeah. It's kind of like throw away software almost, right? Yeah. Sebastian: Yeah. Maybe, maybe one one question on the on this branded slides is the output. put then like, I don't know, a PowerPoint deck or something that you can edit? Or is this like yeah, I don't know, is it always that you would always edit the inputs and it would always then generate the slide deck with I don't know d fully generated images and and ready made Olaf: Yeah, that's an interesting one. So so effectively a few days ago I was in London and I did a talk at at AgentCon and I had my kind of like rough outlines and my slides. And I went into Claude because I know the skill is there, the design skill. So I said, like, okay, use this skill to to brand the slides into the proper branding. Turned out there is actually a a a slide deck editor and presenter mode. in in in in claude and within the cloud projects I didn't even know. So all of a sudden I was in this editing thing and I could click on present and so that it's in in the environment and obviously the AI companies want you to kind of like be in their bubble and and offer as much as features as possible to make it sticky and for revenue. But it was nice. I could do all the editing there. And then this question popped up, obviously, like how do I get it from here into Google Slides? So at some point I changed from Keynote to Google Slides. I don't even know why, but it happened. So I was like, okay, I need to get this stuff in in Google Slides. What to do? And then say, Yeah, I can export it to PowerPoint PPTX. And but what it did, it exported all the images, so it was static. I don't like static. I want to kind of change it five minutes before presentation. I don't know, living on the edge. And so and then it said like okay, I can also make like an editable PowerPoint for you. So that's basically how it goes from from Claude into Google Slides. one of the things that I think this skill needs to add is we have this theme also for for Google Slides, a branding theme that we used in the past. And it didn't include the theme, so it's now black and white, right? It's either the the thing that uses the skill or you d go into Google Slides and then you use the theme. I I'm not sure how you can kind of combine these things yet. So there's room for improvement, but it w it's I mean it's nice, it works. André Neubauer: Yeah, that's one of the things we always realize in our talks, This is so we think just because we can't imagine or we don't know, Claude maybe can't do this or any other kind of LLM. But just asking, it will figure out a way. It will figure out a way. Olaf: Yeah. Yeah. Yeah. Yeah. Yeah. It's so yeah, th there's all kinds of angles and yeah, I s I still think you cannot outsource intelligence or or thinking. It is still for me it's like a super powerful, potent way of automating things with the right guidelines. André Neubauer: Yeah. Olaf: But yeah. André Neubauer: Maybe one follow up question on your tech stack or tool stack. Are you also spending time on that to better understand how to position Circus CI and the products you're offering? is there also Olaf: Yeah. Yeah. André Neubauer: such an intent? Olaf: Yeah. I I tend to play around with everything to get my kind of sticky fingers on on on all the new technology. I think that's part of my job to kind of feel around and talk to people and play with it to just get a feeling for it. and this needs to yeah In my head it starts to kind of bubble and then I at some point I I see this thing or directions or patterns popping up and also yeah, you see my gray hair, not the people listening to audio not, but at some point you start seeing these patterns over and over again. And what I really think is now interesting and it's super relevant to what we do with Circle CI also, you see this kind of expansion of the different VCS systems, right? I mean Circle CI from the origin is VCS independent. So we're we were one or the only ones that basically doesn't include a VCS. So you bring your own VCS and we can kind of integrate with it. But now this is kind of exploding. GitHub has outages so now and then let's let's call it this way. So you see like things like cursor origin popping up as an alternative. We have a feature now that for integration with with origin, but also with Bitbucket, GitHub, GitLab, etc. And now, as we're European, you see this adoption of Gitea or Git Git, I think. Sebastian: Kitty, yeah. Olaf: GitT I I will say Gitea, but I don't know why. or Forge Joe, you know, you see those kind of sovereign oriented alternatives for VCS. So we also start adding those to to the mix. And I think that's that's a very interesting development happening where people start to spread their risks, don't want all their eggs in one basket. So that's super relevant to what we do because it it's already in there, but now it becomes more a broader offering of VCS systems that companies want to integrate with or even combine all these things. Like maybe some projects are super sensitive, they are there. Other ones are less l less relevant, they are in a different VCS or maybe a shadow VCS, right? So that becomes a like a multi-layer thing. And also what is super fascinating to me, a good friend of mine is sales. So he's not not technical at all. And all of a sudden he's like, Hey Olaf, I made this app for LinkedIn so it can schedule my LinkedIn post. Super cool. And then I said, Yeah, that's nice. And I said, Yeah, and now I want to have it kind of like running all the time. So I went to my IT department and they told me like, Yeah, it needs to be in bit bucket and is it containerized? and he said And I'm like, what do you mean you don't know? You wrote this app. He said, like, I don't know. I said, I I asked him, like, is it in GitHub or is it or Git? You know, where do you store the code? He said, I don't know. And then I I asked and and he started looking into what he did. And it turned out that that I think he was using Claude or Copilot. And it told him to create a local Git repo on his on his laptop. So it was in Git. On his machine, but he didn't know he hadn't any clue what is Git. It isn't you know, so you see this this new way of of working where people just do or let the agent do it and then don't André Neubauer: Hmm. Olaf: have a clue, and it that becomes like a friction point because there's the classical IT department or operational department say, Yeah, it needs to be containerized, and we need it in this VCS, and then these people, the five coders, let's call them, they're like And then you see this kind of gap. And for Circle CI, obviously, this is an interesting development because how can we help people to kind of evolve from having code on your local machine, like the old days, right? I'm this solo developer and I develop on my own, nobody else knows how I'm building or testing or compiling, and all of a sudden somebody else needs to take over, or we kind of grow the team. How to get this from a single machine into like a team thing where you can look at the dashboard, right? Yeah. Sebastian: Exactly. Yeah. And and and that's that's actually now almost at our meat already, right? Because we're very, very close to to the main section, the topic-wise, which today is like shift right guardrails, canary deployments for synthetic code, right? So we what we've established and and I would now do the segue actually into the meet because we're already like quite advanced, right? So in the interest André Neubauer: Do it. Sebastian: of time, it and we what we've established in previous episodes is that like technical excellence is the most important driver for agentic engineering or successful implementation and and excellence there and obviously c icd is like at the core of this right and circle c i is at the core of this so you probably have a lot of insights here right so yeah please let us know how how you see this and how you at circle c i approach this topic. Olaf: Yeah. I I I I guess I need to shift a little bit now from from technical to more business because they work together, right? And as I see it is that writing code in any way or form or writing software is a means to an end. It's not a goal by itself. Right? Why do we write software? Because we want to automate things, we might like We want to make life easier for people. So if you work backwards, how do you measure excellence? Why do people need to be excellent in writing code? That's to get the best possible quality of automation or functionality in the hands of end users. When is a software developer happy? Okay, you can love the code or the regular expressions. That's when you approach it more from it's an art form, or you know, it's it's really but André Neubauer: Yep. Craftmanship. Olaf: craftsmanship, yeah, that's the term. Very cool, thank you. but when I think you should be really happy when an end user says tells you like you make my life insanely better now. This software is really helping me to do my job better or my life is is improving because of what you developed. So the value is defined by the end user. They create, they they define if it's valuable to them or not by paying money or using it more and more. And then if you work your way back, I think you should always measure in how you write software and how you test it and how you develop it and deploy and release, where's the value in what we're doing? And it's very much like the lean principles of back in the day, right? You want to minimize waste and you want to maximize value. And I think that's that's where that's kind of like the next wave of of what needs to be done and what's possible because writing code now is become a commodity. Almost right. You let you let an agent do it. But how can you get from this step to getting the maximum value in the hands of end users without spending too much for it? So it's all about right sizing, I feel, finding the right models, the right size models, the right type of models, maybe like local models, maybe cloud-based models, frontier models or very small models for the task at hand. compressing all the stuff that you send to the models, right? Maybe you they only need a small piece of JSON structure instead of like a total dump of your database. so kind of measuring and balancing this. FinOps is also very, I think, interesting party. FinOps was also like a always a separate thing where people would kind of look at cloud cost spent, but now Adding FinOps to your pipelines, your CICD pipelines, and kind of testing if this AI related feature or this AI evaluation step isn't isn't that too expensive for what we try to achieve value wise, right? Automating this process in the pipeline that makes a ton of sense. It's super interesting. And these tools are there. And of course the canary releasing or you know at the end of the pipeline where you can say, Okay, now we have this this artifact, this deployable feature. How can we measure if this thing is actually providing value for end users? You don't push it to everybody in a one go, right? Because that's that costs compute and time and resources and and and patience. Give it, I don't know, to a few percent of end users somewhere and start measuring how they use it. And then based on that information that you collect, you can feed this back into your pipeline and say, okay, this doesn't seem to work because the the value metrics, the business metrics that we want to observe are not changing. And obviously you need to have a multi-layered approach. So it's not only business metrics because only business metrics doesn't tell you the the entire picture. FinOps metrics are also relevant, I think, in these days. So what's the cost versus revenue kind of balance? We want to keep that in the in the green zone. And then you of course have the technical things like the classical SRE kind of metrics, HCDP, memory, CPU, error rates, etc. If you layer these things together, the CI C D pipeline or the platform can basically get these things out in an automated way and then you can go into your canary leasing and and pipe that spec. So I think it's super interesting how we can kind of transform from do this, do this, do this, kind of like a very stupid or basic sequence of steps into a really in a in a smart pipeline that constantly is measuring metrics and then feeding it back. into the agents. André Neubauer: I what you actually, maybe how I understood it is actually also a widening of the loop, right? Because at the moment, the loop is more or less, it ends on your machine. Not all the time, and Sebastian will disagree in a second, I know. But in regard to the example you gave beforehand, right? So people are very much aware of... coding, so they try to solve that, but they're not aware of what's beyond that. How do I get this now into production? And what you're actually describing is seeing this as an entire loop, right? Because if you implement it, you build it, you test it, you deploy it, and then feed the results back, that is actually a large, large Olaf: Yeah. André Neubauer: loop. Olaf: A I mean w with the agents, you saw this Ralphie loop thing, right? It André Neubauer: Yeah, yeah. Olaf: it it's essentially what what this is, but then not on the local development scale. Like make this work and if it doesn't work, do try it again and iterate. It is like the the entire software development lifecycle loop, including deployment and what we call release, which is the thing that comes after deployments basically opening up the door To visitors or users to start using it and measuring things. Of course, the challenge is to align these things, right? So historically you had these silos, developers would focus on technical metrics, FinOps would would focus on financial objectives, and then operations or business would focus on analytics, right? Are we sticky? How is the shopping cart value increasing, right? But now, if you want to automate this and and then constantly feed it back into the loop, you need to agree on aligning these metrics and then getting the right formula on top of them and then feeding it back into your agents or maybe humans, the but it's data, right? André Neubauer: And if we, if we say we, we shift, right, I think shifting right to production is not a replacement for proper testing. Let me put it that way. It's, yeah. Yeah. Extension is Olaf: No, no. It's a it's an extension. Yeah, yeah. André Neubauer: a nice word here. Yeah. Olaf: Yeah, but and and again, talking about value optimization, if you don't do proper testing, then you will most definitely fail in production, right? Which will be costly. Maybe not from a technical perspective, but if your if your customers are experiencing crappy functionality or suboptimal UX. That's also costing you money because your brand value will decrease, right? So it's all about this balancing act. So how much money are we willing to spend on trying this new functionality out? Because you cannot predict any everything in advance, right? You need to get it in the hand of people with the right amount of quality to get the right data back as soon as possible, right? That's the big challenge. and and and I think we're with with agents now with AI, we can kind of reduce the the silo friction because humans needing to a align on things is always kind of tricky. And but if you automate it in formulas and you let the agents do all these steps, then it I think becomes easier to to get this really working and up and running. André Neubauer: Yeah, that would be actually also my follow up question. Then where would you see human in the loop here? Because ideally nowhere, right? Olaf: Yeah, like I said, you cannot you cannot outsource intelligence, and you cannot outsource creative thinking, at least that's my opinion. so I see this change, this shift where these software development lifecycle pipelines become more like a factory, like a software development factory, and the André Neubauer: I like that word. Yeah. Olaf: and and we the humans. We become more the designers or the directors that basically tell the system what we want to achieve and let the factory André Neubauer: Absolutely. Olaf: fill in how we want to achieve it. So I'm I would be very curious because in like when we started FAMP, there was not even Kubernetes, and then we did a lot of things with HAProxy and Nginx, like like the proxy routers. I think if you tell an agent like, okay, I want to do this. blue green or canary release thing create a pipeline for me and set up a terraform infrastructure thing to to make it happen I wouldn't be surprised that they they are actually able to achieve it these days. I would say okay you can use Argo rollouts or Circle CI deploy or you know and this is your by the pipeline definition here. Good luck it will work. And in the past it would be more like Okay, we need the people from this department, we need the people from that department, SREs, ops, you know, and then you all need to be in the same room and align and agree. but yeah, this this it's much less friction, I think, now. Yeah. Sebastian: So it it it almost sounds like there there was this concept and Andre, I think you were like really in interested in this or like found it amazing. The post hoc, right? You remember the the agent that reacts to André Neubauer: Yeah, yeah, yeah. Sebastian: product signals, right? So this sounds like this, a little bit advanced, added FinOps data as well, right? that's that's really interesting. And then if we close this loop, so how would the the signal that that come back be fed back into the pretty much software factory. Then would there like issues being created in the issue tracker? Olaf: Yeah, I guess the I think issue tracking or issues i is almost becoming an audit log, right? It's like okay, what what what actually did these agents do and why did they they do this? So I can take a look. of course, for maybe critical things, you want a human in the loop, like a gate, and sign off on it. Like if you're like a regulated industry, I guess you're required to do that that way. But just with Ralphie loops or that that way of working, if if these things are becoming really like a black box and go really quick, why do you need to be in the loop, right? I if the thing is not going off the rails and the guardrails are around it and and you cap it right on on on budget spent or token spent and and and you let it report every once in a while. I wouldn't know why you would be involved in that feedback. Because what's the difference in telling a user like Okay, we did this five percent rollout of this new feature. It didn't look good because we saw a lot increase in five hundred errors. So and then create a ticket like we need to do a rollback. What's the use of that? You can automate it, right? You can say okay, roll back this thing and then try again. and and of course this trying again that might be where a human becomes super valuable because why doesn't it work? Why are you Sebastian: Yes. Olaf: s and then you need to get the input from the end user? Sebastian: I would I would agree, but this is then the final vision, right? Because I mean, let's face it, the software factory is not here yet, right? So this this full automation flow is not here yet. And for every step in in the software factory, you need to establish trust and confidence in the system, right? So that you can step out as a human. And the same thing would then be the next step. Once the released software is in the wild, it you feed back information. I think the first thing would then be to have the agent analyze the information and summarize it and derive action items from it, right? And then you Olaf: Mm. Sebastian: as a human are still in control and saying, okay, yes, this is the right action item. But fully agree, like, should we be at a situation, or I I think we will be, where we say, I don't know, 10 times out of 10 or 100 times out of a hundred, yes, this is the right action, or maybe even just ninety-eight. times out of a hundred. Yes, this is the Olaf: Yeah. Sebastian: right action. Then we can probably say, okay, happily stepping out if the guardrails are all right. yeah. Olaf: Yeah, as long as you can kind of have an audit log and you can see what happens. It it's like this like the like a senior working with a media or a junior or with an intern. First you kind of coach, you you guide and and like you say, you you you need to build trust and kind of and trust is a two way street, right? So also people need to give the right input for for a person to to do the right thing. Otherwise it's also not working. It's and the moment this trust starts happening more and more and yeah then then I I think you can lean back a little bit more and and and see this thing happening. So I was thinking about it because if you do like electronic designs you have all this simulation right you can you can basically nobody is doing an electronic piece of equipment André Neubauer: Yeah. Olaf: without drawing the thing and then simulating how it works. I think a first Very cool step would be to run a simulation of this fully automated SDLC and not doing it for real, but seeing how it actually would create releases and then measure input from from from users and feed that back. See how it kind of works or if it really goes off the rails depending on on parameters that you you give it. I I think the simulation thing will become maybe maybe even an an industry or or or a a quadrant because if you automate things with agents you want to make sure it works as you design the the process right and before you productionize it but that's a side sidestep Sebastian: Yeah. Yeah. No, but but it it makes sense and I've I've been actually I I've been working at companies and and also like am working at a company where at least production traffic, either like shadow like shadow traffic or like logs of production Olaf: Yeah, yeah. Replay. Sebastian: traffic is being used for simulations. so that that makes a lot of sense and yes, you could this way you could pretty much compare metrics from from the reality with the simulated run, right? And then see how it behaves. And yeah, absolutely. It's just a matter of the size of the system and the the cost associated with having this shadow run. but in I think if you can scope it to the right level of experiment because you're not always changing the whole system, right? So usually Olaf: No. Sebastian: you're working on on one thing or a few things at a time and then you can if you manage to Have interfaces that only like launch instances that need certain data and as an input and then you view the output, right? Then you can probably simulate a lot there. but go going back to your initial thesis, I'm assuming that circle CI is also not there yet, right? Sneaky question. Olaf: No. No, no. I mean I think we we we made great strides in in automating a lot of what we do. but yeah, I mean a CICD platform is is infrastructure, it's plumbing, it's it's crucial. You don't want your your electricity in your house or your Wi Fi to become flaky. I mean that that's what you see with GitHub these days, right? It's it's becoming flaky. And you don't want flaky infrastructure or flaky plumbing in your house. So we we stay very much focused on reliability, stability, performance. So yeah, there's there's there's definitely always humans in the loop to make sure that that what we release doesn't break things or affect So y yeah, it's it's a balancing act, right? To leverage the power of AI without kind of crashing, fa failing fast kind of, you know? André Neubauer: Yeah. But you're sitting in an interesting position, right? Like, so the vision you just laid out, I think it could be a good vision for Circus AI, right? So it's Olaf: Yeah, and and we and we try to drink our own champagne and André Neubauer: Hahaha Olaf: and we use circle CI to build Circle CI. So I mean yeah, so we we are kind of forcing ourselves to to do these to take these steps. But obviously if you have a greenfield application that that doesn't run at enterprises with high compliance around it, it's easier to to experiment. So André Neubauer: Absolutely. Olaf: yeah, so for us it's also very powerful and very valuable to Talk with our users how they are using our platform and how they are leveraging and what they kind of wanna do with it in the future and how they kind of see new features that they wanna use. So yeah, it's it's a li again, a two way street. Sebastian: Absolutely. Yeah. thanks. I think this is a great end to the main segment of our podcast, right? And as usual, we have some smaller segments towards the end. so the reality check. We in the reality check, we usually ask our guests if they have what the fuck moments or also what the latest hype topics for. their or from their perspective are. So do you have some examples there? Olaf: Yeah, yeah, definitely. one example that I vividly remember is at some point I needed to archive I think like a hundred or a a big bunch of repos. So I thought like okay, how shall I approach this? okay, I'll I'll ask the the the chatbot to write me a little Python script or batch script. To archive this with GitHub repos. So it said, like, yeah, yeah, I can do this for you. Here's the little best script. Do you want to do like a dry run? Yeah, sounds good. Do you want to have an audit log? Yeah, sounds good. Do you want to feed it with you know with the with the IDs of your repos? Yeah, yeah. So all these things it it it asked, it sounded very solid. I was like, okay, good. And then it created this this script. And I looked at it, obviously. Does it look good? I want to understand what it does. Yeah, looks good. Okay, let me first do the dry run. And then immediately it it it didn't crash, obviously, but it's a three-one error. And now it's like, I don't know, it looks good. So I started debugging it. Hmm, I don't know, looks very it's no syntax thing, right? And then I started diving in and then And it turned out it was using this CLI, GitHub CLI, and then it uses a flag or a command archive. Turned out there's no archive flag, is it's just not there. And because it looks so obvious, right? You you overlook André Neubauer: Yeah. Olaf: it very quickly. And then I was like, What? Why does it propose this thing? And I said, like, there's no archive command in GitHub CLI. And you're totally right, it doesn't exist, but it would be nice to have I was like, yeah, it would be very nice, but it's not there. And then I I I tried to drill in a little bit more, like, why does it come up with this thing? It turned out there was one person at some point that make it made an issue. Said like it would be really handy to have this archive thing. And I guess it's being trained on the data of that issue. And I basically used it as a hypothetical way of working. It's another way of looking at things like What is in our C L I missing? But yeah, there was a a what the fuck moment. I was like, crap. So it's I took me like an hour to figure out, basically. Sebastian: That's a fascinating example. Also, the fact that there was someone who has been asking for it and then this was probably part of the training data. Yeah, that that's Olaf: Yeah. Sebastian: really interesting. Usually I I attribute this to hallucinations, basically, right? To the the Olaf: Yeah. Sebastian: the LM just thinking it should be there because it sounds logical. interesting. Yeah. Yeah, and did do you also have a wow moment or something where you say, Okay, this is like really crazy good? If not, then that's also fine because we all got so used to working with agents, right? So it's hard to be amazed. Olaf: Yeah yeah so I as a hobby I I also create electronic music, so I have a little studio with all kinds of devices. And I just bought this Aki sampler and I was missing my old Roland synthesizer sounds and there was this guy that made a like a like emul emulation of this thing, and I thought like okay, I wanna basically get all these samples out of this thing, like I don't know how many the thousands into my I kai new sampler. So I basically let the AI agent create a a little plugin or or application and it it was in C or C, I think, and it used this VST plugin library thingy, and then it would do auto sampling through all the the sounds, and then it would create a patch for the AI sampler, and it worked. I mean I never Sebastian: Amazing. Olaf: wrote in C and it and André Neubauer: Yeah. Olaf: it and it used the library I don't I'm not sure what the name again was, but it uses this this like low level library to fire off the sounds. It is awesome. Yeah. André Neubauer: Yeah, it's like, like within five minutes, you now explain both sides, right? And isn't it fascinating, right? Like, so Olaf: Yeah. André Neubauer: situation where it really failed and where you thought, well, this is straightforward. And then where it was so capable. Yeah. Olaf: Yeah. Yeah. Sebastian: But but I think it it's also a good example of you're probably much more well-versed on utilizing git commands or whatever, right? And and interacting with repos that you think, okay, yeah, this is obviously like I I could look the syntax myself, but no, it's easier if I have the agent do it. Versus on the other hand, you would never probably think about creating a VSTI because it's just something that you have no clue of, right? And that's why. Olaf: Yeah, yeah. Sebastian: It's so easy basically for the agent to impress you, right? Even though I have to admit, yes, this is totally impressive. André Neubauer: Yeah, but you were right. Yeah, good point. Olaf: Yeah, yeah, yeah, yeah. It's and and the thing is of course if it if it didn't work, I had no clue how to fix it, right? Yeah, but it Sebastian: Yes. Of of course, yeah. Olaf: worked, but it worked. Yeah. Sebastian: Yeah, cool. Thank you. great examples. Yeah. And yeah, do you maybe the the final section is always we're asking for a prediction. Do you have a prediction for us that you want to share with our listeners? Olaf: Yeah yeah I can I can share. I I think what we're what will happen in in the I don't know when, but I think we will see a a shift from trying to throw as much CPU memory etc. at at the frontier models to more shift into higher efficiency. You already see that happening with Deep Seek, for example. They try to be much smarter and more efficient in in in how they approach it. And I think again this lean thinking, when it's good enough, it's good enough, right? if I run a local model on my machine and it and I need an output that I need to kind of read, I don't need I don't know, two hundred characters per second. That's way too fast. I cannot read two hundred characters per second. It's okay if it does ten characters per second because I can just read André Neubauer: Yeah. Olaf: the lines, right? So I think we will be shifting more into this way of working where we select a model that fits the the task at hand and it's more efficient or smarter in in what it needs to be doing. And and but on the other extreme you now see this like silicon models in silicon, right? That are like insanely fast. But maybe that's another dimension of this thing. It's more like, okay, what does this thing need to do? And it's kind of like a fenced off box, we know what it needs to do. Right now the state of the art is good. We we we we burn it in silicon. I I'm I'm wondering what if FPGAs might be an interesting step there. Like the the FPGAs are like DSP chips, you know, but you can reprogram them. might be an interesting interim media strip. So I think there will be layers of of hardware to open source and anything in between where we kind of orchestrate between yeah. Sebastian: Yeah, I I tend to agree. I'm I'm still waiting for the RAM Pokalypse. I don't know if you've heard of this, but there is a new Chinese manufacturer of really highest throughput RAM that is like getting on the market. I think they released their first product yesterday, still very high price, but the hope is that finally more demand more supply will lower the prices again and then I can finally buy my used RTX three ninety to set up a no a a local Olaf: Yeah, yeah, yeah. Sebastian: gear. Olaf: Yeah, if you're a little bit cynical, I mean Samsung is I think making like insane profits, so that they don't have any le any incentive of changing their production or scaling it up. It it's good as it is, so i they probably need external like pressure, market pressure to start changing it again. Because I think it's not a production or technical thing right now, it's an economics thing. So yeah we Sebastian: Yeah, th th though you have to like also include that these these suppliers, producers, manufacturers have been burned in the past, right? Because there's this in Germany we say Schweinezyklus, so so the the cycle, right? Where it this Olaf: Yeah, yeah, the the yeah, the big cycle. Yeah. Sebastian: is the the the chip cycle and they have been burned several times, I think, trying to ramp up production when demand was very high and then like demand was cooling down. Yeah, we we'll see how how it goes. I think demand will not cool down very soon right but yeah let's see maybe also there are also approaches to have these models run more on CPUs as well right or utilize smaller v RAM like free token and whatnot but yeah let's see interesting and thanks for yeah great great conversation thanks for coming on the show Olav was great to have you yes Olaf: Yeah, thanks for having me. I really enjoyed André Neubauer: Thanks for the invite. Olaf: it. Yeah.