Sebastian: you pretty much jumped ahead, so it's not the review any longer. You say Basically the the pipeline is so good that solid software software is coming out, but the bottleneck actually has been revealed in being before that. So you shift shifted the bottleneck left and and this was the PRD quality, right? Jürgen Helmers: Yes. exactly. Sebastian: Welcome to another episode of our new season of Beyond Vibecoding partnering with Impala Search. André Neubauer: The go to tech and executive search agency in Germany. Sebastian: In this podcast, we explore the transformational change in software engineering and knowledge work in general. I am Sebastian Heidemeyer zu Erpen, CTO at Ecosia. André Neubauer: And I'm Andre, CTPO at Trusted Shops. Great to have you back. today we are talking again to a well, yeah, a very senior tech leader. We are talking to Jürgen Helmers, precisely Dr. Jürgen Helmers, SVP engineering at Undercore, and yeah, many interesting past positions. Sebastian: Yes, Jürgen also went through the Paloa school, like some of our guests already also have been from the Paloa team, and he's implementing loop based engineering across the engineering org at Andacor. He gives detailed explanations of of their approach, and it's really interesting and worth a listen, so hope you enjoy. Welcome Jürgen Helmus to our nice podcast. and we're happy to have you and would like to ask you to introduce yourself to our listeners like usual. Jürgen Helmers: Sure, thanks Sebastian. My name is Jürgen, I'm VP Engineering at EnderCore. It's a small startup here in Berlin. Originally I'm a biochemist, I have a PhD in biochemistry, have done a lot of work that is really computer intensive, scripting and building 3D models. I joined the Berlin startup scene as a Ruby on Rails engineer, eventually became a manager, had several stints at companies. I worked at Wayfair, Talo Fresh, Paloer before EnderCore and now I run FAC at Ender. Sebastian: That was concise. Thanks a lot. yeah, Jürgen Helmers: Pleasure. Sebastian: and s super cool also your move away from the academia and then into into engineering, which actually also makes sense because you told us in the in the pre-chat that you had been coding as well, obviously as a scientist, which a lot of scientists Jürgen Helmers: Yes. Sebastian: do. so it perfectly makes sense, right? But this also like then implies that you are Yeah, d you have been coding, right? So you you probably do code occasionally still or or do you use yeah, probably right. So the the first segment is always the status quo. So how do you work? How do you use AI? What's your tech stack personally? Jürgen Helmers: Yeah, that's an interesting question. the longest time, definitely at Wayfair and at HelloFresh, I was a people manager and building strategy and kind of like thinking about what can we do beyond the quarter. And at Endercore, it's a very different because it's a very different company. Wayfair, we had teams for everything. There's the SAE team, there's the platform teams, as in the distribution of responsibility was actually very wide and there were different teams with a very distinct purpose in the company. Here at Endercore, it's a classical startup. You wear the hat and you do what actually is needed for the overall company to be successful. So when I started Endercore, we had the need for increased velocity in engineering and So it would have been one thing to make plans of what it actually is we need to build. And I also did and still do that. But at the same time, I actually realized that with AI, I'm empowered again to actually contribute to engineering. And what we're using is the entire company is Claude based. So everyone at Endercore has a Claude account and there's instructions on day one on how to set up Claude Cowork, your working style and your role definition in order to get you effectively using AI. And there's a couple of tips and tricks that are part of your setup document that actually encourage you to not spend time on redundant tasks. So every report I have to file, the weekly update that I give, the evening one minute email that I do write, everything is based on an agent. I do a lot of recruiting because we're hiring at Endercoil. And I write a lot of scorecards and they have a certain format. So I created a skill that takes the transcript of every interview, as well as my personal notes where I only indicate yellow or red flags. And a scorecard is automatically created against our career framework. I use it very often, daily almost. I only had to override it twice. So it's actually working quite well. And it's... Essentially, the guide is that, of course, if you do it the third time manually, create a skill. Sebastian: That makes a lot of sense and is also a great guideline for the whole team. And if I understand it's this is true for every single person at an air. Jürgen Helmers: It's literally everyone, everyone, of course, that the idea is that we intact because we use AI very intensively also in the software development lifecycle that we are a bit of expert when it comes to AI. So prompting techniques, but there's few shot, for example, you know, that is that is not really known to a lot. So if anyone has trouble André Neubauer: Hm. Jürgen Helmers: and is struggling in actually reaching their goals, we are educating and sharing our expertise. Sebastian: Yeah, I think this is a very André Neubauer: Nice culture. Sebastian: yeah yeah, exactly. It's a very important topic. And and I think smaller companies have an easier time implementing this because they are so l mostly younger and and like more nimble, right, and more agile. How Jürgen Helmers: Yes. Sebastian: old is Endercore, by the way, and how big? How many employees? Jürgen Helmers: Endocore is five years old and has been in stealth until the beginning of April this year. So it's very recent that it Sebastian: Okay. Mm-hmm. Jürgen Helmers: has come out of scale. We just hired the hundredth person two weeks ago. So roughly a hundred people. Sebastian: Okay. Yeah. Got it. André Neubauer: but still at a decent scale, right? It's not like another ten, twenty people company. So a hundred, like Jürgen Helmers: No, we don't fit into a small room. have multiple areas in the office. So it already has a certain size. And it is also that people use AI just to maybe extend a little bit on the topic that I introduced before of how we use AI. Very recently, a sales agent actually created a tool that solved the problem for him that we actually have on our roadmap. and of course ours is complex, It integrates into our back-up tool, which happens to be Salesforce. It has, as far as transcripts are concerned, a quality gate, and it makes it really, really complicated. Very, very nice if you have the time to actually implement it, so very usable. He created a quick and dirty solution. So he created the skateboard while we were actually planning the car. And André Neubauer: Hm. Jürgen Helmers: I took it and actually built a tricycle that I today rolled out to another sales agent. have a hope that this will actually solve our immediate problems. So it's a clear sign that not only in engineering, but also in sales, people think about AI, they try it, and they literally solve their own problems with it and they experiment with it, which was very nice to see. André Neubauer: This is a great culture. I need to ask Sebastian: Yeah. That's really cool. André Neubauer: for like for on on own interests, like you're running cloud native, like is there any other platform you're running to make adoption easier for non technical people? Jürgen Helmers: We are using Claude Desktop, obviously for everyone who doesn't like the André Neubauer: Sure. Jürgen Helmers: command line, But essentially it's the same thing. When I joined in April, we still had Copilot, but we realized no one is really using it, right? So we had Cursor, and then we realized no one is opening that except looking at markdown files, but no one is actually actively working with it. With the introduction of Claude Code and the harness, it has been completely replaced. André Neubauer: Okay. Yeah. Thanks. Jürgen Helmers: which of course, when Fable was canceled, that's an indication of how dependent one actually makes oneself if one has all eggs in one basket. Because if they should decide that a token now costs 70 euros instead of seven, a million that is, we would really struggle, of course, and would might want to evaluate open source models. instead of a Claude model. You can still use Claude as an application and just plug in another model, but of course you would need to do some testing. So currently we are dependent on... Sebastian: And your thoughts on this are as you just laid out to then quickly use an open source model E via open co sorry open router or something and then Jürgen Helmers: Yeah, I mean, at home, obviously, I have been experimenting with Olamma and I improved my digital personal document filing and the OCR aspect of that with AI by using Olamma because I didn't want my bank statements and my personal communication to hit the cloud. It is awfully slow, but that is mostly due to the André Neubauer: Mm-hmm. Jürgen Helmers: limitation of the machine that I have because I have a Mac Mini, so no big graphics card that could run any more advanced model. I think the trend is an interesting one, right? The EU is completely dependent on the US on this one. And that is something that I think everyone realizes they don't like it. I don't see any larger efforts to actually change it and invest into it, which I think you would need to. cost to develop André Neubauer: Yeah. Jürgen Helmers: these models is tremendous. And I currently don't see anyone who's prepared to actually invest into it just to have an alternative. Sebastian: I mean to be fair, there recently have been some releases of of models, actually even German models, right? I wouldn't say that these are like models that can compete at the highest levels, Jürgen Helmers: Mm-hmm. Sebastian: not even at the not even at the not highest levels, right? but I also am not sure if the model layer actually is the one that would be relevant in terms of sovereignty, right? As long as there are open Jürgen Helmers: True. Yeah. Sebastian: source models with open weights that you can use. Jürgen Helmers: Yes. Sebastian: Yes, you do have to build some guardrails around stuff that you want to do with those models when it comes to pre trained data, right? Because then you probably not all answers about Taiwan or the Tianman place will be correct, but at least you have models that you can use for agentic use cases, coding, whatever, right? So this Jürgen Helmers: Yes. Sebastian: this is probably possible and that therefore you basically just need the just need the the compute, right, in order to run. Jürgen Helmers: You do, I mean, ever since the first of April source code leak of Claude itself, you you do realize it's a smart piece of software and the reliance on the model is actually, you can actually replace it and you can run other models within Claude. Something I still want to do, I haven't really tried it. I mean, me being the scientist, I would want a similar setup to actually kind of like test it properly and have some metrics. So it's like, it's not as good, but... What is the not as good? What works? What doesn't work? So I think if we ever would get to the situation, I think it's definitely possible. Do I currently invest into it? No, because anthropic and the offering is just too Sebastian: No there's no need to. Yeah. Yeah, yeah. Jürgen Helmers: tempting and too... Sebastian: Yes, makes perfect sense, I think, from where you are, just at some point from a business continuity perspective, it probably would make sense to r run some experiments, right? I would assume that the way you're using it, the way you you set up your specialized agents, the loops, the skills, it's probably just fine going for a GLM for a Kimmy K three or whatever, right? So it probably would work. Jürgen Helmers: Yeah. Also, mean, at Endercore, where we have an agentic layer that we actually use in production, there we don't use Claude at all. We, of course, make it work with larger models, and then we ever so slowly walk back to actually the cheapest model that can actually do the task in good enough. So we're actually using Sebastian: Makes perfect sense. André Neubauer: Yeah, makes sense. Jürgen Helmers: flash models, we use OpenAI, we use the Gemini models as small as we can make it. André Neubauer: Okay, but also there you have a certain dependency. You just can choose, but Jürgen Helmers: Of course. Yes. André Neubauer: okay. I'm wondering if that Sebastian: That's a good thing. André Neubauer: might be I'm wondering if that might be a playground to to learn for engineering. So like the other way around, like not not testing, no, not doing testing for production, but the other way around doing something stuff in production to then use that internally. you get what I'm saying? Like Jürgen Helmers: Well, yeah. André Neubauer: yeah. Jürgen Helmers: Yeah, yeah, I mean, I get a lot of calls, of course, from companies that rent GPUs somewhere in Europe. And yeah, you can run the open eye compatible open source models and it's 95 percent is good. But my current motivation to test those five percent difference is not really that high. You know, if you have something that works, you prefer to continue. Sebastian: Yes. All right. let's move over to the main section of our podcast Jürgen Helmers: Mm-hmm. Sebastian: episode. And this time it is about loop and harness engineering really in a distributed organization like you are running it, right? So yeah, please tell us about it. Jürgen Helmers: Yeah. When I joined Endercore, and of course, engineering team is a small team at the time, we're like nine people. They were like, yeah, we're using AI, but they were doing mostly prompt engineering and with a little bit of context engineering by pulling in information from different sources. I then upskilled the team to harness engineering by creating a harness that would connect them to JIRA, which we use for project management. to the Wiki for architectural documents and ultimately starting to put up guardrails for engineering best practices that we want to work API first. I created an engineering vision, which was a markdown document that could be ingested for planning. And I quickly arrived at the need for that already existed in my former companies where you would think, what is the smallest spec file that could hand to an agent to have a one-shot success, that it implements it and I have almost nothing that comes up during a review of the code, that the code is good enough. And I soon realized that mostly during AI assisted reviews, that there were always defects, always defects that were listed as blockers and as majors, as in you couldn't put it into production. One of our Nisho engineers, I think it was Alexei if I remember it correctly, he created a deep review skill within our harness that would actually follow the best practices for Java. Our main stack is Java, we're using Java 25 Spring Boot, so that's what they're using as well and they are currently tasked to actually extract and migrate our logistics domain out of our monolith. into a microservice platform. And they wanted a good and solid code review because AI produces a lot of code and it's very hard to review all of that manually. Instead of actually creating a loop there, they posted the code comments to GitHub as reviews. I saw the chance to actually create our first loop. I changed and took his very good review skill and I defined a loop. I defined the loop as such as that of one agent that actually does the review and loop engineering obviously communicates with the file system. So that one agent that does the code review creates an issue MD file. Then another agent reads the issue MD file and tries to fix the majors and the blockers, then hands it back to the review skill in order to have another round of review. And I defined the gate for this initial loop, and that was my very first time that actually tried loop engineering. I defined the gate as in there's no blockers, there's no major defects. And I defined a circuit breaker that this loop should run maximum three times to not endlessly burn tokens. To this day, we never hit the circuit breaker, so that never triggered. and we have a review skill within our harness that you would still need to manually trigger at that moment in time that would result in improved code quality that would pass a normal human code review and that would greatly increase the chances of this piece of software actually operating as expected in production. Once that was in place, I took it one step further and something that I never really liked with our harness is that it created these waterfall releases for complex projects. So it's like it created really long documents that would cover essentially a roadmap item and it would work until the very end to then spectacularly fail, right? So I experimented then at that moment in time with I asked the agent to create meaningful milestones. Then once I had the milestones, I actually realized I could implement these milestones like we would in the past, where you would like milestone by milestone, and start implementing these. And then I did realize that with the review skill that I could put that, all of that in the loop, that I would have a milestone run and then run the review skill. And if the review skill actually revealed the gates are green, that I could then create a pull request. in order for it to potentially start the next master. I hooked it up to Jira. I made a Jira base so that multiple people could actually work in parallel on a project by having an epic per effort, per loop. And I moved all of the files that are necessary, the Claude MD file, the issue file, the loop MD file that captures where the loop is actually currently residing. that all runs in a specific sub-directory so I can separate concerns and multiple people on the same repository can actually work on different loops in order to prioritize. I would move based on the loops progress, the JIRA tickets, so I had observability in that point in time. I would automatically update and create documentation, which is something that in my experience always usually fails because that's thrown out of the window first. So it would know where our parent page for documentation actually is in order to automatically create it if it doesn't exist or update it based on the progress of the loop. I created this because I'm a big fan. I have to admit I've never done it myself very effectively because I was a better coder than I was a tester. I implemented this loop and the execution of the loop in test-driven development. My friend David Sammert, he... He works for Equal Experts and he's a great coach for test-driven development. He actually proposed that the quality is much better, so I made use of it. I have a red agent that first writes the test, then I have another agent André Neubauer: Mm-hmm. Jürgen Helmers: that operates independently to actually make the red test green, so that's the green agent, and then it's handed off to the review and the orchestrator... then creates another agent that does the fixing. So everything runs in separated contexts. I have an automated stop in there if one of the contexts reaches a certain percentage where it starts not making sense. So a lot of thought went into it to ultimately create a system that if you know what you're doing, you can create actually solid software with it. It is very fast. It is very reliable. However, it has one Achilles heel and that is the quality of the plans you actually send into the system. It quickly revealed that our PRDs have a lot of gaps, have a lot of holes. So it triggered some process changes in our company and our technical organization that we use grill me sessions to actually have an agent poke holes into our plans. It makes them more consistent. It generates a lot of open questions. I realized that an agent produces questions in a way that stakeholders don't understand them. So I built the skill that depending on my input, it takes a PRD, a technical design document and open questions. And it assumes a role. I can send it, hey, this is for the head of logistics, for sales or procurement. So it puts that head on and it avoids any technical language and it starts speaking the language of the audience. so that they have an easier time actually answering the questions that otherwise I would not get any answers to. This way, we're shifting left. We are in a state, and this is very new, so definitely for Q4, we want to make that our standard process is that we plan, we plan PADs. We grill them preferably together. We answer the open questions in the language that our stakeholders can understand them. And that's of course more the business questions. The technical questions can remain very technical in order to have a better understanding and a clear and sharp definition of what the scope is, but not the scope is what the expectations are, what really needs to be implemented, which allows the loop engineering to do a better job in producing a piece of software that fulfills actually the stakeholder requirements. Sebastian: Sounds André Neubauer: I think. Sebastian: impressive. Yeah. André Neubauer: I think we need a moment to digest that. Jürgen Helmers: Sorry, sorry for lecturing for so long. André Neubauer: No, no, no. No, I think it was Sebastian: good. André Neubauer: it was good. Let's let's unpack it. I think this is your this is your point usually, Sebastian. Sebastian: so usually when we talk about where the bottlenecks move because they move away from creating code, right? They usually move to review. In your case, you pretty much jumped ahead, so it's not the review any longer. You say Basically the the pipeline is so good that solid software software is coming out, but the bottleneck actually has been revealed in being before that. So you shift shifted the bottleneck left and and this was the PRD quality, right? Jürgen Helmers: Yes. exactly. And this is something I might refer to Dave again. He has a paper route that I think is a very good one. What is the effect of AI on the software development life cycle on teams? So it's reflecting on how the team behaves, how the team feels about it. But also the fact that AI is greatly accelerating. It's not only accelerating what is good about the software development. And that is you get faster to a finished product. It is accelerating André Neubauer: Hmm. Jürgen Helmers: and identifying every single gap and everything that's actually wrong with your organization. If your planning is no good, it's no longer revealed over the course of a quarter. I have built an audit logging service with ingestion, with querying, and with SDK and Kafka support layer in 24 hours. because I went to the YOLO mode, I experimented with that as well. So it just went ahead and I structured it after every pull request creation, merged it and go to the next one. Very risky, but a well-defined kind of like task because audit logging, you find a lot of like industry standards and good information online, which I included into my technical planning. You can create things extremely fast. So it's no longer a matter of months, quarters, or even weeks or sprints. is actually you reveal what's good and bad in a day. And to the point is that we're moving through our roadmap faster. We also doubled our velocity speed in the team, which also comes with exhaustion, I have to say, and because it's continuous scope changes, I realized at some point I can have four terminal windows open at the same time and run four loops at the time. With the fifth loop, I stop. I stop ingesting what I'm actually doing and I need to read and spend too much time in understanding what's actually going on and I can't answer the questions fast enough anymore. And I overwhelm myself. That is the time where André Neubauer: Mm-hmm. Jürgen Helmers: I realize I start dreaming about my loops. That affects my family. I have three kids and that's not a good thing. So you need to slow yourself down. It is one thing to be able to produce software fast and in decent quality, let's call it. It's completely different thing what it actually does to you as a human being. Sebastian: Absolutely. This this actually reminds me of the Steve Yeeggy article. I think he called it the AI Vampires or something, where he also suggested because of the higher degree of intelligence that our work now needs, so there's much more like less rote g tasks that that we're executing, right? It it's much more i demanding from us and many more context Jürgen Helmers: Yes. Sebastian: switches were exhausting much much faster. He said that at 11 a.m. he's pretty much done for the day and he can go go to bed again. and so he suggested actually a three hour day workday maybe. Maybe that's something for the future. But but then there's also other things at work that can be done and that should be done and that also add value that are less exhaustive, right? That that are yeah. Jürgen Helmers: That's a very good point. At Wafer, I remember I was sitting in the lobby area on our floor and I was thinking and one of my peers came by, like, why are you not working? It's like, I am working, I'm thinking. And ultimately finding the right balance in executing so that the business goals are actually met in time because ultimately at the end of the day, any engineering team's responsibility is delivery and you take care of your roadmap. but at the same time it allows you to actually, while your ages are running, there's no point in staring at the screen as tempting as that might be initially, is to actually think about what do I do next and what actually is good design and brush up on your system design. Because the better your design is, the better your plans are. That is part of it. And you might want to think about you are creating another piece in a microservice ensemble that at the end of the day needs to work together. And how is that actually working? Because the loop is currently not knowing anything what's going on in the rest of the estate. And that then brings me to my hopefully next interest that I will have some time to pursue is how do I make it that everything is becoming more aware of each other? You know, we already struggle with very simple facts like my architectural vision has no record of what's already existing and what's not. So I actually need to update it. So I started experimenting with automatically updating our vision to this already exists. These changes have been made because otherwise it's very quickly outdated. And then the instruction of what the big picture actually looks like is ever so slightly off. And if an agent is ever so slightly off, it starts trying to fill gaps. So I also stored it definitely during planning, I became very, very strict and I included the section of you do not. invent scope. You need to ask questions to kind of like preempt the what was actually happening without these kind of like guardrails that it filled gaps. It put things together and invented formulas when it was coming to like implementation details where I needed to separate where does that actually come from and I couldn't find the source and that needs to be prevented. Is that 15 % of hallucination that actually diffuses and waters down every plan. And one needs to be very careful there. You need to be at high quality of the plan that goes in, but at the same time, you actually need to review the plans. It's not enough to just convert it into a technical design document than into a loop plan. We need to shift left. And this is definitely a learning out of, let's say, well, now almost two months of loop engineering. You need to spend time reviewing your plans because just executing them as fast and easy. But what if you built the wrong thing? It consumes a lot of tokens. It's very expensive. And I got the go from Philip, our founder, that that is fine. So there's a certain, of course, there is a certain limit, but we are not any close there yet. But you really need to pay attention what you're doing. Otherwise, it's by coding. André Neubauer: can I can really I can really relate to that. I think at a at a scale of like a solopreneur, I think that is easy also to implement. What I find all the time very hard is like to keep an organization in sync with that. So what you just described is like a lot of change, right? Like it Jürgen Helmers: Yes. André Neubauer: in in in and the speed at which this change happens is like is actually accelerating. So I'm wondering how to also from an organizational perspective, how to ensure that the organization can keep up with that pace. Welly, how do you do this? Especially since Sebastian framed it in the beginning, you are also distributed. Jürgen Helmers: Yes, we are distributed and we made the mistake, example, our current roadmap to focus on one domain because we wanted to extract it out of the monolith and move it to our microservice platform. But that made it with the advent of loop engineering, it accelerated so much and so many open questions piled up that we had exactly that one stakeholder as the expert who was suddenly faced with 100 open questions and that the engineering team. that is distributed, they're working for initial provider, they're actually located in the Ukraine. They're very capable, they're extremely senior, which means they adopted this extremely fast. They're actually contributing, they're improving the harness, they're making it faster, they're making it more cost effective. But the side effect is that what I actually had them plan to implement until the end of the year, they're now ready next month, which means we need to get our A into G in order to actually come up with more plans so they actually have something to do. It is accelerating everything. It is a challenge not only for the engineering team to kind of like find the right balance between in the past you would write manual code and that was relaxing. It was something you enjoyed doing and it just came up in a retro last week. That's like, this is missing. Now it's just like everything is fast and it's extremely fast-paced. And it trickles back. to our entire organization that we need to think about faster in more detail what it actually is we all want and then map it back to our vision of where we actually want to go, which means we need to make faster decisions, not only in tech, we need to make fast decisions in business with our stakeholders and they are sometimes overwhelmed by the speed in which we actually want to roll out because suddenly we have these three things that need to go live on the same day and you're like, no, you need to stagger it. because otherwise you don't have André Neubauer: Mm-hmm. Jürgen Helmers: the time to monitor things and it's simply too fast. André Neubauer: I need to ask an follow up question on the technical thing. like is there one harness to to rule them all? I do not know your technical footprint Jürgen Helmers: you André Neubauer: at and a core, but how complex is that? And you have then one set of, I don't know, skills to to run the loops or how is that how Jürgen Helmers: No, it's a combination of skills that are ultimately nothing but markdown files, right? And there's some connectors, they're MCP skills, or we actually have API credentials in a .env file that gives access to tooling. It is a collection of skills. that ultimately are interacting with each other. So there is an André Neubauer: Mm. Jürgen Helmers: agent that is orchestrating and fanning out. So ultimately it's a multi-agent setup, right? So there's this agent that's always running the other skills and is delegating work. In the loop recipe itself, I also implemented the use of different models for different tasks. We use Fable for planning, because that's where it comes. Then we use the Opus models for the implementation as well as the review for anything that is interacting with like third party kind of like interfaces, whether it be Jira, whether it's GitHub to create pull requests, we're using the cheapest Haiku models. André Neubauer: Well thought through. Sebastian: Mm-hmm. thanks. you you mentioned that the speed has increased so much that it's almost going too fast and that that it's hard to come up with yeah enough plans to keep the team busy, which is like sort of a luxury situation, right? Even though it it it like reveals an another bottleneck. And Is this and you mentioned that the engineers and the in the nearshoring teams are very senior, very experienced, and that's why they're adopting Jürgen Helmers: Yes. Sebastian: this very quickly. Is this true for the whole organization or do do you need to like do certain things in order to make the whole team keep up and yeah, stay in sync, basically? Jürgen Helmers: Yeah, no, that's a very good question. We struggled with this already at my former company where we thought we show how it works. So kind of like in a frontal knowledge sharing and then everyone will be amazed and will be adopting it. That was absolutely not the case. We had very few. that jumped on this bandwagon because they were excited to fiddle and they were excited to actually do something new. Ultimately, it's change management. So you have the classical characters. You have the naysayers. They're like, no, I'm still faster if I write my code manually. I'm very good at it. And you have the yeasayers. They are the ones that really invest, but the majority of the group, the median, they can go either way and you really need to convince them. at Paloa where I worked before, went from frontal to actually a hands-on workshop. I still remember Pedro actually sitting up front and guiding everyone through and actually taking everyone by the hand. And I learned from that, I adopted that. So I started with knowledge sharing session. I then took an example of, hey, this is how it actually works. And I realized that... People have very personal access to the eye. They connect very different things with it. While some actually do like to see them still as a craftsman that manually creates something. And I was very much drawn to software engineering because I like woodworking. So I also play different instruments. So I actually play the guitar. So my real goal is I will build my own Spanish guitar. And... Here we had some that really liked the manual aspect and they insisted on it for their own benefit and they think quality can only be generated manually. With the advent of models getting smarter and better and us creating a harness that ultimately is self-improving, where you can actually tell the harness, I noticed this and also collect some metrics during these runs, so I know where there's a lot of time being spent. Just recently I actually I thought I reviewed and improved a contribution from our near shore engineers that were supposed to save tokens and make it faster and I completely made it worse. I burned token like ice in the desert and it took forever to finish a single loop. So I realized I miss also someone that truly takes ownership of the harness to actually guide it and have a focus on that every pull request is actually making it better. Currently, it is a group and team effort that is distributed in itself without anyone truly owning it. People still assume because I am accountable because I'm the VP engineer, it's like, Jürgen, you need to review it. But sometimes this was a pull request for like 94 files changed. I didn't look at all of the 49 files. I didn't have time for that. So it's like, yeah, well, Alexei, knows what he's doing. I just fix it a little bit. turns out I was actually making it worse. So in its sense, is a system that by now you can actually ask it, and this is the smartness, of course, of the models. This is what I observed. This is the expected outcome that I would like to have. Improve yourself. And that actually works quite nicely. It's very effective. Sebastian: And then basically to go back to the to the knowledge sharing, so knowledge is shared via the harness. So pretty much people are using it as long as the interface to the human doesn't change the same way, but it the outcome is is better. Or I mean also the interface will change, right? And other things. Jürgen Helmers: The interface we're using in the terminal, so it's a terminal application that we're using, so the interface is old school, right? I did notice, that also a lot of people, or some people, they struggle with precision. It's like, thank you, and could you please? I mean, that's just a waste of context from my perspective. I mean, it's very important on how we interact as a team, so that's another challenge. It's not only a scope challenge, it's also a challenge on How do we actually communicate effectively? And crisp and very clean instructions to an agent are extremely beneficial. They're the death of every launch meeting because that's not the way you communicate. So that's another scope change that our team has to go through. But using extremely precise language and be very specific what you expect and what you not expect allows you to actually have the self-improvement. and ultimately have the results you're really after. And this is something you need to learn. This is something the way every harness itself is ultimately a interesting network of different skills that have a certain behavior. And you could argue that the Google search prompt, I know how to search things and find things very effectively. My first wife or my dad, they don't and they They describe things in a way that leave a lot to the imagination and they're very ambiguous, which means you're not finding what you actually truly after. And the agent is very similar and the plans are very similar. I, for example, noticed in one engineer, he used a plan to create an RFC. Out of that, he wanted to create a technical design document, but he actually did a loop in itself. He created an RFC based on an RFC. because he didn't use very precise language. That was a lesson for the entire team very early on. It was only an hour that was wasted, but it was a good learning. Ultimately, what worked for me to go to that way back to the beginning of the question, Sebastian, that you asked is how do I got the team? And this is why I got back into actually hands-on engineering myself. I did realize I can talk about it as much as I want to. I can share my, at that moment, to the team apparent theoretical knowledge, whereas they were actually asking me, it's all nice that you talk about it, but we don't know how to get started. Can you help us to get started? So I was sitting myself down with a hands-on workshop and then literally with pairing to support them to just get started. Because like everything in life, the first steps are the hardest. You need to overcome this kind of like hurdle before it becomes actually self-fulfilling, but they then realized is that It's very verbose, it's a lot of code that's being produced, but the outcome actually, it does what it needs to do. Is it giving us this wow effect? Like in the past when I got to review a piece of code by a staff or principal engineers, like, wow, these few lines of code do all of what it actually is supposed to be doing. And that was an amazing feeling to just look at it. I don't have that with code that is produced by André Neubauer: Hm. Jürgen Helmers: any agent or definitely not by a harness. Is it actually working? Well, I made sure it's working because of this test driven development that is not only unit tests, but actually integration tests. It also gives us the opportunity once a recipe or loop file has been created, a loop planner should say, that I can inject the test plan that our QA engineer, Yoya, has been created in order to further specify with each milestone what the gate really is to ensure. that what we produce is really what our stakeholders want to then greatly increase the chance that what is coming out of it during regression testing, not only can we click a button and we run the test suite, we know that what we wanted to do is actually being provided, but that's being done in the nicest way. Well, we are startup, right? So we need things rather faster. We are not really building yet enterprise software where we look at, okay, how much time do we need to spend on maintenance? It's important to have features forced at this moment and in this phase of the Sebastian: All right. André Neubauer: Amen. Sebastian: Yeah. and and the the topic of getting back into coding and pretty much understanding what you're talking to with the team, I can so relate to that. That was also my impetus. So over Christmas I spent a lot of time with the agents in order to understand how to best use it and how to apply it in our domain and then share it with the team. Absolutely. Yeah. Jürgen Helmers: Yeah, mean, is something that is talking about it is one thing, but if you then can't do it yourself, I mean, it's a little bit of this like street cred. What was wrong before? Because, you know, at HelloFresh and at Wayfair, people were like, you shouldn't be coding. What are you doing? This is my job. You do something else. Here at EnderCode was a little bit different. I talked about it. I painted the picture of what we could actually accomplish with it, which means we deliver our stuff faster. We have more time to actually improve our plans. our vision of what it actually is we want to create. And they realized, yes, the speed is possible because I've shown them, especially with my audit service or my three services with the audit service suite, should say, audit logging, I need to be precise. That was literally built in 23 hours. And they were like, wow, that is possible. And can you actually use it? Well, I had an integration test that was cutting across these three services and working together, that's like, yep, here you can see it. I tested that it actually works. And now we're using it, we start rolling it out in production. And it is one thing to see it, but like I said, they needed help to actually get started with the harness because you need to be able to express yourself in a certain way in order to use it effectively. Sebastian: Yes, amen. All right. André Neubauer: I I think the secret source to that, and maybe I just want to challenge that. also want to get challenged that by you, Jürgen. like we often we often have these like people in in in the podcast and I'm I'm wondering like how you like what's the secret source to be able to step down? Because this is actually what it is about, right? Like so not not not people management, but like getting back to software engineering. Right. And I'm I'm wondering, is that the knowledge about software development lifecycle and all the techniques? I think in one of the recent episodes we talked with Ben Hoskins about that. And he was also like Yeah, Jürgen Helmers: Yeah, my former manager. André Neubauer: and I I know that's why I'm I I I brought this up. he we talked so much about the techniques, and I think as a like I think we are more or less all in the same age group, right? We we know these techniques. And I think this Jürgen Helmers: Mm-hmm. André Neubauer: is maybe the secret source why we like this group is able to step down. I'm I'm other than that, I'm really wondering like what is the reason why managers is are capable, like really doing these kind of things, right? Getting the hands dirty. You have an opinion on that? Jürgen Helmers: Yeah. It's a very good question. yeah, I do. It is ultimately when I used to be a Rubin Rails engineer, right? So new versions coming out, I always thought, ooh, let me just try that. But just the setup of tooling, installing your gems and dependencies and getting something up, you could actually get into this flow of I can actually produce something was cumbersome. It was hard. Now firing up, you know, an agent that has access to context seven, every document ever is like, hey, set myself up for a Rubian Rails in this and this version. It just does it. So it actually reduces and takes down this threshold that before it was hard to jump over. And it also required time. You needed the time to actually do that. Me having a family, having multiple hobbies from playing instruments to going cycling, is a... Andre, you know that, you're cyclist yourself. I sometimes disappear for these four or five hours to the... to the dismay Sebastian: Yeah. Jürgen Helmers: of my family, I simply didn't have the time to actually fiddle then around with a computer and trying to set up the latest version on Ruby on Rails. Now it's become awfully easy and I find myself on the sofa on Sundays and I just, I can play with it and I like building something. Me not being a software engineer and a computer scientist by training. by actually being a biochemist, I like building things. I like to find out how things work. And for me, this harness is a system of multiple elements and multiple agents working together. And during my time as manager, it was my job to create the perfect team. And here and there, I actually succeeded. had the team that everyone was happy working. I made them effective. I invested into the scrum ceremonies and I made it happen that this group of people was actually working in perfect harmony to the point I still remember one of them, Runeish, coming to me, like, Jürgen, I think we can do anything. And I think that was still to this day is the nicest thing that someone of my team, to me as a manager, ever said, because it meant all my investment into this group of people actually had paid off. They felt they can do anything. And that was what management was all about. Now this is changing. How did you ever become a senior if as a junior you have Claude Coat at your side. How do you make these long lists? How do you gain this long list of failures that ultimately made you into who you are as a technical person? That we, the new generation of engineers is missing the most important experience that life has to offer. And that is failure. This hands on. interaction with something that doesn't work, you go to stack overflow, you try all sorts of like really old suggestions and you somehow make it work. That goes completely away because the agent is abstracting this complexity and you become very early on this conductor of an army of smart senior engineers. Whatever that will do to the profession of software engineers. Sebastian: I mean we'll see where it goes, right? But I think the level of failure just moves up, right? So your your v your errors or mistakes are happening on on a different level. Not on the code level, but maybe on a different level, right? The plan is shitty or whatever. Yeah, yeah, exactly. Jürgen Helmers: True. And your plans are not good and so you need to invest into system design, which means you need to think about system design already as a junior. If you just start going, you most likely will do something that is not in line with what is needed. Sebastian: Exactly. Exactly. And you'll you'll learn it the hard way also if the system doesn't scale, if the system isn't Jürgen Helmers: Yes. Sebastian: secure, if it's not reliable, right? absolutely. Yeah. All right. let's slowly but surely come to the end of this episode. I just want to, before we move to the reality check segment, I want to move Jürgen Helmers: Mm-hmm. Sebastian: sorry, I want to mention one thing briefly. You mentioned that You're taking in the the questions from the from the grilling session that are very technical and along with some other helpful documents like the PRD and such with the help of an LLM that has the head of a specific stakeholder on, you then create Jürgen Helmers: Yes. Sebastian: a document for the stakeholder to actually consume and then Jürgen Helmers: Yes. Sebastian: be able to answer questions. I heard something very similar from from two folks actually at a CTO craft event that I attended Jürgen Helmers: Mm-hmm. Sebastian: like a month back or so. So this seems to be something also like seems to be a pattern. Like using Jürgen Helmers: Yes. Sebastian: the the the tools in order to create context for non technical people. Yeah. Jürgen Helmers: Yeah. I also realized, at least in our case, that the PRDs have a different structure system, uses the paragraph symbol and a paragraph 6.3, for example, while the technical design documents, it's actually numbered with just 6.3. So references in these very technical first version of the question, were intelligible. You needed multiple documents open at the same time. You would, the first reaction of every staker was, don't understand the question and that it needed me to first explain the question. So thought, I do that every single session. That cannot be. So now the agent actually does the translation and I still have, I think I have two more reviews open to tell it how it actually does. So I built that in, hey, let me judge what you're actually doing in order to perfect it a little bit or improve it rather. And so far the reaction is we are moving not, we used to need like, we used to answer like, and I question every 10 minutes. Now we need two minutes per question, which means sometimes it's like, yeah, it's that. I know exactly what you mean. And I still have the references, so I can go to the document if something is unclear, but it does that for me. So essentially, I replaced myself in the session. There's someone else is explaining it. I can take it eventually asynchronously offline and I can hand them the document. I set an ETA deadline. Please fill it on by then. I replaced myself. Sebastian: Yeah, that makes perfect sense. All right. Now let's come to the reality check. So w we usually ask our guests what are like what the fuck and wow moments that they experience still with these LLMs or AI agents. Yeah. Do you have some of these? Jürgen Helmers: Well, there's probably two. One I already mentioned that was the RLC that was created another RLC. That was a good learning. I have another one that my front end engineer, Madhu, actually shared with me. He was one of the slow adopters of the harness. And he was like, yeah, I don't really like the it produces. it was like, yeah, I have this issue. And it created like... I'm making the number up. It was multiple files to create something. He just looked at it as like, need three lines of code and this issue is fixed. But the AI didn't understand it and literally created multiple files in order to fix something that three lines of code would have done. Of course, an argument and water on his mill that he yep, we probably don't want to use it really. The game changer for André Neubauer: Hm. Jürgen Helmers: him was he was using Fable. And Fable for him was... That was the breaking point. He realized there is something, it's not necessarily AI itself. There's different models that can do different things. And I just resigned, okay, let it be more expensive because Fable is more expensive. But if it gets him to actually produce code using the harness and therefore using the standardization that we have built into the harness, that's good for me. And then of course, then multiple Sebastian: Example. Jürgen Helmers: loops, me trying to fix the the pull request that was supposed to make things cheaper and faster and I made it so absolutely worse. I was actually running it and I saw the count going up and other engineers were coming to me. like, hey, you know you did this and we observed. they waited a little bit because I did the fix and they thought, Jürgen knows what he's doing. So that was a no harm moment and probably also wake up call for them that, I'm probably not necessary, you know. the person that should be the gatekeeper for these pull requests. It was literally, it didn't accomplish anything in two days. Normally it would have taken 20 minutes. So I really made it worse. So looking what you're doing is always a good Sebastian: Absolutely. All right. Thanks a lot. And then final question. Do you have a prediction for us that you want to share with our listeners? Jürgen Helmers: I actually think having observed this a little bit from toy status to being really, really useful is that whatever is the coolest rage, and just three months ago that was loop engineering, I actually do a regular reality check and I'm listening to AI podcasts and read some blogs and there's always something newer that comes up. I mean, you take it from loop engineering to graph engineering, which is something that I'm looking forward to in experiment, because I think it could actually make it happen for Endercore what we actually want to accomplish as a company, as an orchestrated marketplace, where I have sub-agents that ultimately own domains and operate within their domains. But of course, if they operate in isolation, they pre-optimize very quickly to their own success metrics. But the success metrics is ultimately that we have happy customers. How do you abstract that? How do you measure that? Right? So you need multiple levels of gates that are interacting with each other, but someone needs to kind of like own the overall direction of everything. And that we currently do not have. Even when we create software, it's still very simple in an essence because there's one loop running. And if it is, of course, parallelizable, I have two implementations running at the same time. They don't know of each other. they operate completely independently. adding another layer of complexity, I mean, this is ultimately evolution, so I'm back where I started in my career as a biochemist. Evolution only allows you to create more complex system if you actually have an advantage. So it is on us to actually make that advantage happen, that the outcome of the code is actually measurably better. Otherwise, the simplest system will win. And that would... be something that I'm very interested in, that I want to actually experiment here at Endercore, make it happen. I'm currently recruiting folks that are mostly interested in contributing to the AI layer. We have the backend and the migration of the backend to Java microservices. That's the initial team. But even there, I have folks that have an AI background. I want to hire people to actually help me experiment, to ultimately create something that is smarter than what we currently have. definitely is agents controlling other agents. And how many layers you need on top of that? Well, that depends on the capability, on the success metrics, and at the end of the day, on your budget, because it will not be cheap. Sebastian: Definitely. All right. Well, thanks a lot, Jurgen. was a pleasure speaking to you. Jürgen Helmers: Pleasure. Thanks for having me. Sebastian: Yes. yeah. Thank André Neubauer: Thank you. Sebastian: you. André Neubauer: So now to the recap, what stuck with us. Sebastian: Yeah, for me, the most noteworthy thing was that they are clearly past the review bottleneck, basically. So their loop based approach revealed rather the next bottleneck in the planning. And that's what they are currently working on, right? And it's from what it sounded like, the loop based approach seems to be very impressive, seems to be really solid and producing solid software. really curious to see where this is going. André Neubauer: Yeah. And I also would say that at the end you need to fix you need to start fixing at the beginning. so that this like starting or shifting left absolutely makes sense because you could also easily say shift should in, should out, right? so how can you trust the result if you cannot have high confidence in the beginning? so in the spec. so f actually a bit Obvious, honestly, right? We talked about the review topics so often. but that is actually also a good segue to my recap. For me, I need to think about that longer. But I think what we are seeing in all the interviews is there is a pattern that as an tech manager or tech leader, you need to have a deep understanding of the software development lifecycle. And yes, also that. Sounds obvious. honestly, if I look back, a lot of stuff I did during the last years was also on organizational level, like people management, all that stuff. less about the specific software development lifecycle because that was set. But in the agentic era, this is now like so crucial. and without that knowledge, it will you will have a hard time being a tech manager. So this One of the core requirements. Sebastian: Yes, I tend to agree. And what it sparks now in me is that what has been hard, or at least what seems to be hard, is usually taking the flow or the new life cycle that you created and taking it from one engineer to a team. He explained, right, that he did this or they did this with their harness, so they have a harness that provides this. life cycle provides this this infrastructure for their flow for for the loops and he has these integration points in order to improve the plans right with the other stakeholders where then the questions need to be answered that that are coming up through the or during the grill me sessions. But yeah the harness really as the the main tool in order to distribute or scale the the the the agentic software development life cycle. That's very interesting. That's it for today, thank you and hope to see you next time. Bye bye. André Neubauer: Hi.