Sebastian: welcome to Beyond Vibe Coding, the podcast where we explore the transformational change in software engineering and knowledge work in general. I am Sebastian Heidemeyer zu Erpen, CTO at NorthIO. Csaba Tamas: hi everyone, really nice meeting you. So ⁓ my name is Chaba and ⁓ currently I'm leading ⁓ product both product and tech at Palloa for ⁓ roughly one and a half years. ⁓ prior to that I was working at AWS. ⁓ originally I was leading the ⁓ solution engineering team for digital native businesses. So these are like companies like Zalando or HelloFresh. ⁓ Sebastian: The Beyond Vibe Coding Podcast is a project by Sebastian Heidemeyer zu Erpen and Andre Neubauer, partnering with Impala Search. The content is created by us and our guests. André Neubauer: Was hat mitgetragen? ⁓ Sebastian: Ja, definitiv ein tolles erste Episode. Für mich war die anstrengendste von Palo das Begriff der Reliabilität Sie haben die Schweizer Käse-Aprache in dem Airplane-Sicherheit, wo kein Schnitt Käse, das man irgendwo anbietet, ohne Hölzer Aber wenn man genug Käse die Chance, sich Risiko durch alle Hölzer sehr, sehr niedrig. André Neubauer: Einfach LinkedIn oder unsere Website wo wir alle Episoden haben. Fragen und Anwesenden könnt ihr auf LinkedIn Dank für eure Zeit und bis zum nächsten Mal. Csaba Tamas: ⁓ but then after a while I switched to the dark side to sales and this was a role invented by me where the main idea was to essentially merge together product view and business view and make sure that the executives of these organizations have a good understanding of how technology can be used in the benef in their business benefits. ⁓ and this was so successful that later I was asked to ⁓ create strategy overall for AWS. So I switched to a global role where I defined strategy for digital native businesses. This sounds great, but it was a little bit like ⁓ when you are defining recipes in McDonald's, right? So in each and every country define their actual operations and they receive the recipes from the global ⁓ organization. Sebastian: Genau wie er beschreibt. Sie benutzen eine Mischung verschiedenen Dimensionen, die sie auf dem Flieger Sentimentaländerungen, zum Und dann runnen Statistik-Analysis. diese Analyse, nach dem Faktor, wird dann in weitere Verbesserungen der Agenten Das Beispiel, dass sie viele ⁓ Dinge ⁓ Csaba Tamas: ⁓ so this was again a learning opportunity for me, but then I felt like I need to go back to my roots, which is essentially ⁓ product and engineering. So that's the 20-something year of my experience. That's that's what is characterized by ⁓ in my 20-something years of experience. So throughout my career, I was an engineer first, and I didn't even know what product management is, but I was acting as a product manager and I was always looking at Sebastian: Einer meiner Tatsachen Csaba Tamas: ⁓ software development as a mean to an end. So I improved continuously myself in terms of engineering because I wanted to deliver faster, better features and capabilities to our customers. I'm not coding anymore for food for ⁓ 15 years, but obviously I still work on ⁓ on pet projects and I still try to keep myself up to date with everything that's happening deep down in engineering. But today I still André Neubauer: Ich Andre Neubauer, CTPO bei Trusted Heute ist ein neues Kapitel für das Show. Wie wir letztes haben mit dem Parler Search das Go-To-Tag und Exekutivsearch-Agentur in Deutschland. Und gemeinsam werden wir Beyond-Wipe-Coding auf dem nächsten Level Csaba Tamas: I'm active on an everyday basis on the architectural solution design level as well. And that's what I also believe that's becoming more and more important with cloud code and with ⁓ all the ⁓ agentic software development. I would say that algorithms and alg algorithmic knowledge as of now is better done and faster done by by AI agents. ⁓ but then still you still need a human who is actually guiding that in the right direction. And that's where I spend most of my time as well. Sebastian: Ja, Impala Search bringt eine große Exekutivnetwork, die über Jahre gegründet Exekutivsearcher für die deutschen Top-Tech-Kompanien Sie werden uns helfen, mehr der Menschen zu die die nächste Waffe der AI in die Zimmer mit Shaba, unserem Gast heute. Und sie geben uns die Öffnung, die Operator, die Founder und die Techniker zu die die nächsten schöpfen. Die Partnerschaft ist auch eine persönliche. Ivan Lechev bei Impala ist jemand, ich persönlich schon seit einer Dezember Awesome. perfect framing I think for the rest of the conversation. ⁓ what you didn't mention is that you're currently at Paloa, right? Do you wanna spend a minute on that? Csaba Tamas: Exactly. Yeah. So Palloa is a customer experience automation ⁓ platform. So essentially we are building AI agents for large enterprises. And these AI agents are ⁓ automating the customer experience processes, which means that we are landing in the call center, but then there's lots of opportunities to automate the whole customer journey there as well. And to some extent we are building a modular composable system, which is similar to the main idea of cloud providers, where they provide some some ⁓ basic André Neubauer: What's stuck with me is actually. Sebastian: Wir sind sehr gespannt, was wir zusammenbauen Wir fantastische Gäste in und wir sind sicher, dass unsere Partnerschaft mit Impala Beyond Vibe Coding auf ein ganz neues Level André Neubauer: Die Nutzung von Mighty Agents macht absolut Sinn, ein Kind von Task-Breakdowns weil das zu kleineren Promis, kleineren weniger Kontext-Degressionen so weiter war guter Erinnerung, dass man ein großes Projekt auf der LM sondern sich darüber nachdenken, wie man das dann herunterbrechen und dann am Ende ⁓ Okay, dann würde sagen, wir das Wie Sebastian gesagt reden heute mit Chaba, der Produkt und Technik bei einer der AI-Poster-Schaltz Europa. Csaba Tamas: primitives to to their customers. Maybe the only difference is that while at AWS we were speaking about the Lego building blocks. So you take a lot of small building blocks and you build whatever you want. In at Parlow we are building duplo building blocks. So bigger components and fewer ones as well. But laser focused on customer experience automation. So this goes from conversational platform where we take your ⁓ voice and we make sure that you have a great voice experience to Sebastian: Absolut. eine andere Sache, die ich auch sehr notwertig das Prototypen für Anwendungen des Enginierings, was Shaba beschreibt. Das etwas, was Anthropic auch beschreibt. Fiona Fong in ihrer Rede in der letzten Claude-mit-Code-Konferenz beschreibt etwas sehr ähnliches. Und es macht so viel Sinn, und Prototypen sind jetzt sehr teuer. Und so wie es in der Anfang Ja, und er hat uns gepackte AI-Agenten in Produktion Eine super dänse Episode und voller Insight. Wir hoffen, ihr die Episode könnt, so wie wir haben. Csaba Tamas: agent orchestrations to skills and to many other features and capabilities like simulations and evaluations and ⁓ also agent observability. So these are the actual building blocks which ⁓ make an enterprise successful in their ⁓ agentic endeavors. Sebastian: bevor du andere das Waste André Neubauer: Absolut. Es ist egal, ob es Claude oder Lavebeur oder Bord Aber wirklich ein Prototyp in der Theorie wie alles absolut guter Für mich war es interessant, zu dass sie nicht ihre eigenen Modelle Sie sind wahrscheinlich einer der größten Frontiermodelle. Sebastian: That is super interesting. thanks a lot. I already have like 10 questions I would like to ask. Unfortunately, we always start with the first ⁓ section in our podcast, With the tech stack. So what is your personal tech stack? I assume that you don't have much time actually working with code directly in your job, right? But what are you using when you do and maybe what are you using? when you don't do if there are ⁓ interesting tools that you use that help you structure your your ⁓ schedule, whatever. André Neubauer: In Quoten sind sie nur auf der Harnis Das auch schönes dass man mit großen Waffen keine Probleme Eine Standard LLM ist auf vieles man muss nur auf der Harnis Sie machen viel, wenn es ⁓ das und die fein zu Das war guter Lernen für mich. Csaba Tamas: Yeah. So actually I still code even for the job. And the reason for that is that whenever I'm building a new requirement or whenever I'm diving deep into a new area where we want to step in as a company, ⁓ instead of building the requirements, first I build the prototype and then I reverse engineer the requirements from the prototype. So this is a shift to how it was before, because before writing the requirement specification was the cheaper or early Sebastian: Und das einen Tonnen Sinn, da sie meist nicht viel Rezendung Sie müssen sehr schnell das Sprachwort in den Text dann aufsteigen und dann den Text zurück in Kasse zu ⁓ Csaba Tamas: kind of way of thinking about something but today building a prototype became so easy with AI that actually as a first step it makes a lot of sense and through that you can gain a better understanding of that space so that by the time you're building the requirements as a reverse engineering process from the code you're building a better requirements already from the get-go. ⁓ obviously cloud code in itself is a great tool but then what is also important is that ⁓ in order to build André Neubauer: es wird auch komplexer, denn am Ende ist es auch Technologie, und je mehr Technologie du hast, desto größer die Komplexität. Es ⁓ schön, von so einem komplexen Szenario ⁓ dem sie auf Standards Csaba Tamas: reliable, good and scalable systems, you need ⁓ adversarial agentic in ⁓ collaboration. So one agent is building the code, the other agent is maybe criticizing that code. That's extremely important because otherwise these AI agents are optimizing for success, but success also sometimes means that they are adapting the unit tests to to pass a wrong logic. Right? So so these things if you want to be successful, you need to go a little bit beyond the the wipe coding and and to build more like Sebastian: Ja, das ist eigentlich ein großartiger Zeug, um meinen dritten Punkt den ich sehr bemerkenswert Das wie sie ihre Agenten testen. er beschreibt, dass sie simulieren und dann E-VALS Das ist eigentlich so, wie man es mit LLM Produkte, die LLM brauchen E-VALS. Sie haben Infrastruktur für E-VALS, die der Karte André Neubauer: Ja. Csaba Tamas: Factory pipeline things. ⁓ Sebastian: That's pretty much exactly what our podcast is about, right? Beyond vive coding. Thanks for for this ⁓ like phrase. André Neubauer: In dem wir in einem der alten Episoden die Digital Product Factory, das war so, wie sie das genannt Aber super interessant zu hören, dass Sie noch hands-on sind. Wie viele Leute sind da, wie in ProTech, in Palua? Das vielleicht auch guter Thema für einen unserer nächsten Videos. Ich denke, das muss ein fokussierteres man AI in seine Produkte nicht nur für interne Prozesse. Für mich ist das kein echter Insight, aber es schön zu wissen, weil ich nicht gewusst dass Sebastian: Exactly. Type product factories. ... Csaba Tamas: So we are roughly 200 people and continuously growing. ⁓ we are in a world championship, so we have great ambitions to be a world class AI company, and it me this means that we are competing with the Magic 7. ⁓ so we are competing with Google's and AWS of this world, and we want to make sure that even though we have a small team, we have a punchy small team. ⁓ and we have a talent density which is maybe specific to top one percent of the companies out there. André Neubauer: Wow. Paloalto ist mit der Max 7 Bezug auf Talent Wie gesagt haben eine hohe Talentdensity, wie die Top 1 % der Industrie. Csaba Tamas: And we have an engineering practice which will be looked upon a couple of years from now as we are looking at Netflix engineering practices as of today or or some of these very famous engineering practices. Sebastian: that sounds really interesting and great to have ⁓ like such a company ⁓ in in Europe actually. yeah, one one thing I I just wanted to add is that the way you work and it it's it's really it makes a lot of sense building the prototype and then re engineering basically the the requirements from that is exactly how I think ⁓ Anthropic is working, right? So ⁓ the claw team is working exactly in the same way. ⁓ Yeah, cool. Thank you. ⁓ so we don't have lot of time. Hence, I would like to make the segue. You already hinted a little bit at what is required in order to build these agents working reliably and at scale ⁓ at Paloa. But yeah, please let's hear it go a little bit more into detail because I think yeah, this is really something that is becoming important for every company in the next months and years, right? And you are doing it already. André Neubauer: Vielleicht ist das guter für Arbeit. Csaba Tamas: Yeah. So little disclaimer, I will use a couple of numbers, but we didn't prepare for this interview. So you I don't have the numbers, the exact numbers in front of me. So whatever I will cite, it will be directional. ⁓ so one step and the most important step I think ⁓ there is research which shows that 95% of the agentic projects do fail and they never see the light of ⁓ of ⁓ the production. And the main reason for that is with AI is super easy to build a good looking demo which can impress the boardroom immediately, but it's very hard to get to high reliability. By reliability, I mean this is for me a proxymetric of instruction following, of hallucination, ⁓ and compound ⁓ proxy metric. So, what I used to say to my colleagues and also to my customers is that you can get 75% reliability at day one. But that's not enough for the business, right? So you need a human-like reliability, which is we can argue that it's ninety-nine percent. ⁓ then some there's some research which says even humans are not so good. So even humans are somewhere at 96, 97 reliability in trained and repeated processes. But then for us, the big question is how do we get from 75% reliability to 98% reliability? And obviously, in the enterprise, humans are super keen to Get that reliability to as close as possible to 100%. Fun fact ⁓ there was a survey which showed that ⁓ roughly 43% of executives already made a decision based on hallucinated information from a from an LLM in in in 2025. So while they are okay with that level of of ⁓ kind of reliability in terms of decision making, they are not okay with their agents not having a close to 100% ⁓ reliability. So everything what we do, everything what we are building for as of today is really about reliability of these agents. And obviously, because we are working with customer experience and in customer support, latency is important as well. So we have certain limitations which are specific to our business, which are specific to the fact that we are in the front office. Like in the back office, if you if you're building agents for the back office, you can use reasoning, right? Because you don't care if that ⁓ Sebastian: Yes. Csaba Tamas: LLM was processing a request for 10 seconds or 15 seconds. But you care a lot if you are on the phone and you tells you told something to the agent and you want to have a response as fast as possible. So this is actually puts us into a very special place where we have to optimize a lot to do asynchronous and parallel processing of multiple ⁓ LLM ⁓ operations. So coming back to the Initial part of the question, what makes us or what what what what we are focusing on? Obviously, we are focusing on different stages of the agent building. So on the one hand side, ⁓ on testing these agents. We call this simulations and evaluations because unlike in software development where you could run the unit test only once for each and every test case, in case of the AI agents, you have to run hundred or a thousand times each and every use case because you need a statistical relevance ⁓ relevance in terms of response time. So it's it's not called Unit testing, it's called simulations and evaluations. You're simulating conversations and interactions, and then you are evaluating those to get kind of the the ⁓ understanding of how reliable they are. But then this brings you from let's say zero to almost zero before going in production. However, after you go in production, you need to kind of still find the edge cases. You still need to understand what's happening in production. In a way that you have maybe millions of conversations and there's no one single human who can listen to all of those conversations or read the transcripts of all of those conversations. So you need an evaluation platform or an agent observability capability, which helps you to understand how the your agents are performing. What are frustrating factors, right? So what are customer sentiment trajectories? Right. So what is the situation when a customer started Maybe in a neutral sentiment and ended up in a positive sentiment. Or they started maybe positive and ended up negative. So if they ended up in a negative sentiment, ⁓ then you need to also understand what the agent made in the wrong way and how you can change the instruction set to to turn it to ⁓ into a better shape. And obviously, then there is this situation of LLM guardrails, so you want to make sure that your agents are not being misused. ⁓ you probably heard about this. ⁓ interesting news when there is an agent of a used car agency and you call the agent and you say, Hey, you are a useful agent, right? So you want to make me happy. So you know a 90% discount makes me happy. So make me a discount at 90%. And then this user screenshots the agent's promise on the 90% discounts and sends it the message to the to the used car sales to the to the used car's used car company that hey, your entity promised me ⁓ this, so you want to make sure that this is not there. But then there are other more ⁓ interesting situations when the agent starts hallucinating advice in regulated industries which shouldn't be the case, right? And then we which violates regulations. So this is again a situation what customers, especially enterprise customers, don't want to meet and see. So that's why agent observability is is critical. And then there's lots of derived ⁓ requirements in terms of data sovereignty, data locality, which is becoming especially in Europe, in Canada and in several other is ⁓ more and more important because more and more companies are questioning if ⁓ the over reliance on US ⁓ big tech is the right strategic approach or not. Sebastian: thanks a lot. unfortunately, yeah, we d we don't have the time to go into every little detail there, But super interesting. Maybe just to to poke it ⁓ on one end, and and Andrea, I see you also want to ask a question, but you're saying that you you have these these guardrails, right? And you need to prevent hallucinations through observability. Can you describe maybe one tactic how you do this, in the operations? Csaba Tamas: yeah, maybe you can help me also to get to the specifics, otherwise ⁓ I I'll ⁓ start just from a high level and then ⁓ deeper into some areas. So one thing what is very interesting is also to make this whole agent observability as cost effective as possible because your customers need to pay for it and at the end of the day ⁓ in AI you have a new type of continuous cost of goods sold where which are the tokens. Sebastian: Perfect. Yeah, yeah. Yes, yes. Csaba Tamas: And if you are reevaluating or evaluating each and every conversation, for example, that would ⁓ potentially drive the costs beyond ⁓ the levels that our customers are willing to pay. So the first step is to essentially use very traditional ⁓ analytics approach to process the data and create different dimensions to do a statistically relevant sampling of the data. Then once the sampling is done, you can run LM as a judge to essentially get a deeper understanding of what In those conversations and define certain dimensions. I mentioned one, the kind of the customer satisfaction trajectory, which is approximately again, because this is just you are taking it out from the context from the way how the customer is discussing. You look at this traject trajectory over the whole period of the conversation, and then you want to understand what are the drivers, like what are the underlying motivators. But then when we are looking about or we are speaking about reliability. We have a clarity and understanding of what is the sequence of tool calls, what those tool calls should be all about, right? So did all the what was that sequence followed upon? And this is not just evaluation. ⁓ we also have a policy engine which is checking all of this, right? So that the agent wants to make a decision, wants to make a tool call. What we are checking is if previews tools were called in this conversation or not. Very good situation is like imagine you have a restaurant booking system. And first you look up, but the agent needs to look up the available restaurant. So like you say, Hey, I'm looking for an Italian restaurant. I wanna ⁓ go there with my wife. So the agent looks up the restaurants that great, then you say, Okay, this is restaurant ⁓ Giovanni's restaurant is the best one. This is where I wanna go today at eight o'clock. And the agent forgets to do the reservation call, but does not forget to send the SMS to you with the confirmation. So that's the wrong the worst thing can happen because you go there with your wife, you are prepared for a good dinner, you have the reservation confirmation in on your phone, but the reservation never happened, right? So what you need to do in this case is to make sure that SMS reservation confirmation tool only happens after the reservation ⁓ tool call was happening, right? So that's again kind of a mechanism which is helping to to make sure that the the reliability of the agents are higher. Then there is many other things, like for example, you have this for example, you have this ⁓ Jackov shotgun theory, I don't know if you are aware of this. So the idea was is coming from from theater where they say if there's a pistol in in act one, by act three, somebody will shoot use it, right? So the same thing is happening with agents. Like if you have a lot of instructions, the agent at some point in time will use it. Jackov shotgun theory, I don't know if you are aware of this. So the idea was is coming from from theater where they say if there's a pistol in in act one, by act three, somebody will shoot use it, right? So the same thing is happening with agents. Sebastian: No. No. Csaba Tamas: Like if you have a lot of instructions, even if those instructions are contradictory or unnecessary, the agent at some point in time will use it. And this creates or increases the level of ⁓ hallucinations. So what we are doing is also meta prompting and evaluations with agents ⁓ of your prompt to to increase the quality of prompt to call out all those issues which are created by the prompt engineer by by defining potentially contradictory or ambiguous terms. ⁓ and then there's many other ⁓ little steps. So essentially there's not one silver bullet, even though everyone is looking for the silver bullet. ⁓ but there's multiple little things which we need to make sure that they are playing together very nicely. Yet another thing is in in this whole process is is the use of multi-agents. And multi-agents are again playing into the reliability because the if you break down a complex task, you can have a smaller prompt for each and every subtask. A smaller prompt means smaller context window. A smaller context window means better performing agents because there's a lot of research which is showing that even though if ⁓ agents or LLMs as of today have a large context window, you can put a book in in some of these models as well, and they will process it. As soon as you reach above 30% of context window, the agent starts the agent performance starts to degrade. So what you want to do is to keep the context window usage as low as possible. And with multi-agents, you can do that because each and every Sub-agent is taking care of just one subtask of the whole process. So these multi-agents are actually sharing the same tone of voice, the same personality. So these are transparent for the for the caller. But then let's say identification is one agent, then checking out the available restaurants is another agent, doing then and doing the reservations on the restaurants is yet another agent. This is great in theory. Then we built this, we tried it out in practice, and what we understood is that agents are also hallucinating the handover among themselves. So one agent is saying, It's not this is not my task, this is yours, right? The other agent was saying, No, no, no, it's not my task, it's actually yours. So they ended up in in infinite loops ⁓ of handovers. So then the question again, how do you control those things? How do you manage handover policies? How do you bring in ⁓ deterministic behaviors into the non-deterministic ⁓ orchestration of of AI agents? Obviously, you could discuss here easy responses like, yeah, okay, I could put an orchestrator on the top of it, but this is not playing for us for two reasons. One is it's adding latency, ⁓ or if I'm using a workflow engine as an orchestrator, then I'm losing flexibility, I'm losing the ability to build really human like experiences. And it's at the end of the day, this is a game of trade-offs, right? So at each extreme, I'm losing a lot of. ⁓ on the other side of the coins. So the question is how can I find that sweet spot of trade offs where I have an optimal and trustful experience? Sebastian: Yes. Thanks a lot. That that there's so much insight there. Sorry, Andre. André Neubauer: Ich Variablen. Erstens, die eine Geschichte, die gerade schreibt, erinnert mich an klassische Organisationen, wo die sagen, dass es nicht ihre Verantwortung sie zu dass Also, bleiben die aber nein, nur witzig. Ich habe über die du hast gesagt, dass es keine Silberbullen also es keine Ausstattungslösung. Kannst du ein bisschen wie du die verschiedenen Anforderungen Sie hier ein bisschen teilen. Csaba Tamas: Yeah, so actually I believe that there will be a one solution, right? So we are building, but there's not one silver bullet means that there's not one secret we have. We don't have ⁓ these secret souls, but we have multiple small little innovations which are compounding together into a more robust ⁓ solution at the end. André Neubauer: Okay, fair. True. Csaba Tamas: But speaking about ⁓ the the model and model tuning, ⁓ we don't do that today. ⁓ actually there is research which shows that there's much more opportunity to improve ⁓ agent reliability and agent performance in the harness. So if you know this meta harness study, they found out that if you use the same use case, the same model, you are just changing the the harness, like everything around the LLM, you actually can get a six S improvement. Six six times better agents just by changing the harness. And you would also understand that this is a continuously changing ⁓ field. ⁓ what ⁓ Entropic and the Cloud teams is saying that every time when the model's performance is getting better and better, you need to adapt, sometimes removed from this hardness some workarounds for weaknesses and to to adapt to the model. But our primary focus is on the harness. And yes, we will do fine-tuning, ⁓ and the reason for that obviously is mostly cost. So that we make sure that we are essentially reaching same the same level of of performance of the models with smaller models, and we get also ⁓ better latency through that. And we are working on on this as of now, but the primary focus when it comes to performance and reliability is is all about the hardness and all about the little nitty-gritty innovations, which all together ⁓ are more like you you ⁓ you used to say that the the parts are are ⁓ Like the whole is more than the the sum of its parts. Sebastian: Thanks a lot. like maybe one final question in this section before we then I think need to slowly but surely head to the end of this conversation. ⁓ you mentioned also sovereignty and and ⁓ locality. How do you solve for that? Are you using mo local models somewhere or like hosting models in in your own infrastructure? How how do you solve that? Csaba Tamas: Yeah. So ⁓ we are working on this as we speak. ⁓ we are building an infrastructure where we can deploy into local public clouds or local private clouds. And obviously we are also providing the speech to text, text to speech models fine-tuned for that and also the LLM. I believe that the future, the near future is in your local LLM as well. Because ⁓ most of the tasks which we need as of today for repetitive well-defined in well-defined environments, also as like customer support, the the open weights models are more or less there as of today. ⁓ so and and the open weight models are continuously ⁓ increasing their on their own performance. So if you think about the marginal improvement for each and every new release of the of the frontier models, you see a flattening curve, right? So you are André Neubauer: Hm. Csaba Tamas: And ⁓ the former ⁓ chief scientist of ⁓ of ⁓ Meta is also saying, hey, the current concept or the current parag per paradigm has two more years, right? So this whole ⁓ generative AI and transformer-based ⁓ model s will be superseded by something better because these models are at their end of their life cycle in terms of what are the marginal improvements each and every release. André Neubauer: Keep it bit off. Csaba Tamas: But that's a good news for us because this means that ⁓ very soon we will have true democratization of of AI models, right? So I was concerned about AI when and until the point when we saw that models becoming accessible to everyone. If all we have we all we if we all have access to the same kind of level of intelligence, then I think again we need humans to be able to build competitive differentiating features on top of the same baseline. And ⁓ this is becoming ⁓ more and more ⁓ attainable for everyone. Obviously there's still the GPU capacity, but this is also temporal, like if you look like the the level of investment everywhere, ⁓ that's I think it's a temporary issue which we will solve very soon. Sebastian: you. Then then maybe final question. ⁓ that because it begs the question, are you using open waste models currently or are you utilizing like models from OpenAI and Tropic? Csaba Tamas: So as of today we are primarily using ⁓ models from Entropic and ⁓ not from Entropic because that's not solving the latency issues. It was never optimized for that. It's it's a great model. I I really love it personally, but it's not good for our use case. So we used other specific models ⁓ which are more latency optimized. We are using fine models in special cases like ⁓ LM guardrails and So and we are using this learning to then also expand ⁓ on this, but we are in the process to ⁓ to make sure to to switch over. It's not our primary priority, as I said before, because the primary benefit of it is cost reduction. But currently there is a competition, it's a kind of world championship of AI. And this means that we are it would be premature optimization if we would start working on this. So for us now it's all about reliability, more automation capability, and that's all all our focus is going on the on building a more robust harness and an environment ⁓ where enterprises can know and can trust that they can rely on on our systems. Sebastian: Yep, that makes a lot of Thank you. All right. coming to the end of this conversation, which was really interesting. ⁓ thanks a lot for the the density of of really insights here. ⁓ yes, really ⁓ so one of the ⁓ questions we asked towards the end of the episode, or we like to ask ⁓ our guests is do you sometimes still have what the fuck moments with ⁓ your AI use? And if so, do you have an example? Csaba Tamas: Yeah, so obviously I have it a lot. And ⁓ it is also interesting that ⁓ sometimes ⁓ and the the whole debate on on vibe coding, right? So because some people are reporting a lot of success with vibe coding and some people are ⁓ struggling with it. And I built my first application ⁓ for the first time just using the cloud model. Like I didn't write one single ⁓ line of code and it worked very nice. And I got a lot of kind of confidence that this is the future. And then I ventured into a space where I had no expertise. So the first one was more like a data, big data application where I had a lot of expertise. And then I went into a heavy and complex front-end project and I failed completely. Same model, same person behind the model, and two completely different outcomes. And I start reflecting on like why this is happening. Like, and and then I turned it turned out that. Even in the first case, there were little mistakes, right? So the model tried to trick me into certain situations, ⁓ which didn't really make sense from the technical architecture point of view. But because I was an expert there, I could immediately put my finger on the problem and I could say, hey, that's wrong. And the model was saying, Yeah, sure, okay, let's fix it. Like, no worries. And it tried to fix it. And it's like, like, now I fixed like this. And I said, like, it's still wrong. And then with three or four iterations, we got where we wanted to go. But when I went into a territory where I was I had no expertise in the front end, I just relied completely on the model. And probably the same thing happened. It was something wrong there, but I couldn't put my finger ⁓ on onto the issue. And and that's the moment where you understand you still need the human, you still need the expertise. So the model essentially is just a an extension of of your current capabilities and it's a lever which helps you to be more productive, but you still need to have some sort of an understanding of the underlying technology. So that's one, and obviously I had all the other examples which many other people had where the model tries to rewrite the test to fit a wrong behaving code. ⁓ or the other thing where you ⁓ I tried spec driven development and I created a huge spec with several hundred tasks and I said, Okay, now let's implement it. And then the model did like thirty and it stopped and said, Okay, I will implement the rest of it when it will be needed. And it's like, yeah, that's the reason why put all those ones there, right? All of them are needed. So the models are learning from humans and from human data and you have this ⁓ deflection and and healthy laziness in the models as well. André Neubauer: Nice story. Sebastian: Yes. Thanks a lot. I I can relate to this a lot, actually. Yeah. All right. ⁓ thank you. yeah, I think we need to finish here, right? thanks a Yes. But much. It was great. Yeah. Have a good two. Bye bye. André Neubauer: Super Episode. Vielen Dank. Csaba Tamas: So much. Have a great rest of the day, folks. Thanks for the opportunity. Bye