speaker-0: Yeah. speaker-1: The big challenge and and kind of the bottleneck for us to kind of switch towards ⁓ fully open weight bottles is the infrastructure layer. Cause unless we would build this ourselves, we can get to the combination of speed, reliability, and caching that we need. speaker-0: So at the moment, actual learning that what's running in production with actual learners is completely hundred percent frontier models. speaker-1: One-on-one tutoring is the best way to learn. It w it it just wasn't available to the vast majority of people. So so that's what makes me super excited. It's bringing very high quality education to millions or hundreds of millions of people. speaker-0: At every moment of the teaching ⁓ process, the system has to decide what are the relevant behaviors for right now. Because if we cram all behaviors into the model at one time, that's over 120,000 tokens, it's it's massive. speaker-1: But if you do the math and you kind of start realizing one hour of AI tutoring can cost several dollars, you quickly realize that that means ten, twenty, thirty, forty million dollars on a hundred million dollar ri run rate just in kind of AI cost for the tutoring. speaker-2: I'm Kim Eisenberg, Superintelligence Editor-in-Chief with my co-founder and co-host Peter Thume. Today we have two guests from Datacamp, co-founder and CEO Jonathan Cornelison. I hope I pronounce it correctly. And Chief AI. Thank you very much. And Chief AI Officer Yusuf Saber. Welcome both. Thank you. speaker-1: Excited to be here. speaker-2: Great. So here's why we invited both of you. Many companies are still discussing, in theory, whether to build on frontier models or on open models. Datacamp is already doing it at scale. The AI tutor runs for millions of learners, and Datacamp pays those costs every month. Yusuf built that system at Optima. which data camp acquired. Jonathan runs the company, so between you, you can tell us what it really costs, what really breaks and what a company should choose. Jonathan, we will ask you about strategy, Yusuf, we will ask you about the technical default. So let's speaker-1: Go. All right, guys, Jonathan, ⁓ let's go back two years ⁓ and start where you did with the interview beginning in twenty twenty four, ⁓ where you warned that AI once it gets ⁓ going into production, there are gonna be three key problems that people face. ⁓ you can't control the costs, it's hard to change ⁓ systems if you become too dependent on your provider's models and and you can't really control how ⁓ you ha it's hard to control your systems or keep control over your systems. And so I'm curious now that you're running a tutoring business at scale, which of those has proven to be the most ⁓ difficult in reality? sure it's a ver it's a very good question. I would I would maybe add one thing in the context of the tutor. I think and this is mostly the Optima team and Yusuf who kind of went through this phase, but I think getting to a high high enough level of quality, ⁓ to excellent teaching, I think that that is kind of the first breakthrough. Like can AI actually solve this use case? For us, ⁓ the priority has kind of that still remains a priority. How do we build the absolute best AI teacher out there? ⁓ but for us the priority ⁓ has shifted towards cost quite a bit. ⁓ Just to give you a sense, there's more than 10 million hours of learning happening on Datacamp last year. ⁓ and and we're expecting to reach around a hundred million in ARR this year. ⁓ but if you do the math and you kind of start realizing you can like one hour of AI tutoring can cost several, several dollars. you quickly realize that. That means ten, twenty, thirty, forty million dollars on a hundred million dollar ri run rate just in kind of AI cost for the tutor. It it it it it it kind of paints picture how important cost becomes. And just to be clear, we don't yet have ten million hours of learning on the tutor. ⁓ that's on the platform as a whole. A lot of that engagement still is on kind of the datacamp original experience. ⁓ but our goal is to shift all of that engagement to the AI tutor. And so cost is is a huge bottleneck ⁓ to make that happen. And it's one of the reasons we we started looking at at open weight models. ⁓ And and I've always been a huge proponent of open source. So yeah. Yeah and obviously cost is is is the focus right now, but I think control is a close second for the simple reason that speaker-2: We all are. speaker-1: This is so strategically important ⁓ to us that it it it it creates all types of issues not to have control ultimately over the model side of the the house. Yeah, we we've had quite a few conversations with people talking about the rel the control of some aspect of the business, whatever business you're talking about, is really a huge issue. ⁓ and it's funny what you said about open source models being excited. I think right now. The vast majority of people are excited and there are some people who are very afraid. Yeah. Yeah. Yeah. Depending on the I'm definitely in the in the camp of the excitement. Let me ⁓ g given our timing constraints, let me pass it back to Kim so he can ⁓ move us on and through this conversation. speaker-2: Yeah, you know, I actually would would love to to keep keep the discussions about open sourceness and ⁓ because I'm really in favor of open source myself. But let's start actually with this simple question. ⁓ Jonathan, for people who don't know the product, right? What is the AI tutor and how is it different from normal online courses with videos and exercises? speaker-1: Sure. So so in the old world, I think I think there's been three phases in education. ⁓ the first phase in online education was really the the MOOCs, the massive open online courses where you had static video content that was brought online. So think about Coursera, Udemy, and they did an amazing job just bringing static content online. The second wave ⁓ was companies like Duolingo, Datacamp, ⁓ Code Academy, many others. The second wave was all about how do you use kind of traditional software engineering to build more engaging learning experiences. ⁓ and for Datacamp, that meant we focus on learners spending 70, 80, 90% of the time actively learning, doing exercises. But conceptually, those learners still have the exact same content they go through. And so what's different with the tutor is is you now we now we don't create actual kind of a a a a completed course. We have experts who create the ingredients of a course. And then the tutor will adapt the course to the individual learner or to the organization that learner is part of. And it enables so many things because it it means if if the tutor notices somebody's new to the subject and they struggle a little bit, the tutor can slow down. ⁓ or if s if the tutor notices somebody really gets it, they c the tutor can speed up, the tutor can make it relevant to that individual, by by understanding their role, the country they're in, and so on. And so so it becomes a super personalized experience. There's been an enormous amount of research over the years on kind of what delivers the best educational outcomes. And it's fairly well established that like one on one tutoring is the best way to learn. It w it just wasn't available. to the vast majority of people. So so that's that's what makes me super excited. It's bringing very high quality education to millions or hundreds of millions of people. ⁓ I'll stop there. speaker-2: So Yusuf, you build the system at Optima, right? So please explain how it works step by step. I mean you you you already introduced your product a little bit, but a student gives an answer. What happens next inside the system before the tutor decides what to say back actually? speaker-0: That's a very good question. So in in our discussion so far, we have been referring to high quality tutoring, right? So that the tutoring that actually helps us achieve those learning outcomes is ⁓ there's a lot of tutoring, right? So it has to be at a certain level of quality for those learning outcomes to be achieved. And if you look at a great tutor, like we all had great teachers over time, they do an incredible number of things really well. So the way they explain, the way they choose the right example for the person that they're talking to, the way they answer questions, which questions they choose to answer now versus hold on to, how they decide when to step in and help versus step back and let the learner kind of keep trying. So a great user does so many things and does them really well. So if you think of building that into an AI system, we think of these things as ⁓ behaviors. So we call them ⁓ tutor behaviors and there is a very large number of them. So we have over a hundred of those behaviors today. And at every moment of the ⁓ the teaching ⁓ process, the system has to decide what are the relevant behaviors for right now. Because if we cram all behaviors ⁓ into the model at one time, that's over 120,000 tokens. It's it's massive and the model just can't compose all those instructions. So once the learner responds at every moment The system will decide what are the most relevant behaviors for this very moment. And there are three tiers. So what we call ⁓ persistent behaviors, always active, like how to address the learner, how to communicate their behavior, they're always there. And then what we call triggered behaviors. So how to respond to a question is only relevant if the learner is asking a question, right? Or how if a learner expressed that they know something well or gave a great response. There's a behavior for we call adaptive coverage, which will decide. speed up and maybe skip some of the material. So these are ⁓ triggered behaviors. They're only needed when the trigger is is has happened. And then third, we call content-driven behaviors. So based on the nature of the content being taught, there is ⁓ there are behaviors that are are triggered. So is it now teaching something or is it now helping them practice something? Very, very different behaviors are needed. So once the learner responds, a process would run and decide what's the subset of behaviors that are most relevant for right now. Those will happen. And then there's a layer. ⁓ people have been using systems like Cloud Code would sometimes see it. We call them system reminders, which are essentially the system knows where is the tutor likely to make a mistake. And there are certain clues from this the session so far. And based on those, it will trigger a set of system reminders which help keep it on track. And then it will compose that with the learner's profile and with the content that's currently being taught. And then it based on all of that, it will deliver. a response. So imagine you know teacher behaviors composed with learner profile, ⁓ current content, a set of system reminders, and out of that the teacher will return back with a response to the learner. speaker-2: Fan fantastic. Really exciting. I I would actually move a little bit forward. So now that we're understanding how it works internally, I wanna be precise about two terms that are often used as if they mean the same thing. When when Datacam says open source AI, right? So in the beginning we already talked a little bit about open source, but do you do you mean fully open source system or do you mean open weight models in this regard? speaker-0: Yeah, so so the the answer is ⁓ open weights mostly, ⁓ but ⁓ open source ⁓ models are are definitely ⁓ on the table too. But open weights mostly ⁓ is where the the the tier of performance we look for ⁓ tends to be. Cool. speaker-2: And ⁓ which models are actually running in production today? ⁓ y you don't need to give me exact numbers, but I I would like to understand roughly how much runs on open weight models, how much runs on frontier models, and ⁓ why you divided that way actually. Yus Yusuf, I think you're the expert there, the AI expert. speaker-0: Yeah, yeah, yeah. Happy to happy to. ⁓ so at the moment we are ⁓ the that actual learning that the what's happen what's running in production with actual learners is completely hundred percent frontier models. We still haven't put ⁓ open weight models in production with with actual learners. ⁓ so today ⁓ 100% of that is ⁓ frontier models from the the top labs. speaker-2: Right. And ⁓ on top of that, which new open weight models have you tested for the tutor recently? Because there is I mean, ⁓ there is a huge debate actually happening about all the new releases, right? We read about it, Kim K free and and whatnot. So ⁓ what what have you tested? Gemma, Lama, Quen, Mistral or any others? ⁓ w where of them was better than the other models you already use and and where was it worse actually? speaker-0: Very ⁓ very good question. So so we have actually tested all of those. So all of those have been ⁓ have been part ⁓ of the process. ⁓ so so far Gemma ⁓ have has kind of outperformed so the the the thirty one the thirty one billion parameter dense model. But it's a very, very interesting model. So it's ⁓ it's a thirty one billion parameter model, which is tiny compared to You know, you have ⁓ multi-trillion models now, like 1.6 trillion with with Deep Seek, which was one of the models we tested. But it's a dense model. So all 31 billion parameters are always active, as opposed to let's say, you know, ⁓ Deepseek, the same the V4 version we tested. 1.6 trillion, but you know, about 50 billion are active at any one time. So ⁓ the the Gemma ⁓ 31B is 31 billion ⁓ parameters, so it's ⁓ significantly easier to host. and all 31 billion model Parameters are active, so we get comparable performance to some of the the very heavyweight models, ⁓ especially for for our use case. speaker-2: ⁓ interesting. But if this is so exciting because ⁓ when I attend at the Google IO, I had the opportunity to talk to the ⁓ developers ⁓ of of Jammer. So so this is like a really interesting insight ⁓ to me. So and yeah, I I think it's a great model. But to add on on this, do you send different learner questions to different models? For example, a cheaper open weight model for for simple cases and a frontier model for for difficult ones? And if yes, how does the system decide in real time? I mean there they're has to be like an a rotor in between, right? speaker-0: So today ⁓ all the usage of open models is ⁓ our internal testing. ⁓ so we've been running it so we have this simulated environment where we can play back ⁓ a lesson as if that like we take what happened in in real life, what happened with actual learners, and we rerun it ⁓ in our simulated environment as if it's happening now. So in essence it's it's doing teaching, but it's simulated teaching, but with real data, w exactly the responses we got from real real learners, ⁓ not ⁓ but not with actual real learners. So our aim is ⁓ sometime at the end of this year we will start to actually have some of that running in production. ⁓ we can talk a bit about some of the the challenges that we need to work through to to get there. ⁓ but as of today, 100% of the learning is on ⁓ frontier models. The actual speaker-1: And Kim, maybe maybe just to to kind of give you a little bit of context on the why. ⁓ because like if we if you look at our evals from a quality perspective, Gemma is actually beating the the frontier models that we ⁓ that we are using today. so if we could we we would switch over, ⁓ if only for the quality. On top of that, if you think about the price kind of the theoretically kind of achievable price you would pay it would be way lower ⁓ for for Gemma for relative to what we're using today. ⁓ the big challenge and and kind of the bottleneck for us to to kind of switch towards ⁓ fully open weight models is the infrastructure layer. ⁓ because unless we would build this ourselves ⁓ we can get to the combination of speed, reliability and caching that we need. ⁓ and we've been we've been in touch with several vendors. We're kind of figuring this out. We would rather not build this ourselves, ⁓ because we feel like eventually it's not probably not gonna be a core competency. We wanna rely on on kind of umf infrastructure focused vendors in all likelihood. But that's the bottleneck. ⁓ and then there's some questions from some enterprise clients, but the vast majority of our engagements I think would be comfortable with kind of a Gemma for as long as it's hosted in in s certain certain regions. ⁓ but the bottleneck really is the infrastructure layer. It's not the model quality. ⁓ I just wanted to be clear about that because if we could we would switch I think tomorrow. ⁓ if seriously it would it would be so much cheaper. I have one quick question that's wasn't on the plan, but I want to ask it of both of you guys, whoever wants to answer it. When you say we'd switch tomorrow, what does that actually mean operationally? Like if you what is the process of saying, okay, we actually are excited about this, now we have to evaluate it. Is that like a a one week process or a six week process or a quarter long process? How long does switching actually take and you know, for for excited operators listening? speaker-0: Yeah. ⁓ so the f the first the first line is ⁓ evaluation. So if ⁓ a team that has ⁓ a solid, comprehensive, reliable evaluation setup, ⁓ switching over can be a matter of like hours or or days, not not not longer. So ⁓ we we we have this eval system which is deeply comprehensive. ⁓ it it run it takes hours to to run across so you know these hundred ⁓ plus behaviors we mentioned, each one has what we call a set of expectations. So if the model meets those expectations for each behavior, then it passes the bar for us. And then we have, you know, hundreds of what we call like fixtures. So these are fixed instances from actual lessons. And we know what the correct behavior in this instance was. Like for instance, ⁓ when a learner asks a question, their the teacher should to do one of four things. If it's completely irrelevant, like the learner will ask something about like a political question about the weather. So it should not engage. It should say, you know, we stay on topic and should not even, you know, get into that at all. ⁓ second, if the learner asks something that's super relevant to this moment, it should say, this question, if I don't address it, their learning is impaired. I have to address this now. Third is something that is coming later in the lesson will be addressed properly. So it should appreciate their curiosity and give them something, but not derail the lesson. Or something that's relevant but totally beyond the scope of the course. So it will you know, give them something to kind of keep them ⁓ asking questions and being curious, but try not to derail the lesson too much. So a model can make a mistake. So it can easily g mistake one case for for the other. So we have about maybe 50 ⁓ test ⁓ cases that make sure these expectations we have four expectations here are all met. So when we try a new model, this is for example we we run those. So there are ⁓ four expectations, about fifty cases. And each is run five times. And then it has to and there you can check does it pass any one of the five times or does it pass all of the five times. So pass all of the five times is the standard. So imagine running those. It it if you multiply those numbers, they quickly get into the thousands. We run all of them. And we do that for all of the models. So it could take maybe a few hours and we can test the model out and we get an estimate of how it performs and where it fails. Like you know, we're We love Gemma, it's it's been ⁓ outperforming, but on certain things it's actually underperforming in certain aspects of of that. And some of them are quite important. ⁓ we can work on them. So our prompts were optimized for you know the the frontier models. So Gemma has a bit of a disadvantage. The model the prompts were never optimized for it, and it's still winning on in on on many cases. So we so that's that's is the big part that enables ⁓ a cutover. So you'd run that. The model passes, you quickly get a sense of of cost, speed, in addition, of course, to task performance. ⁓ so initi internally for our ⁓ admins, ⁓ they can they have a drop down on the right side of the platform where they can switch to any model and continue learning. So it's already in production, it's already integrated, it's in the infrastructure. So switching that con is a config that we can make and learners would be learning on a new model ⁓ in in a few minutes. speaker-1: Thank you, Youssef. I'm sorry, Kim's gonna kill me if we don't get to the other question. So speaker-2: No, it's a great question, Peter. But ⁓ yeah, so Jonathan speaker-1: Maybe if I could if I can add one one more thing. ⁓ it's also just strategically it's important to us to make it fairly easy to switch between models. ⁓ we're still in the rollout phase, but we are live with some of our enterprise customers ⁓ with the AI Tutor. ⁓ and it's actually getting enabled very fast now within our ⁓ enterprise customer base. But there are some very large ⁓ organizations, think of financial institutions, governments that have a list of approved models they're allowed to work with. and it might be that the optimal model that we choose for our broader audience is actually not necessarily on that list. ⁓ and so so it's core functionality in in our mind to to have an ability to switch ⁓ kind of ⁓ the the the model layer ⁓ between different clients and and over time as well. speaker-2: Okay. Mm exciting. But ⁓ I would still like to focus a little bit more on the on the on the product, right? So we we talked a bit about the models behind it, but ⁓ looking at the model itself, how do you measure whether the tutor is actually good? And by good I mean ⁓ four things. Is is the answer factually correct? Is it deductly good? Is it personalized? And did the learner actually understand? How do you test each of These factors. speaker-0: Yeah. So ⁓ I think I'll I'll borrow part of the the conversation we just had. So the evaluation system we discussed ⁓ a few minutes ago, ⁓ that's our core mechanism. So ⁓ we look at lear like teaching as a set of behaviors, and each of those behaviors, ⁓ whether the model is or or the system is able to deliver on them, we represent that with what we call expectations. So a behavior would have maybe anywhere from maybe five to maybe twenty expectations. So what do we expect to happen so that we can say this the the the model of the system exhibits this behavior to the degree that we want. So that's one part of the system. We we define the behaviors and we define the set of expectations that must be met for this behavior. We say that the model is truly exhibiting it. And then second is the actual test cases. So there's a ⁓ we From actual learning. So for example, one one of the behaviors we have is that the the model will manage practice right. So manage practice right means what? Means ⁓ it would when it's sharing the question, it would not give a hint. So models are trained to be very helpful, right? So when when giving the question to the learner, it would then proceed and add a paragraph after the question that essentially gives away the answer and just spoils the opportunity to practice for the learner. So that it it fails that expectation. The expectation is to deliver the question without giving any spoilers ⁓ that that kind of give away the answer. ⁓ on the other side, a model that lets the learner struggle way beyond the struggle was productive is also failing. So it at some point needs to say, okay, this is a moment where it can ⁓ start to provide hints. ⁓ another thing is some learners want to keep trying and some actually are got to a point where they'd rather, you know, get some help. So we have it that the learner chooses whether we give them a hint or share the solution or not. So the model doesn't just volunteer the solution, it asks them, ⁓ you have tried this long enough. If you'd like, I can share the solution with you. Because this prevents cases where ⁓ learners who who want to keep trying, they say you spoiled this for me, I was just very close to it, I I almost got it. So we want them to make the choice. So a model that would not follow that, that failed an expectation. So there are like Over a thousand of those that ⁓ are tested. And each is tested with ⁓ anywhere from three to about five ⁓ different tests that that test the same ⁓ expectation. And we run each ⁓ up to five times. So that's you can imagine the number of runs because sometimes the model would get it right by chance. So we have to make sure it's and we always have a baseline ⁓ overall and for every one of those expectations, and every change we make, we run the whole suite against that. We have three tiers. So this the first tier runs only the critical ones. It runs in a few minutes and its cost is like a few dollars. So we run that first. If that fails, no we can't proceed. Like that's critical. If that passes, only then we run the next one. And before like full deployment, we run the complete ⁓ suite of ⁓ of expectations. speaker-1: On top on top of that, if I if I can add two two more things ⁓ which I think are are worth mentioning in this context. ⁓ so once once kind of the tutor is teaching live in production, ⁓ and and this is one of the things we're super excited about, we now have over three hundred thousand learners that have already spent quite a bit of time with with the tutor. We expect that to to scale to millions in the next six to twelve months. ⁓ But that's an enormous amount of kind of interaction data between the tutor and the learner. and we already have review systems that kind of understand like when is there friction here, when is there something happening that maybe it's deliberate, maybe may maybe it's the right thing to do where the learner is struggling. But oftentimes it's an opportunity as a teacher to understand like what could have been done better. So you have an AI review system. ⁓ and this is where I think over time there's going to be a massive data advantage because as we kind of see all these learner interactions, we can understand what works, what doesn't work, almost in in in near real time. ⁓ and we can adapt either the tutor system or sometimes we can adapt the underlying content to be more effective. ⁓ and then the last part, and this is something we don't do yet, is you can kind of look at the data set exposed and and maybe have external benchmarks on kind of like standardized tests people do at the end of of some of these programs or certification tests. ⁓ and I think there's a tremendous amount of academic research that that ⁓ would be incredibly interesting and valuable. ⁓ so if there's if there's any academics listening or reading this, ⁓ we'd love to to set up some collaborations. speaker-0: We we have lot. Yeah. Yeah. And if I might add a tiny little thing here, most of those expectations, like I'd say about sixty percent, they were not things we came up with. They were discovered through the mechanism that Joe just described. So we have a system called Sentinel and its job is to detect any friction in the process. And ⁓ the the bigger majority of those expectations came from that. So we think of it like, you know, any engineers ⁓ the audience would relate to test driven development. So we like to think of it as like evaluation driven development. So the teacher like drops the ball somewhere, we see it through the friction lens, and we create an expectation for it and then we fix it and make sure it passes that expectation so the teacher doesn't lapse in that way after that again. speaker-2: ⁓ I'm just curious. ⁓ I mean y three hundred thousand ⁓ learner is is a lot. ⁓ w how how was the feedback so far? I'm just curious. ⁓ what was the most ⁓ positive feedback and where were they kinda critical? speaker-1: ⁓ so so the feedback overall has been very positive. ⁓ we're we're a very data-driven company, as you can imagine. ⁓ and so we we we've really kind of we've we've done A B tests, we've looked at the data in every possible way. and and maybe a couple of data points that that are the most convincing. ⁓ so Datacamp on the consumer side has a freemium model ⁓ the first kind of twenty minutes to to sixty minutes. of every learning experience is is free. If you look at the f the first twenty to sixty minutes spent with the tutor versus in the old experience, the conversion rate to become a subscriber is significantly higher in cut for the tutor experience. ⁓ the other thing that I found the satisfaction ratings are also higher. And and the other thing I found really fascinating is ⁓ engagement tends to be higher but your kind of distribution shifts ⁓ Because what happens is you have learners who already know the topic and they they are able to move through ⁓ and get to the the learning objectives faster. but then there's learners who who who kind of struggle with the material or who are very curious. And in either case, there's a lot of questions, a lot of back and forth. And so they t tend to spend way more time, but they end the course actually deeply understanding. kind of the learning objectives and and what what what they set out to learn. ⁓ and so this is very encouraging. So we don't have formal academic research yet that this is actually better, but you c we have enough data that makes us strongly convinced that like didactically this is not just more satisfying, not more just not more engaging, but also just more effective. speaker-2: I mean I mean it it makes totally sense, right? So so looking ⁓ at how schools usually works. You you have like one teacher who has to educate the whole class at the same time so in a way to generalize. And you all I mean we all went to school and we know that there were ⁓ fast learner and ⁓ slower ones. So I mean making use of this technical revolution where NAI that has all the time it needs to teach like new stuff, new learnings is a fantastic way to to really make use of this yeah, intelligent this raw intelligence, right? So it's so fascinating and it's so exciting at the same time. It's it's like a whole revolution in the education system, right? It's it's fantastic. speaker-1: Abs absolutely. I I th I think one of the motivations, I think both for myself as well as for Yusuf to do this is the frustration ⁓ like I had with the educational system of like you're forced to be there. ⁓ there's it's like a factory system. You move through the system at the same p pace and sometimes you're bored. A lot of times you're bored, sometimes you're lost, but like it moves at the same speed for everyone and it's quite Both as a as a learner but also as a teacher, to be honest, 'cause it was quite frustrating as a teacher to see you see it. Like some students are lost and some students are are are are kind of bored. speaker-2: Because it often even kills ⁓ curiosity and excitement when you go to school. So having like a a personal tutor in AI intelligence who ⁓ supports your curiosity, who engages in you, it's it's a fantastic new way of of really learning and supports curiosity, right? So this is what really ⁓ excites me. I I have a son myself, right? And looking into the very bright future, this is so great. And I I really love to see and hear it. it right so yeah it's it's fantastic. But but yeah so but but moving on, right? So ⁓ l a qu a quick one on on to ⁓ the the the product and and the way speaker-1: ⁓ absolutely. speaker-2: ⁓ yeah, ⁓ the schoolers learn. So usually there is a trade-off when it comes to speed, right? A faster model can give weaker explanation, although it really depends on the underlying hardware, right? So I I know we we have like inference chips, so even frontier models can now serve very fast inference. But usually ⁓ it it depends on ⁓ the quality you decide. Bigger models, slower, ⁓ smaller models, quicker answer. ⁓ how how do you handle ⁓ this ⁓ process of yeah of ⁓ of of speed and quality. Wha what's what's your way of handling it? I'm I'm just curious. speaker-0: That's ⁓ that's a very very good question. So ⁓ if I if I zoom out, ⁓ you can think of like three major ⁓ p parts to this framework. So first, what you said, definitely the ⁓ bigger models, ⁓ like a model with a larger number of active parameters would just need more time to to respond. So your time to first token ⁓ would would be slower. But in a real production system, that's usually the smaller ⁓ piece. You have two other pieces that contribute so much more to latency. First is your own backend. So ⁓ like your basic database reads, ⁓ how your ⁓ infrastructure is set up, that can be a massive ⁓ contributor to latency before like models are engaged or after models are engaged, how you process the structured output from models. So there is a lot of latency that goes in the system. And third is the the architecture of your AI engineering system. So one one simple example. If your model makes many model calls, that like multiplies latency by by a factor. So every every sorry ⁓ makes multiple tool calls that significantly uses the time. So every tool call essentially is one full round trip that you have to wait for. So when we are building these systems, we look at all three. And usually ⁓ the the s the second and the third are where most of the latency is. So roughly for us, for example, now about ⁓ 50% of the latency is in our own software and 50% of the latency is with the model. And when we are doing our like latency reduction projects, we find that we can bring a lot of performance improvement on the software side. But where most of the latency comes from is in the AI engineering side. So for example, for us now, we have no model calls where ⁓ sorry, no tool calls where the model has to wait for a tool's response. So we replace that with a mechanism we call system request. So if the model needs something done, it will trigger a system request and move on. ⁓ And if you think of we have this asymmetry in how time works for our learners. So the time for for a couple of seconds, it's extremely time sensitive. So the learner sends ⁓ a response, they're sends a mess, they're waiting for a response. So here's extremely time sensitive. And then the learner takes time to process, to write their answer. And then we have like sometimes 20, 30 seconds of there's no one waiting on anything. So we use that time to do a lot of work. So we like we cache what then what would the learner likely say and we pre prepare responses in in most of the cases. You can guess with very high reliability. So we use those twenty to sometimes thirty seconds to do a lot of the work that takes time. So that when the learner does respond, most of our learners ⁓ have audio on. So they hear as well as read. And then to generate like text to speech, that also takes time. So how do we so we we first take the the the shortest complete sentence and transcribe that with very fast systems very quickly so they start hearing. So as soon as like what would be ⁓ like time to first token is when the LM responded, but then now we have to convert that to speech. So that's an added latency. So we take like the the smallest possible part of the of the response, we transcribe that so that what the learner ⁓ they can start hearing, and then the rest is done in real time. as as the as streaming happens. So streaming and transcribing happening and lockstep so that there's no delay, there's there's no waiting. So my I think the core of that is I think this ⁓ this last part part, the architecture of the system and how it works, I think is where most of the latency gains are. So you can still use a very powerful model ⁓ and ⁓ iterate here so that you can bring down a lot of the latency. speaker-2: Okay, exciting. Mm and it makes totally sense. ⁓ Peter, ⁓ I would ask one follow-up question and then pass it on to you. I have so many more questions, but I I'd like to focus on on on the really important one for for this time. Because so I mean I I I'm Big and for myself and probably for you as well, we're living in c some kind of AI bubble. So we're we know about the frontier. But most of the people ⁓ they hear about AI and they're they are worried about, for example, hallucinations, right? And especially when it comes to education, the hallucinations would be like the the biggest fear of them. So in in most products, ⁓ a wrong answer from an AI is is usually only annoying, but in in a tutor it it means a student learns something that is false, right? ⁓ I mean we already digged a little bit into this, but what do you do to stop the model from explaining something incorrectly but confidently? So ⁓ I mean if if you can can give us a g ⁓ a a brief explanation about this, this this would really switch off the the ⁓ the anxiety I I would assume for many people that are like I said, not within our AI bubble, right? So yeah. speaker-1: Yeah, the the the the high level answer is there's we mentioned in the beginning there's ingredients to the course. You can think of that as the the kind of factual reality and and the ground truth. ⁓ and so the tutor has quite strong guardrails to go back to the ground truth and to be quite cautious to kind of go to to and explain anything beyond that. ⁓ so we we did we put quite a lot of thought into kind of making sure we understand what should be in there. ⁓ Yusuf, any anything you would add to that? speaker-0: No, so I I think that was kind of problem number one. That's the first problem we had to solve. Because to your point, a learner is is vulnerable, right? They they the fact that they're taking this course is most likely they'll know this material. So it's very easy. Like when you're using like Cloud or or Chat or or Chat GPT, a lot of the time you're having it work with you and your you know s the space that you know. So you can quickly detect if something is off, you can push back, but a learner has far less facilities when it comes to that. So we had to be very confident that we are not, you know, ⁓ putting a vulnerable learner in a place where they are kind of being misled. And the the first part was the model, we want it to use its intelligence, but not its knowledge. So ⁓ this is a very core part for guardrails that we wanted to use the intelligence it has to be a good teacher, but not to use the knowledge it has. All knowledge comes from what we provide. And given the the very short latency window and ⁓ a bunch of other concerns, there's ⁓ it's not allowed to like search the web. So given that some of that search the web and it's not allowed to use its knowledge, we have to provide all of the necessary knowledge up front. So there is a bit of work where it has the core knowledge for the lesson plus a s ⁓ kind of a ⁓ a repo of supplementary knowledge that if someone is learning this, these this is kind of the the surrounding space of knowledge where if someone is learning this, if they're curious, they typically go around into this space. So it has access to to that as well. And it can request additional knowledge from the system for in the rest of the course. So if someone wants to learn something that's down the the the line in in a further in a lesson in the future or even a different course, it can request that knowledge and get it from the system so it can share, share some of that. But the second and more important thing in our opinion is it will actually say, I don't know. So like what one of the major ⁓ one of the major ⁓ expectations that we test, actually this is one that ⁓ we need to work on to get Gemma because Gemma was actually underperforming on this one, is it was slightly less likely to say I don't know than the frontier models we were using. So it has to check, do I have an answer to this? If I don't have a reliable answer, I will say I don't know. speaker-2: Yeah. The first time a model really was like, hey, I d I just don't know the answer. ⁓ this this may i be incorrect and the whole community was like mind blown because this was like b the the ⁓ it's it's like ⁓ You know, as we were saying, it's trained to give an answer, but being honest about ⁓ n not knowing something is is ⁓ is so important in a way, right? So yeah. Sorry, I did interrupt you but yeah, this is so important. Yeah. Mm mm. speaker-1: I gotta say it's very hard and realizing that we have eight minutes and I wanna talk for another 80 minutes. ⁓ I'm gonna skip through the whole next set of stuff and go on to the next section of things I wanted to talk about with you guys and get straight to ⁓ a question I think is really important. ⁓ on your website it says that people learn at six times the speed ⁓ they would normally learn. and l let's say I'm a buyer, what I would wanna know is is how do you how is that number calculated? And by when I get to the end of that, when I'm, you know, I'm consuming one-sixth the amount of time to learn that information, ⁓ am I learning more or am I just processing more stuff? Like am I actually le walking away with more of it in my head? I think you kind of answered this already, but I'd really like to hear that specific answer, I guess Jonathan. Yeah, Peter, I th I think if I'm not mistaken, I think you're referring to the The claim we I know we make is ⁓ you our course completion rates are are three to six times higher than kind of the traditional learning platforms. I I may have I may have mis misspoken or misphrased that. So I don't want to run to misstate your own claim. Sorry. No no worries, no worries at all. ⁓ but ⁓ but essentially I think the the ⁓ the fundamental problem in online education has historically always been that If you're just passively consuming information, if you're just watching videos, yeah, it just tends to be like quite boring and not very engaging. And so so the key problem is really engagement. ⁓ and I think the the the key challenge is how do you make it really engaging while ensuring people are still learning? Because you have some companies that ⁓ take it so far in terms of gamification, you're barely learning. And that's not That's not us. Like we have a lot of enterprise clients. They actually care about the skills they learn. We have a lot of consumers who come to Datacamp because they want to advance in their career, because they want to get a new job or get get a promotion. So the actual learning really matters to us. But we also face the same problem that engagement is is the real challenge. And so the fact that it's interactive and the people go back and forth was always part of what Datacamp does. And that was creating a lift of anywhere between three to six. ⁓ The exciting thing is that with AI Tutor, it feels like w you can go further and for people who are fast learners who who who already might know something about a topic, you can also reduce the time they need to get to their kind of learning objectives. So so you can have the best of both worlds due to personalization. I think that's incredibly exciting. Because it w before kind of Jenny I entered the scene, we thought about this and we tried a lot of things. And the technology just wasn't there to be able to do this. ⁓ but now it is. Now you can actually create a more engaging experience that's also more effective per per time unit. But it would be too early tell. And so it's not it's it's about an insurance policy or you know, insuring yourself against government ⁓ intervention rather than Then helping people actually get smarter or better on a topic. Totally. And I think it it it's never mattered more than today, because the the reality is AI is changing the skills people need faster than probably any technology historically has. So there's a tremendous amount of reskilling and upskilling that needs to happen. And so it's not just about ticking a box in compliance, it's it's really about like can you create a behavior change Can you get people to the other side and c can you kind of can you make them more productive and be more competitive as an organization? Right. Kim, I'll throw it back to you 'cause I know we're running out of time. speaker-2: Yeah you know, I I would like to close ⁓ the last questions with a little bit more practical view. So I'm curious what is the worst mistake that tutor has made in production and what did you change afterwards? I'm just so curious. speaker-0: Worst mistake. So I think ⁓ for us the worst mistakes are when it derays the learning process. ⁓ so for instance, ⁓ we have ⁓ courses where there is a virtual machine that needs to start. So for example, you want to learn ⁓ a tool like you know Power BI or Cloud Cowork. So we want the learner to experience the actual tool, not just watch videos of it. And the way we do that is we start a virtual machine. And we run the actual tool in the virtual machine. Now we trust the AI user to do a lot of things here. We trust it to trigger when the virtual machine should start. It decides what files to load on your behalf so that the tool opens with the right files. ⁓ It gets constant screenshots of the virtual machine. So it decides if you're learning the right thing. So it would say now let's do X and then it opens a file that does Y. It's not the right file that the learner should should learn. ⁓ or it's a point where okay, the virtual machine needs to start now and it says here's the virtual machine, but it doesn't start. So learner is like, okay, where's the virtual machine? So those were some of the issues. And these are like especially painful because they break the learning. The learning is like w what what happens now? The good thing is it or it ⁓ immediately adapts. So once the learner says, I don't see it, it will quickly ⁓ self-correct and fix it, but we don't want this to happen in the first place. So this is where we rely heavily on what's perhaps one of the most powerful techniques in all of AI engineering, which is system reminders. So the system knows that the virtual machine should come now and it will remind us to actually open it. And in cases where we are we know with high certainty we actually have the system automatically open it. In cases where we think it's it's highly likely, we we still need intelligence, we need the model to make a call. We tell it it's very highly likely that the virtual machine should be started now. And when it has that reminder, it's exceptionally less likely to make mistakes of this sort. So all of these mistakes where it's likely to derail the lesson are likely cases where there's a clue in the session that code can detect, but you can't do it with code only because it's not a hundred percent. But a system reminder kind of bridges the gap. The model gets a strong reminder and then it doesn't make those mistakes. So those are usually the the worst kinds of mistakes because they really break the the experience. And we came up with this whole system reminder mechanism, which we rely on now for quite a few things to address exactly that. speaker-2: Okay. ⁓ can I ask one last question, Jonathan? ⁓ I know we're so very short on time, but just very quickly, what was the best feedback you ever had? What was the best ⁓ positive outcome? And where do you see yourself and ⁓ the company in one to two years? speaker-1: All right. ⁓ I have another interview, so so I really have to jump. But the short the the short answer is the thing that excites me the most is when people say this feels like a real teacher sitting next to me. It's the best educational experience I've ever had. that's ⁓ what we set out to do. The vision is that we are very focused on data and AI skills right now because that's what what the big gap where the big gap is. But fundamentally we're building the best AI teacher. ⁓ and so if you think about w what's our goal for the next few years, it's building the system that can can teach any any type of skill ⁓ for professionals initially and then kind of beyond. ⁓ so so I feel like that's an incredibly exciting ⁓ thing to work on for for the next few years. speaker-2: Perfect. That is a fantastic closing. Thank you so much guys. This was excellent. We're going to share obviously all the information in our social medias and we highly suggest everyone is going to check it out. So thank you so so much. It was such a pleasure talking to you guys. speaker-1: Thank you so much. We couldn't ask. Jonathan, I hope we get to meet you somewhere here in Manhattan by Likewise, likewise. Let's do it. Thank you. Bye everyone. speaker-0: Thank you for having us. Mm-hmm. speaker-2: Yeah. speaker-0: They're