🔒 8 more in the full analysis
Build custom AI agents for enterprise customers by understanding a specific operational use case, making architecture trade-offs, and deploying a system with controlled voice, language-model, tool-calling, reliability, and evaluation layers.
Full plans for 1 idea. Inquire for details →
🔒 21 more in the full analysis
Searchable transcript of ⏭️ Forward Deployed: Voice AI on what works in 2026 — Latent Space (36:31). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by Latent Space. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 All right, we're in the remote studio. This is a special one because we're launching a fourth podcast on Dayton Space. Uh people don't uh necessarily keep track, but we cover different things. Basil uh came across my radar because you're doing all these like dinners and gatherings and panels and podcasts on for deployed engineering. You hosted the FD track at AIE and it did super well.
00:21 So, welcome to the pod. >> Yeah, thanks. >> So, you've been running your podcast for a while. What what made you decide to focus on FDE and you know what's a what's your typical self intro? >> Yeah, so I started as a product manager at in I was specifically working on credit karma for a couple years. I worked at a small venture studio after that and then I ended up starting a consulting business called Exoflop Labs where we were working with like retailers, insurance companies, that sort of thing just building agents
00:48 essentially. This was like 2 years ago. So it just felt like a natural extension of a lot of the product work that I was doing earlier. It was like, hey, I'm working with customers. They have a specific type of project that they want built for some use case and I'm going to understand what they want. I'm going to make trade-offs and I'm going to help them build that.
01:04 So, yeah, I just felt like a natural extension of what I was doing. And so, back in January, we started working with a couple private equity firms and a lot of people were just like, "Hey, like I'm following on what's going on on Twitter, what's going on LinkedIn. I don't know what's marketing BS and what's not." So, I was like, "Why not just bring on people who are working on the cool stuff at cool companies?
01:23 we will talk about what they're working on like we'll do a deep dive and we'll record it we'll post it and hopefully people can take learnings away that they can apply to their businesses and that's why I ended up starting doing a lot of the fireside panels and the podcasts that I've been doing since then >> can you give like rattle off just your greatest hits like what you've covered >> yeah so our very first episode was on the future of agentic engineering so we brought on companies like factory cognition composio
01:49 sound grap we've done panels on voice agents which is actually the one that we're going to be showing here. We've done some on computer use agents. We've done agents in the enterprise. So, we've done a ton of stuff just to get nitty-gritty into the details of a lot of this stuff. >> Yeah. And I guess you know, you picked the for for our sort of first feature, you picked the voice agents one.
02:08 What are we about to listen to and what stood out in particular? >> Yeah. So, we brought on some people from Decagon, Vapy, Retail, Daily, and a company called Smallest AI. Basically, we just talked about like what is the state-of-the-art in building voice agents cuz that's a very hot use case. I think uh Sierra even talked about this is this is like one of the most competitive markets in AI right now is like building voice agents.
02:31 And so I thought it would be useful to talk about how they're actually built and some of the things that you know engineers who are building in this space still have to contend with. So for example, we talk about how like the state-of-the-art right now in building voice agents is a cascaded pipeline. So we talk about how it's like a three-step process.
02:47 you go like speech to text LLM then text to speech and we talk about like why is that versus why don't we just have a voice to voice model like why don't we use that like the reality is those just are not super reliable right now we also talk about like the trade-offs that you have to make when you're building voice agents so you have to trade off like the intelligence of responses that you're getting versus the latency so you can get very intelligent responses but you're also going to trade off it's going to be a
03:10 slower response and so for different use cases that might be good or bad we also talk about like reliability so there are some LLM companies that may not be like super reliable in terms of like their infrastructure. And so like what do you do when opus goes down? Like you need to have a waterfall of models to basically pick it up so that your voice agents don't just stop working.
03:28 We also talked about like turn taking and how like it seems like a trivial problem where hey like you're talking to a voice agent and let's say I pause. So how does the voice agent know that hey like I should I should interject and I should actually like respond to you versus oh the person is just taking time to think. So turn taking is actually like not a trivial problem to solve.
03:47 And so we talk a little bit about that. And then also we even talk about like how um I think Pipat built a voice agent for the World's Fair, right? So we even talked about that a little bit and how like I think I mean you can talk about this a little bit but I think like you guys built a voice agent to like if anyone calls the voice agent it'll answer questions about the conference and then you got a ton of like real world information and turn that into a benchmark.
04:08 >> I mean to be clear I have no idea actually. I never looked at the analytics so I have no idea how many people actually called. probably you know people just kicking the tires you know like it's not it's not that serious like we it's a conference people know what it is yeah but you know I think it was a good deployment of use case daily is a sponsor so why not >> I you know I I would say that vast majority of voice stuff is support and like yeah oh my god there's you know so many call centers out there hundreds of
04:33 billions of dollars spent on this stuff and they're all you know bad so hopefully we can sort of raise the the state of the art which you know I think that's why Decagon and all these things people exist To me, you know, it's not as interesting if a bunch of people who are obviously selling you voice pipelines telling you that voice pipelines are the state-of-the-art.
04:51 Like, yeah, duh. But, you know, having having the the people actually focus on customer support and all those things also basically conclude like, yeah, like the models are not there yet. They may never be. And actually, this is just the way that you got to do it. And the the people that really feel it are not the researchers cuz researchers just always want bigger models to solve everything.
05:10 And the product engineers will try to do it, but the FDES, the people actually dealing with customers will be like, "Dude, they made this like horrible mistake. Like, this can never happen again. How can you guarantee me that?" Well, >> yeah. Yeah. Yeah. So, like actually we even talked about that a little bit. So, like we talked about like the inbound versus outbound use cases cuz we even did that once.
05:30 We like ran an outbound use case and like what you find is whenever somebody like picks up the phone and you do an outbound use case, a lot of the times they just hang up as soon as they realize it's a bot. So we even talked about that like how do you how do you solve for that? >> All right. Uh that's a good teaser. I you know I admire your work. Uh I'm excited to feature on space.
05:45 We'll be doing more uh in future together but uh this is just an intro to FDE for uh at least for at least a l space audience. >> I think this might be a good place to give everyone a 101 lesson on what does a simple voice agent architecture look like. So maybe that's a good question for Brun because you guys uh run Pipecat. So >> yeah. So there are many ways to build one nowadays because the complexity is the nature of things.
06:12 But the simplest one which is called a cascade model is typically uh voice input. Uh it comes through a transport. It could be web RTC, phone call or web sockets. It goes into the mo into a speech to text. Uh the transcription happens. There can be some additional models right there for background noise removal, voice isolation. So if there are few people in in the foreground, it will isolate to one person, remove the background noise, it there's a turn detection model, unlike text where you know when you're done, you
06:47 press enter and the message goes and the LLM starts to like execute with voice there is no it's not walkie-talkie or pushto talk. So there is no additional uh information apart from like if you're looking at the face you could actually figure out that the person stopped speaking but typically use something like a voice activity detection and a smart turn model to like figure out that the turn is complete.
07:14 Um if I pause mid mids sentence that's different from like if I finish speaking and pause uh you may use a speech to speech model right there. Uh then when you have the text you uh from the transcription then you send it to an LLM. The LLM may respond with an inference. It may actually tell you it may infer to tool calls. It may do a bunch of things there.
07:37 You get that when you have the full output you text to like you do uh text to speech DDS and uh it's usually faster than real time. Uh so then you have to like figure out how to like stream it back to the end user. It'll go over the same p transport that that you spoke over and then it'll play it back on your speakers. Uh but things are getting more complex as people, you know, this was like 2024 25 people wanted to do this.
08:07 If you're doing an outbound use case, like you do an outbound call, debt collection is quite commonly a use case for this. Um you get a call, they tell you you have like an unpaid bill. um they ask you if you would pay it or you know give you a certain set of options. Uh for that the bot has really like if you go off the rails like the human goes and starts saying random stuff it can just drop the call.
08:31 It's not obligated to stay on the call in any way. So guardrails are easy. Um it knows what the inputs it can accept. It knows what inputs it can it doesn't need to respond to. It flips when you have an inbound use case because the bot has no context like has some context of why it exists but doesn't know why you the person who's calling it is if you're a mom and pop shop like a flower shop or or like a barrier or something like that okay there's only like finite things that you can do there but if you are more complex
09:03 let's say you're Amazon right there like a billion products on your platform um and you have different policies for different things is it uh you know is it a physical item is it a cheap physical item is it uh an expensive item u is it an electronic item like all of that so it needs to figure and then when you said something like you know you can't put all of that inside a single prompt because it will like LLMs are you know goofy in that sense that always remember the first 4% and the last 4% and everything in between
09:37 they kind of forget uh so if you have a return policy right in the middle of that context, it's very likely it'll hallucinate. So you have to do more smart things and that's where you have like more models come in. Uh you can have a compaction model like we're all using coding agents now, right? So you know like if you're using like a 1 million context model like at 25% you should become nervous at that 25% mark.
10:06 is nowhere near 70% mark but at 25% mark you're like okay should I start compacting and like saving my work so that this thing does not go off the rails right the same thing I think voice AI has been like one generation of ahead of coding agents like we've all in the last two years solved things that felt like so alien but very tangible for us which now when we see coding agents do we like man compaction we we were doing compaction from day one like We knew like after five tons or 10 tons and your context was only 250
10:39 tokens or 250,000 tokens or 50,000 tokens, you know, models were really small like before and at 10,000 tokens it would like go off the rails. So, so a lot of our work is I think about like making sure that the bot does not hallucinate and trying to like keep track of the conversation as the the turns progress. And the turns are fairly short because the either way the bot is more likely to speak more than the human did.
11:10 So you have to like kind of keep track of like what did the human say? Where are we in the conversation? >> Yeah. So um I guess this is a question for one of you three. Uh so do you guys use this like cascading model for your guys' voice agents? >> Yes, that is one of the major offerings that we have. Cascading and speech to speech. >> Yeah. I was going to ask like that sounds too complicated.
11:34 Why not just go voice to voice? Like why not? Like why do you have to have this like crazy three-step process? >> I feel like the easiest way to to answer this is you can call into a, you know, really nice voicetovoice demo and you're like, "Wow, it's like listening to me laugh and it's like responding to my tone and it's so snappy. It's so fast." But then, you know, I I tell it that it asked me what day I want to schedule my appointment for, and I say, you know, next week.
12:01 And it says, "Great. Uh, is that June 10th?" And I'm like, "No, it's the year is 2030." And it's like, you're right. It is 2030. So, let's schedule this for June 10th, 2030. And I'm like, great. Sounds good. Um, and so from that perspective, I am actually very curious, especially on the on the smallest side, um, how that's evolved over time because we are very much keeping our eye on how these speech-to-pech models are performing.
12:23 Uh but overall I think our current stance is uh the cascading model just allows you to enforce so many more um rigid guard rails and just tight control over the ability to say hey this input is going to go through the same supervisor models to detect for you know prompt injection or social engineering. Then it's going to go down the conveyor belt into intent selection.
12:46 Then we're going to uh optimize context by checking conditions ahead of time to say, you know, in this complex um process that we're going through with this user, uh maybe this prompt is applicable if they're this type of customer, but this prompt is applicable if they're that type of customer. So let's, you know, figure all this stuff out ahead of time, compact and compile a good system prompt for our message generation model, and get a response back.
13:12 And then we can take that response and we can check a whole bunch of other things, right? uh we can check is this grounded in truth, right? Is it 2030 or is it 2020 20 whatever year it is 2026 um and then it goes back out over the line, right? And so the obvious constraint to that is how do you make that performant, right? How do you parallelize as many of those steps in the conveyor belt as possible?
13:34 I think the last six or so months has been uh a really amazing feat from our engineering team at least to find and shave off like 10 milliseconds at a time across every single part of this pipeline to make it feel snappy even though there's a lot of things going on behind the scenes that uh you know you don't necessarily have to do with a speechtoech model.
13:53 So that's at least my take but yeah I am curious for for the rest of the group's thoughts on this. Um you want >> yeah um no I think you brought a question around like cascaded versus speech to speech like let's talk about like why did people start thinking about speech to speech and so the initial idea was simple that hey if you convert speech to text it's going to lose the emotional information right so if you say hey um you might be saying it in a sad way or an excited way the bot is going to answer in the same
14:26 manner right so that was the obvious reason people started I think uh at some point of time at least at smallest like that has evolved to uh speech to speech is uh a more natural way in which the human brain operates. Uh so for example uh when you do the cascaded thing you do speech to text then you send the prompt to an LLM and then it responds right so we call that a synchronous architecture like it's happening one after the other but our brain is thinking while listening so as I'm speaking to you you're already
15:02 forming your thoughts and if I'm talking for too long you'll interrupt me right or you might be taking notes in the back end right so you might be essentially doing tool calls while uh I'm speaking to you. And so the whole idea is that if you ever want to pass the Turing test of how the human brain operates, you need something that is working asynchronously and can not just understand emotions but also operate uh like take in speech natively and give out speech natively asynchronously and and so that's how why we have
15:33 been sort of building Hydra. Now Hydra is our speechtoech model. Um now in terms of accuracy and and interpretability and and all those things I think whenever there is like a new architecture that comes out it's often sort of good in one parameter and then regressed a little bit in other parameters right like so for example um the speechtoext accuracies for a just a speechto text model might be way better than the encoder of a speechtospech model that that you have and so I think The challenge is that while you made
16:09 progress on the um you know uh making it more natural and operate more humanlike etc etc how could you keep the accuracy bar the same and so I think a lot of that comes down to interpretability of um such models like because you don't want it to be a black box so how can you understand where it is lacking and um so that that's a lot of research that we do and how to make speechtospech models more interpretable Um the other constraint is I think initially the speechtospech models like I think sesame etc that came in
16:41 that they were just speech to speech like they literally took in speech and give out speech ours is like multimodel so it takes in speech and text gives out speech and text both so it can do tool call and uh take in text parallelly. So if you want to put guard rails if you want to do all those things that constraint does not go away. Um what we are also seeing is there is still at least in enterprises a lot more cascaded deployments compared to speechtospech model but speech to speech I think will be like an eventual
17:11 future that that's my opinion. So my take is that when you have like competing things you end up in a hybrid and uh I think the answer for the midterm will be some form of hybrid because the speech to speech are improving. There are parts of the conversation loop where you would say like oh my use case of my workflow is complex. Uh I'm going to use multiple LLMs anyway.
17:39 So for the active loop like I'm talking to as a human you called me. I'm answering your questions. But whenever I have to do some kind of lookup or something, I delegate to another LLM which will then do the cascade stuff. So the voice in the first part of the loop keeps running. Uh and then you have like interesting kind of semaphors or like you can think of the as threads.
18:03 uh you want to interrupt the the speech-to-pech model because you realize that this is a complex question and you want to before the speechtospech response you say actually you cannot answer this question right and delegate to the cascade and and so on so forth so typically I kind of say there are many use cases where you know if you wake me up in the middle of the night and ask me a bunch of questions there are definitely some class of questions that I can answer without thinking and we all do our jobs in a certain
18:31 way where we can in autopilot like 50% of our time, right? So as these use cases become emergent and uh you know you're fully deployed with a customer doing high volume, you could essentially train specialized models that fully understand like that 50% use case very well. Uh and you could you could always ask the model like hey what is my return policy and say like in the simplest case this is my return policy and it applies to like 70%.
19:04 And I know which SKUs or or products it belongs to or it applies to and if you ask me outside of that then I have to like do all the crazy work and if I don't then I answer it right. So, >> so on these cascaded pipelines, what what models are you guys using let's say of the frontier models of the you know GPTs of the of the uh what opus and sonnet are you guys using the latest ones or are you using you know like GPT4 because it's like the right balance between speed and intelligence.
19:38 If you turn off thinking then you can actually use any of the like if you want a fast model which responds to like the prompt and does not need to think because if you wanted to think you actually start to think about um paralyzing the work because you want the the first model to be super fast. I still like my Gemini 2.5 really well. Like it's so fast.
19:58 It's like the 3.5 which they launched is like okay it's slower but the 2.5 is so good. Uh even the haiku is really really good. uh they're still slower than what you would expect, but yeah. >> How do you parallelize that? >> How do you parallelize that pipeline? Because it feels like, oh, well, I do have to first know what text the person said and then the LLM needs the text to do anything.
20:25 And then you can only generate speech once the text has been generated that you want to wanted to send, right? you you just spend the you know you can send the speech like in the speech to speech model case you would send the speech directly to one place fork it into another place run the st there um and if you want multiple models the text output from the first forks into multiple LLMs you can think of it as like a waterfall like you know it comes down >> because you don't know what the output exactly will be >> yeah
20:57 you put a put a gate at the bottom which XR or whatever fancy thing you want to So which says like the if all of them say we are right then you put another LLM downstream which one is better like you can go really crazy if you love computer architecture 20 years we had nothing to do now you can do all this crazy stuff from 20 years ago. >> Yeah. Um so how do like accents and different languages all fit into this?
21:27 Maybe Sudarian because you guys are building speech to speech. Um yeah so at least for speech to speech right now we are just fully focused on making it more intelligent in English and you know getting it really good in English like we don't want to do any other we don't want to introduce any other variables because the technology in itself is I would say quite frontier asynchronous is not yet like mainstream etc right uh but in terms of like cascaded yeah we've seen like a lot of demand in so we have a lot of presence
21:57 in India so we see like a lot of demand from India Um we see a lot of demand from Latin America etc. Um I think US is more like mostly English and then there's some Spanish and then there's a lot of accents to it. Um and I think yeah US at least we have not had any troubles in terms of um technology. I think it's yeah I mean noise cancellation is probably the last mile of problem that is pending uh in terms of handling those things and but yeah that has nothing to do with accents I'm curious what you guys have seen >>
22:34 u maybe I can add uh some examples about that um because on multilingual uh I like to definitely leverage cascade model because there are multiple levers you can pull uh to talk about an example um it's a customer we were deploying uh for Japan Um, so you could typically go uh you know to your uh 4.1 on OpenAI or your deepgram for transcriber and that typically worked well.
22:59 Now the challenge became on the voice uh piece. So it turns out uh what we thought it was state-of-the-art for um something like English, Spanish, Portuguese um those type of uh that are typical in the US didn't work at all. Um so one of the benefits uh about this is that you can swap. certain pieces that doesn't quite work for you. Um, at least on Vapy, how we solve it is that we let customers bring their uh custom uh text to speech server.
23:27 Um, so they actually have like there's some startup there's a lab um in there and there's typically uh you will find that in every market. Um, Arabic is also really hard to get it right uh in pronunciation of brands, addresses uh and that really we don't uh cannot build expertise and optimize for every single use case. So that's one of the reasons why I like that as well.
23:54 Cool. Um, also actually question on the like sesame model. Yeah, like they came out with this like really insane like demo and then I've never heard of them since like what do you guys know like what happened and what's going on? >> Yeah, I think uh so they are building hardware is what I understand like they are putting uh those the voice into some sort of a hardware.
24:15 I'm not sure if it's glasses or what are they working on exactly but uh I know a few people who got hired there and they're all sort of focused on the hardware and voice sort of backgrounds. Yeah, the answer I got was that the CEO had already made a ton of money and he just really wants to play around basically. So I guess he can do whatever he wants.
24:35 Um, cool. Um, so is it standard practice for in this like cascaded pipeline for you to just give just have one giant like system prompt that you're giving the LLM on these are all the rules for how you should be responding? um or is it more of a um like a workflow approach? How do you guys think about that? What's the standard best practice right now?
25:06 >> So I think it highly depends upon the use case and also the size of your prompt, right? So I think I have a counter approach to Baroon in in terms of that like usually in inbound I find you know you you're usually having dedicated lines and you know specifically where the workflows are going to go right and so in these cases if you know a little bit more about and you can predict where the conversation's going to go then you can have a more of a node or graph builder kind of approach but if you don't know what's
25:36 going to happen typically I think in outbound yes there is a dedicated like message but what the caller says back you never know because they could just be frustrated they could be angry or they could be I don't know like annoyed and in that case when they're erratic when a person's erratic you never know so having it in a way that is like all in one one prompt allows you to have an more like a central brain in terms of I can pick up components from my prompt that may not have been in a dedicated flow but actually can
26:10 be utilized iz to enhance the response. And then on top of that, you have to think of your knowledge base and knowledge the things at your at your expense, right? Like in a inbound flow, maybe you know exactly when you need to pull a specific part of your knowledge. So you don't need to extract every single piece of context, right? Whether it's your websites and documents, but sometimes in outbound, you should utilize all that information to get a better response.
26:39 then you will rely more on cosign similarity and stuff to make sure that hopefully your retrieval is good enough to respond to them. I think from my perspective um the limiting factor has been and and will continue to be although it's getting a lot better over time just uh to you know some of the earlier points made uh the ability for these uh frontier LLMs to be able to take a massive wall of text and not just read the first three lines and the last three lines but actually be able to properly uh reason and like
27:14 localize themselves into you know we are halfway through this procedure and these things have already happened and these things haven't happened yet and these things are true and these things are false about this situation and so therefore like this is the one thing that I should be really zooming in on right now right um so from our perspective I think um we kind of take an approach that depends on the customer um but we do within our execution engine have the ability to um make some of these rules or prompts or
27:40 guidelines more situational so that we can you know check if these are uh relevant right now before we even uh include include them or or uh uh don't include them in the system prompt. And so that was our solution that was especially relevant, you know, like 7 months ago when the LLMs really were just like consistently skipping step 4 A1. Um but we have found over time now that uh sometimes the reliability of deciding when these things are relevant or not is actually worse than just giving all of it to the model
28:10 because the models have improved a lot since that time. And so now we're going back, I think, swinging a little bit back towards just give the model everything and more or less it will figure it out. Um, but at the end of the day, you know, with nondeterministic systems, it really just ends up being customer by customer. uh we have to really understand their use cases and at the end of the day just build it and then test it a whole bunch of times and then use LLM as a judge to tell us you know did it work was it
28:36 reliable or do we need to add a little bit more um pre-processing and and context optimization ahead of time in order to make it reliable. So the the real answer is most likely always just test and find out and then test again and then find out again and you know it goes on forever. I'm assuming that yeah maybe the models are able to handle those edge cases now like just in a giant prompt but that probably also increases latency.
29:04 So first could you talk about like what is a good latency look like for a voice agent and is that even true that yeah the models are getting better but they also require more time and so for the end user that's actually a bad thing. >> Yeah that is a really good question. I mean I I think the definition of what like benchmarking latency stats look like that are good have uh changed a lot and um expectations are you know getting better and better and better from our customers.
29:32 Overall, I think it it is also important to mention that um as much as fillers and contextual fillers is something that uh we don't like to hear just as a noun, um it is normal for in normal conversation for humans to say a few words while they're thinking. And so we've actually found that when we have zero contextual fillers, sometimes it actually feels more rigid than when we don't.
29:56 And so we tend to find that uh having the the right amount of fillers that um kick in when you know we are experiencing some sort of latency because of some very large prompt or some very complex thing that the model is thinking through uh most commonly you know running a tool call where our customer's API becomes the bottleneck and we're waiting you know 5 seconds for something to come back.
30:16 Um being able to say hey just give me a sec uh looking at that. Okay, cool. Here's here's your answer, right? There's actually nothing wrong with that. Uh when we've when we've tested with our customers, um obviously the the limiting factor in that case then becomes how good are your fillers and how natural do they sound. Um and so that's been a place where we've invested a lot of time as well to say how do we create an execution engine where we really don't need fillers, but in those moments where we do, they're
30:42 really good. Um, so that's at least my take, but again, I know folks have probably a lot of opinions here about latency and and how to find that trade-off because in my opinion, I honestly think that that's the hardest problem to solve in voice deployments is just that trade-off that's always going to be there of like performance versus um stability and reliability in in general.
31:00 So, >> um, yeah, I agree with all of that. Some of the ways that we think about it in Vappy, uh, is offloading, uh, some of that instruction. So break it down. Uh sometimes you don't really need um that you know 10 15 step workflow in the in the prompt. Uh one example I can talk about is um this uh collections uh use case. Um we do allow users to capture um businesses to capture credit card information.
31:31 Uh but not every call is uh about that. Some people already have a payment method maybe added or that. So uh one way of doing that offloading into a uh specialized agent uh so we can start thinking about multi-agents and there are many architecture and patterns in there um is to constrain the uh instructions that you give it to to that particular moment of the conversation and once that achieves its goal it can go back to that u more free form approach which has loaded like uh your longer uh context.
32:02 Something we're also been uh experimenting lately um and it's more about the guar space is um to uh your point where you can have uh think about it as a waterfall you can actually stream to multiple places. So what if you have uh streaming down to a a small language model which can uh do inference in a very little time. So think about classifiers to maybe collect the intent um while the pipeline is still maybe adding a filler word as well to keep a consistent experience but uh kicking off a background process that is
32:39 still intelligent still um thinking. So those are still things that we'd like to invest a little bit more research on. Um so yeah constantly investing on that. >> Sure. I think of one thing that we haven't also factored in in terms of the latency uh tradeoff is cost right now that you're putting such a big large prompt into your your LLM. You have to think about those calls where many people just hang up and predominantly calls just hang up 10 seconds into the call.
33:14 Right now you're having to pay for all those tokens to just get inputed in. So that like Stephen mentioned splitting out your prompt, you are also playing to that strength of lower latency and lower cost. I'll speak just to the benchmarks. So we publish u turnbased ST benchmarks, LLM benchmarks and TDS benchmarks. The whole point of that is that as I think today Neumatron 3.5 launched and there's already an ASR um benchmark out did very well the Nvidia one.
33:51 So um the idea there is that all our tooling is open source. So with the HTT benchmarks, the TDS benchmarks and the LLM benchmarks all the turns are actually open. So you can run it against any new model that comes out. Uh you can look run it locally. you can look at the outputs from from our graphs and compare that to what you're getting on your infrastructure if you're self-hosting or you know elsewhere.
34:16 Um so that's one important aspect to like uh consider if you know what your turns will look like you can just update or you know just download our benchmarks update the turns and see how that like goes for your own flows. Uh usually for the LLM stuff, it's really useful because you can throw out the ST and the TTS and just run the LLM inference loops to like check against uh a customer.
34:42 So yeah, I think eval are really important. Um, and if you have like a like while talking to customers, if you can evaluate like what type of conversations or like what type of workflow they're going to have, like those are real direct inputs into like what you can directly start testing even though you're not live with them or even they've not shown any intent.
35:05 Uh, so I think that's where like the sees and the forward deployed engineers can actually come to for like you we all have cloud running in the background, right? So you can just say here's what we had a discussion like can you build an eval like based on our eval suit like can you like just run a smoke test with the conversation information that we already have and that helps a lot I think.
35:27 Yeah, just just adding to that like we are seeing so we while we focus on models we also have all the model sort of orchestrated on our platform. So you could build a voice agent on smallest. Um we are seeing a lot of customers who just pair our models with our own like small language model. It's called electron. fine-tune one of those to make it work for realtime voice use cases and see a lot of folks using that with better knowledge basis memory etc over GB 40 4.1 you know the popular realtime models um just a from a
36:02 cost perspective it's lower b from latency perspective it's like way lower u and then just more reliable I think I'm not sure if everyone faces this but like um open AI APIs spike and in um you have no control on those latencies. Um with if you sort of have a self-hosted model um you could scale up scale down based on your requirement. So that's one thing we are seeing a lot with our customers.