← All transcripts

Context Engineering for LLM Apps: What Worked, What Didn't Transcript, AI Summary & Key Points

InfoQ · 4 days ago · Science & Technology · 47:35 · EN

Watch on YouTube

AI Summary

Context engineering makes LLM applications more reliable by deliberately controlling memory, retrieval, context selection, ranking, compression, token usage and caching. Ricardo Ferreira built My Jarvis, an Alexa skill backed by an LLM, to stress-test the open-source Agent Memory Server on Redis. The application encountered context poisoning, distraction, confusion, rot and clash; upgrading to a newer model with a larger context window increased cost but did not fix the experience. Solutions included date-and-time tools, session memory, long-term vector memory, tenant post-filtering, separate knowledge-base retrieval, query compression, few-shot examples, reranking, token windows, summarization and semantic caching. These mechanisms increased the number of LLM calls, so token growth and operating cost became a separate architectural problem. Context engineering belongs in the architecture of serious AI applications that need humanlike interactions.

Key Points

  • My Jarvis connects an Alexa skill to a Lambda function, LangChain4j, and OpenAI; Alexa interactions must respond within 8 seconds.
  • Context poisoning produces nonsensical information, context distraction supplies excessive irrelevant data, context confusion mixes unrelated information, context rot degrades growing context, and context clash mixes conflicting versions or timings.
  • Upgrading to a newer, larger model with a larger context window increased cost but did not fix the context problems.
  • Chapter 1 — LLMs do not reliably know the current date and time; a date-and-time tool supplies that context for reminders and appointments.
  • Chapter 1 — Session memory is bounded by a time-to-live. A five-minute TTL caused the assistant to forget Ricardo's name after five minutes, motivating long-term memory.
  • Chapter 1 — The built-in chat-memory implementation caused a tool-calling loop that exceeded Alexa's 8-second timeout, so a simpler working-memory implementation replaced it.
  • Chapter 1 — Long-term memory combines stored text with vectors and uses semantic search to retrieve relevant memories for later questions.
  • Chapter 1 — Multi-tenant vector retrieval requires post-filtering by tenant ID and owner ID so one family member's memories do not answer another member's questions.

AI in practice

Used for

Agents

  • My Jarvis — Answer questions and carry out tasks such as reminders through an Alexa voice skill. 2 held 03:06

Tools & resources

9 items

ANo. 4970
AIAINotes.us AI product

Agent Memory Server (AMS)

In the AINotes directory

Agent Memory Server (AMS) is an open-source memory server built on Redis for AI applications. It provides short-term session memory and long-term vector memory, with support described in the source material for multi-tenant post-filtering, query compression, few-shot examples, reranking, summarization, token-window management, and semantic caching. It was used behind an Alexa skill to test memory handling in an assistant under an eight-second response constraint.

Mentioned in
1 video
Kind
AI
ANo. 4971
AIAINotes.us AI product

Alexa

alexa.com

Alexa is Amazon’s voice assistant and conversational interface, available through the Alexa app, browser, Echo devices, Fire TV, and other compatible devices. It supports natural-language requests, answers questions, maintains conversations across devices, and can turn information into actions such as scheduling, creating checklists, planning events, and booking services. In the cited use, an Alexa skill placed an LLM assistant behind an interface subject to an eight-second response limit, allowing the assistant’s responses and context-handling behavior to be tested in everyday use.

Mentioned in
1 video
Kind
AI
ANo. 4973
AIAINotes.us Tool

Amazon S3

aws.amazon.com

Amazon S3 is a cloud storage service from Amazon Web Services. In the cited use, it stores an AWS Lambda application package when the package exceeds the size limit for direct deployment.

Mentioned in
1 video
Kind
Other
ANo. 4972
AIAINotes.us Tool

AWS Lambda

aws.amazon.com

AWS Lambda is a cloud backend function used in the described system to connect an Alexa skill to an application.

Mentioned in
2 videos
Kind
Other
CNo. 4968
AIAINotes.us AI product

Cohere

cohere.com

Cohere is an enterprise AI company that develops language, speech, translation, embedding, document-parsing, and retrieval models and solutions. Its Rerank model is a retrieval-optimization service that scores and reorders vector-search results so that more relevant items can be selected as context for an answering model. Cohere also offers enterprise deployment options including dedicated, private, on-premises, or isolated-VPC inference.

Mentioned in
1 video
Kind
AI
LNo. 4967
AIAINotes.us AI product

LangChain4j

langchain4j.dev

LangChain4j is a Java framework for building applications that connect language models with tools, memory, retrievers, query transformers, routers, context injectors, and content aggregators.

Mentioned in
1 video
Kind
AI
MNo. 0415
AIAINotes.us AI product

Model Context Protocol (MCP)

Open source · modelcontextprotocol/modelcontextprotocol

Model Context Protocol (MCP) is an open, standardized protocol layer for connecting large language models and AI agents with external data sources, hosted infrastructure tools, and other contextual data. It defines formats, metadata, protocol schemas, and APIs for sharing information and tools, attaching, referencing, and validating context such as documents, embeddings, and provenance, while supporting scoped authentication and permissions. MCP translates JSON requests from an agent into calls to service APIs, including CRM, container-management, and GKE capabilities, and can provide design-system context and related tools for generating consistent applications. The official project publishes its specification, documentation, and protocol schema; the schema is defined first in TypeScript and also provided as JSON Schema for broader compatibility. The protocol was created by David Soria Parra and Justin Spahr-Summers, is hosted at modelcontextprotocol.io, and is licensed under the MIT License.

Mentioned in
25 videos
Kind
AI
ONo. 0105
AIAINotes.us AI product

OpenAI

Open source · openai

OpenAI is an AI model provider whose language models are available through an API and can serve as one provider in multi-model routing architectures. The videos describe its models being used for AI replies and extraction, Amazon-review analysis and customer-service email generation, legal workflows, complex software development and debugging, and experimental or auxiliary agent tasks.

Mentioned in
21 videos
Kind
AI
RNo. 2252
AIAINotes.us Tool

Redis

Open source · redis/redis

Redis is an open-source in-memory data structure server, cache, key-value and document database, and query engine for real-time applications. It supports strings, lists, sets, hashes, sorted sets, JSON, time-series data, transactions, scripting, queues, streams, publish/subscribe messaging, full-text and geospatial queries, aggregations, and vector search. The project is used for caching, distributed session storage, event brokering, real-time analytics, and as a supporting service for the support desk shown in the video.

Mentioned in
3 videos
Kind
Other
🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Context Engineering for LLM Apps: What Worked, What Didn't — InfoQ (47:35). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by InfoQ. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:05 Glad to be here. Uh my name is Ricardo Ferrer and uh I'm going to be your next presenter today. We're going to be talking about like this context engineering stuff. And uh this presentation is actually going to be like a mix of like a lessons learned that I've kind of when I've developing a specific solution that I'm going to talk in a minute, I kind of learned the easy and the hard way.

00:25 And also is going to be like a mix of like storytelling because this application that I've developed that actually has a story and um I'm actually going to try I was setting up my Alexa device which is kind of has to do with what I built and uh I just set it up in the last minute. So that's a good sign because it was taking a while to change the Wi-Fi for it.

00:44 All right. Uh in the bottom you can see my handle this place that's that's the same handle for LinkedIn X, Twitter, Instagram. all my social net GitHub all social networks your same so if you want to get in touch uh feel free to do so uh I work for Radis so I think it's pretty clear all right so let me tell a story about what what I actually have built so you can understand where I'm coming from um I work inside Radis from this part of the sim called developer experience right and one of the thing that we started doing

01:15 probably I think about two years ago is to build an open source project called agent memory server or AMS on top of Radius, right? So the idea of this open source project was to kind of create a very quick fast memory layer on top of Radius. So we can developers can build either shortterm memory and long-term memory on top of it, right? So build more natural humanlike conversations with it.

01:40 And then my my our team uh that developed this project, we were also being asked to develop applications that put would put this open source project to use. Right? So we've released open source but we also try to test internally. So the the the the basically the requirement that we are asked was build something that can put some stress on the on this type of open source project.

02:04 And I come from a background where I've developed some Alexa skills in the past, right? And I know for sure there are some very unique and interesting challenges to when you develop Alexa skills. One of them is that not sure how many of you know this uh the Alexa interaction has a timeout fix a timeout of 8 seconds, right? So it needs to respond within this time frame.

02:29 So I I thought it okay, I'm going to try to use this like for in this open source project. So the first thing that I tried to do right and that was the first motivation. The second motivation is that probably on this slide that you're looking at right now, how many of you have interacted with Alexa devices in the past and then most of the time spent like not knowing things that actually knowing it right.

02:49 So I think if my Alexa is ready right now, we can test it. Alexa, can you explain what is context engineering? >> I'm not quite sure how to help you with that. >> All right. So, how many of you have spent like I listen to the same thing? Okay, point is I kind of had this idea. What if I could put some brains behind Alexa, right? So, what if like I put an LLM behind Alexa and then I create this uh skill that I've called my Jarvis, right?

03:15 So, in a minute you get the reference for where the name came from. So, the way it works is like this. Um my Jarvis is a skill, right? So basically what you have here uh is the Alexa device and this is going to make this interaction this backed by a lambda function right this is usually how Alexa skills have the background backend spar and then on top of the Alexa lambda function I have two layers one of us is the actual SDK from the Alexa skills right so that's a bunch of wrappers on top of JSON payloads that comes

03:46 from the Alexa interactions and then I use this framework called link chain forj which is a java version of the link chain framework, right? Kind of ish because actually technically link chain for is not a part of the lane chain projects more like a lane chain uh along with yama index along with spring AI. So is a is a more complex framework but point is that's what was kind of a backing the uh LLM which in this case I use an open AI right so basically I kind of was able to do this Alexa ask my Jarvis to explain what

04:21 is context engineering. Context engineering is the art and science of designing systems that interpret and respond to contextual information. It involves creating algorithms and models that enable machines like AI assistants to process and prioritize relevant data from user interactions and environment cues. This discipline is crucial for developing intelligent systems that can provide personalized and accurate responses by understanding the user's intent and the situational context.

04:52 That's a very like insightful answer. Thank you. Thank you. >> No, I was asking thank you Dan. So, but thank you too. Yeah, obviously. And thank OpenAI for the answer. Right. So, the point is that's how I kind of felt after I start developing this like so the name my Jarvis came from the famous like AI character from the Marvel universe and I'm huge Marvel characters fan for those of you that know me.

05:14 And then as I was kind of got excited with this development, I started like asking more questions, more elaborated questions and try to incorporate this skill in my daily life. Right? So what I ended up realizing is that uh yes the skill is able to retrieve answers but it's kind of a fairly easy right once you actually make a call to open a uh to the LLM is is going to be virtually able to respond any possible question.

05:41 But the human aspect, the mannerisms when you ask questions, when you interact with the let's call the model is what takes place most of the time. Right? So what I found that is that the experience of dealing with that was kind of odd at first and then I started realizing that okay I am running into a set of problems that has to do with like context poisoning which is like a things that didn't make any sense at first or context distraction too much data too much information that was not necessarily spending more time

06:15 creating elaborated answers than I could creating the right answer right context confusion which is like a bunch of irrelevant data that was kind of a mix and match. U also like the famous problem of as you content grow the context rot right. uh and ultimately the context clash which is kind of a that happens a lot like a mixing of oh this is true this is false this is one uh the versioning and timing of the answer and all of this mutation of the data kind of a creating a a very poor answer that was given right so

06:52 probably I'm I'm not going to be make a bat here probably the first instinct that most software developers do in situations like this is all right I'm using I'm writing on a back end and if the back end is malfunction and it's not keeping up. What I'm going to do with the back end I'm going to improve it or I'm going to change it right like what we used to do in the old days of uh databases right relational databases on prem right oh if the rel relational database on prem is not keeping up with the workload what's it

07:18 going to do more CPU more memory or put a cache like in the side so I tried to use a newer model for chatpt like 5.4 before. Yeah, increased my cost by doing this, but also my assumption was if I use a larger and newer and a a larger context window, it was able to kind of create more elaborate answer. Silly of me thinking that because changing the model didn't do a thing in the experience that I was having, right?

07:46 So, I think that's was the first like lesson learning of this that not all the time upgrading to newer models in your AI implementation is going to solve your problems, right? Yeah, I know for sure that in my case it didn't solve mine. What I actually needed was and I didn't know at the time that's what one of the fun sports about this implementation.

08:04 I didn't know that I needed context engineering, right? That was that implementation started probably like a one year and a half ago. The therm itself was originally starting to catching up, right? But just like starting to catching up recently, harness engineering, right? uh but then uh at the time what I actually have felt is that all right there's a bunch of situations that I need to fix it right as an engineer as a software engineer I need to fix this each one of them pragmatically right and as I was fixing them

08:37 right I was using some instinct approaches right some others I actually was collaborating with the rest of the team and then oh no Ricardo there's a name for this uh this is kind of a context uh reranking right because my background is not for AI engineer. I forgot to say this, right? I'm a software engineer, right? So, uh most of those practices that I would assume that are pretty common like in the AI world, it wasn't for me, right?

09:04 So, the huge majority of the time was kind of a trying to fix them instinctively, right? But if you ask me, not the industry, but if you ask me recorder, what I have felt when I actually fix all the things, I would categorize and I was conceptualize context engineer exactly as this. It's something that goes beyond prompting, right? Because that was another instinct of mine like let me create a more elaborated system prompt that would help the agent to think a little clearer and better, right?

09:32 Another mistake. What I ended up creating so was a a set series of very intentional, very well-crafted approaches to pragmatically solve each one of those bad experiences. Right? And if you ask me, this is what context engineering is about. Right? I know that that there's a very more conceptuals and more elaborate answer that Alexa just gave us, right?

09:55 But the point is I think that's how I felt at the time. Right? So this presentation I'm going to break it down in two chapters. Like the first one we're going to walk I'm going to walk you through and all the single problems that I went through and how I solved them and some of the lessons learned that I kind of learned it the the easy and the hard way.

10:13 And chapter two we're going to talk about cost, right? which is the incidental type of cost that by using context engineering I ended up having which is by itself my most important heads up for any any of you that would like to start like implementing context engineering as a practice in your applications right there's a cost uh involved all right the first problem that I I kind of went through was this uh probably this a no-brainer for everybody but LLM have a hard time deciding ing what day and time is today, right?

10:48 So, how many of you have learned uh have used an LLM and then all of the sudden you rely on an answer that oh yeah because we are in the 2021 this has happened this this and that. Okay. No, we're not in 2021. We're in 2026. Right. So, the first thing that I actually had to do and because one of the first use case that I've implemented was a reminder system.

11:08 Right? I I wanted to use my skill to remind it uh myself about things. Actually, let me do a test here. Alexa, ask my Jarvis to remember that I have to say hi in 2 minutes. The me >> that you need to say hi in 2 minutes. Would you like me to set up a reminder for this? >> Yes, please. Okay. So, what you're looking at right there is >> what you're looking at right there and let me try to make a little uh larger so you can actually see it.

11:40 Right? So, this is uh the actual what we call LTM or long-term memory stored internally on radius. Right? And this long-term memory is basically a combination of the text which is the actual information or memory that you're supposed to have and the text vector, right? Because there's a bunch of vector searches that happens behind the scenes in order to do that.

12:02 Right? But with that said, the first experience is that like I mentioned before, LLM has no concept of time. How do I solve this? Using tools, right? So I created a tool that I was able to use it in order to kind of teach the LLM along with the context about dates and time. So that's how uh hopefully that's how it's going to remember the exact day and time and that became even more important when I had like reminded me about the appointment next week on Wednesday.

12:30 So those things really matter, right? So for those of you that never worked with link chain forj, this is as simple as to create like an agent in agent for link chain forj. basically associate the models the list of the tools and you build it that's AI service shed systems and then you can start invoking your queries it's that simple right and then that's what actually start to happen like uh the reminders start popping up exactly in the right time that I was supposed to kind of do the presentation right uh then here

13:02 is my second problem that I went through I kind of start asking questions that would require a subsequent remembering of those questions right so for example yeah my name is Ricardo nice to meet you Ricardo you know my name yeah apologies I don't know you right because at that time I haven't introduced you the context of short-term memory right yet so uh probably you did some of you didn't hear this but that was the reminder that just popped up like 2 minutes later but apologies but I should have set it up my Alexa

13:38 device in advance but you you saw this right? All right. He he is approved that it actually works. So you can talk to him, right? You hear this at all right cool. Uh so another like uh incidental like realization was LLMs are stateless and I think this is no-brainer for everybody right so what I had to do is to create uh first of all I used this open source project called radius agent memory right that has support for long-term sorry short-term memory we call session memory right because they are session bounded and

14:09 they are supposed to be short-lived and TTL bazid uh link chain forj the way it works you basically implement that interface called chat memory store and then you create whatever persistent vector store you you want to create right so basically create a wrapper to call that either via MCP or via REST APIs I think originally I use the rest APIs and then I switch to MCP right so the way this session memory works and let me try to show here just I think it's good we're all developers right I think it's good for you to

14:42 have like an idea about how this work so if I refresh this those are all the interactions that I had with Alexa so far and as you can see here this is a JSON uh payload right uh so it's a forever appending forever growing JSON payload so this is how a session memory kind of looks like behind the scenes uh so this is the radius edited memory architecture in a nutshell bas so basically what it does is a thing wrapper on top of radius and also uses LLM behind the scenes to create long-term memory memories out of

15:18 short-term memories in a background. So as you keep having conversations, the reg agent memory keeps like using the LLM to what is important here to realize to uh promote to long-term memory. So it kind of does this automatically, right? Um and this is how I you can associate it like uh for for with link chain for day. The interesting part here this implementation of a shaft memory here is how you can get a hold of the last always the last 10 message that were exchanged between the user and the system right so you can

15:50 kind of create an idea of what I have said before but one of the problems that I found with this is that uh so yeah I was able to fix this like I started remembering uh my name but one of the things that I realized with this are first and foremost at the time the building implementation of that um chat memory was kind of a malfunction because uh LLMs also make birectional calls between tools right so that implementation was causing a forever loop of calling all my tools so as you can imagine the timeout for Alex is 8

16:24 seconds and then it started like having more than 10 seconds because of the forever loop so I had to replace it with my own implementation that I've call working memory chat that was a simple pass through between the message it didn't get a hold of the last 10 like a buffer and it was kind of a forever growing the chat memory. So it solves two problems at a time.

16:43 The first one was the forever loop and the second one is that my conversation would be there kind of a forever forever uh between codes because session memory you associate with a TTL. So my TTL was five minutes. So kind of a five minutes was an okayish time for you to remember the last thing that you have uh talked to, right? But then if I would ask Alexa 5 minutes later what was my name?

17:08 Okay I don't know that right because the TTL had expired so it had no recollections of who I was anymore. So that was the trigger for me to start implement what long-term members LTMS right so the context must use long-term members as well. So the way you implement this in lane chain forj right and arguably with uh lang chain itself in python uh you basically has this concept of a context retriever right a context retriever like a that's a direct reference for rag right um is is an abstract implementation that you

17:42 given an input you can create a list of outputs right so I create my own implementation of get long-term memories I create my memory server that was able to talk with my radius agent memory This in the end of the day, right, is a simple vector search semantic search call, right? Here's the prompt. Give me the top k results that would match this on a vector search level, right?

18:05 Using hnssw uh index type, so highly efficient at scale, not necessarily very memory hungry. And then you would retrieve all the recent memory. So I can correlate this and try to provide a more uh relatable uh answer, right? So yeah, I was able to start like okay I have an appointment next week and then even many many minutes later if I ask the same question again yes you have a dentist appointment next week.

18:31 So can you see how my experiences was increasing as I was putting more information and functionalities at the device but didn't that didn't come for free at first because first my naive implementation was I would just put the LLM here and the LLM would take care of everything. So I think you're getting a feeling of how when you start like a dealing with a more humanlike experiences all of those things have to be a little bit more intentional right doesn't come for free uh so another lesson learned about that is that um

19:05 first and foremost I start using this kind of Alexa skill like inside the house so my son and my wife start using as well for their own memories their own interactions so I have to figure out a way to use like a multi-tenency Right? So I ended up having to use another technique of the AI word which we call uh vector search with pause filtering. Right?

19:24 So what is vector search post filtering? You basically do a vector search you retrieve a bunch of results and then out of those results you're going to do a post filtering basin on a metadata which is my tenant ID and actually owner ID at a time. Right? So Alexa gives you you know when you approach Alexa and then it kind of detects bas your voice or your picture who you are.

19:46 So that internally is an idea that I was kind of using for do this futuring right. So I had to do this otherwise the questions that I would ask could be answered by members from my wife right or my son which is 15 years old and be quite dangerous these days right so I you have to separate this tenants over there and secondly uh probably this is going to be like a very simple thing to understand but like how many should be set to the top k five 10 15 h how much is enough memory to use to formulate a right answer right

20:18 so this is what I I ended up using Another use case was not necessarily everything has to be about memories. So, uh, what I mean by this is that for this is was an actual use case. Dad, can I actually teach Alexa how to kind of decorate my the the code that you keep changing for the garage door for how do I do this? I could create my own memory like Alexa, remember this?

20:41 Or I could maybe fill a PDF document with a bunch of instructions about everything that has to do with the house, right? So I ended up creating something that I call it knowledge basis, right? Which is yet another category of long-term memories, right? But knowledge basis would require uh a little bit of more elaborate like usage of how I would implement this, right?

21:07 So now I would have like two retrievers like one to retrieve the long-term memories, the user specific memories, right? and another one to actually retrieve all the general knowledge base and as you can see here is shorter because because of the slide length but it would have a more descriptive description that's redundant I know but a description of what each one of those use cases kind of a do and as you can see here I use in link chain for the own LLM to decide based on the question okay should I retrieve this bas

21:42 it on the user memories or bas it on knowledge base depending of the nature of the question, right? So if you remember the nature of the question, right? What is the wireless garage door to for the touchpad? Oh, you can unlock your garage door with the code 7121. So that was the result of the null debate. I changed this code already. So this is not a secured concern.

22:03 Okay. Um so this is interesting because uh first of all like I you see how I am using the LLM with a yet another use case not necessarily to answer question but as a subsidizer to actually sustain an ongoing answer right so that was kind of a an inline call to the LLM just to make the actual call to the LLM so in order to to decrease the knowledge basis so this is another example of you have to craftedly and intentionally designing your implementation, right?

22:36 Um the other use case was a very interesting one was that all right uh ask my doctor to recall who was John do I have met last week John do was blah blah blah and then uh if I you I use a subsequent question uh where he lives because I am in an ongoing conversation right human beings would be able to capture real weakness but then Alexa would say who lives who where lives who I just told you don't go so that kind of a confusion has no longer to do with memory but has to do with how the LM is going to uh do like some

23:10 associations based on the context right and link link chain forj terms you use a tech which ended up real I realize that is a technique called compression of a quer transformer right so you can use the actual LLM given the prompt and then it can decide to do things like this so this this replacement and associations on the fly so he in this context text is John Doe.

23:39 Right? So this is one good example and compression is another technique of context engineering that for a human being it would be so natural but you have to teach those things when you interact with your AI services right so that's was another uh practice from context engineer that I implemented using something that was like pre-built uh with Alexa with sorry with lane chain forj that was was interesting uh for example can you confirm what is my favorite color and then oh yeah your favorite color is black which is true

24:14 right and you also write live writing code in Java okay why he answered that right can you guess why he did answer that as what I've explained before about when you do a vector search it retrieves the top key results before answering the question so basically right ve vector search is not an exact search the name tells everything right so by the time I actually asked this question about favorite, right?

24:41 Two entries had to do with being favorite, the favorite programming language and the favorite color. So those two entries came with the results and became part of the context that the LLM uses to respond to question, right? So I can't blame the LLM. It gave a very complete and elaborate answer. Yeah, you write you love writing code in Java and you love the black color, right?

25:03 But again for the human experience right you have to solve that problem. So how did I solve this? Uh yeah the model dumping everything that retrieved it needs to be a little more selective right. So how did I solve this problem? I actually used another technique that's not necessarily popular in the context of context engineering. It's more is more popular from the context of AI in general AI engineering but is a technique called few shots design.

25:33 Right. So what is few shots design is when you give a set of examples positives and negatives contributors to the LLM in a form of a system prompt right so you can actually better explain the notion of intention from the user in that particular domain model that you are designing right so first and foremost I had to use a little bit of a more structured uh context so you can link forj allows you to use this with a context injector Right.

26:02 So you can create actually a straighttory input for this and then you can provide there was a bunch of them like I think I created like 17 in my original implementation but of course I couldn't put the 17 here in my slides but this is one example of okay giving the question what programming language do I use the context would include favorite color birthday October the 5th enjoys Java so it would select this right so that would be almost like a hint to the LLM about how to behave not to think but how to behave right so

26:35 it's more like a behavioral engineering type of thing the user of few shots right so when I implement few shots I kind of I start having more objective answers like which is cool right which is ended ended up with what I was having right uh another important one which is kind of a variation of this confusion that LLM does every time and now and then uh Alexa what is my priority for the week right priority right now.

27:02 So adding support to Jason and going to the dentist are my priority access right now. >> Now is the time that I can put you on mute. So every time the Alexa wake up word is going to trigger, right? So in this example here, obviously like I since I mentioned the word PR, right? I assume as a human that you would understand that I'm talking about pull requests, right?

27:28 So if I'm talking about pull request, it would be logical that the answer would be adding support for JSON, right? Not going to the dentist because this is not a pull request, I think, right? You don't open a pull request just because you're going to the dentist, right? Uh so this one I had to kind of a create a more like a teachantic meaning of the context itself, which is yet another challenge that I have to undergo.

27:55 Right? So in this case what we're talking about here like the technique for context engineer that I ended up using was uh reranking right so it's a very popular technique but the reason why I had to use reranking was uh remember before I had those bunch of uh bunch of memories that was retrieved out of the top K results right and I have to teach the LLM about the reasoning about uh importance and priority right So I've I've solved this with the few shots technique, right?

28:27 But here there might be two or more memories that are correct and they're supposed to be used in the context but some of them would take higher priority. So it was about the ranking or reranking as we call it right. So luckily link chain forj also has support for reranking. basically you instantiate something called scoring model right and then you in this case I use khir right which is I don't know if that changed but at the time was the only like open uh LLM model that actually supports ranking and provides an API

29:02 for that I don't know if there's something new but I'm still useful here because it works very well and super cheap right and then I ended up associating this with this uh construct called content aggregator so which kind of makes sense you aggregate content Right? But all of this happens in flight if you think about it. Right? So the actual invocation to the LLM to answer the question haven't happened yet.

29:28 So all of this is happening like a in pre-flight mode. So you create a content aggregator that's going to pick up all the top K results from the vector search. It's going to perform a ranking and a minimum score which is kind of a 80%. Right? So I had to kind of manually calibrate this. I played with different numbers at the time and then you start you start seeing this that my pipeline execution now I have a context injector I have a query transformer for the compressing thing I have a query router to decide to know

29:56 for the knowledge base or the long-term memories and now I I also have a content aggregator to perform reranking right and again it all started with a simple assumption I just want the L&M to answer the questions right so That is a very good example and what I've mentioned before I'm going to actually stress this term throughout the presentation how when you apply context engineering things has to be deliberately intentional right so you kind of have to mix everything intentional as you as you implement so this was the

30:30 example of what I ended up getting as an answer after implementing the technique of reranking which is another very popular within the scope of context engineering Okay. Uh two lessons learned about this exercise. The first one was a more like a operational one because I was the remember I told you that the back end of Alexa the skill is a lambda function right on AWS.

30:54 So lambda function has this kind of a limitation. If you upload your LMA function directly uh through Terraform or through the console whatever uh the jar file has to be top of 50 megabytes right and if you upload what which I ended up using this approach which is using S3 buckets there's a maximum limit of 250 but none of this was enough because my original intention was not to use cohhere right I ended up using cohhere because of this problem which was uh I use it one of those like uh MES Marco M LMM L6 models which

31:32 is Ominx format right so you could right conceptually you could deploy your own models just to do reanking and that model specifically solved all my problems that I want to use but the jar file ended up like being like twice what is supported by the actual uh AWS infrastructure so I couldn't use it right so that's why I have to kind of a switch to let me use a service which is kind of a coher lightweight.

31:56 I don't have any dependencies and then my jar file ended up being like probably it's like less than 100 megabytes which is very lightweight right and also something that I've noticed as well and probably this is a heads up for everybody that doing this process if you change your models this has a direct impact on the scoring model or mean score that you're going to use for the re ranking uh technique right oh Ricardo so you're saying that the same 80% that you were using with like MS Marco didn't work with coher

32:29 exactly right. So I have to calibrate those numbers. So yes changing the models have a direct impact on the calibration process that you do for the reanking process. It's almost like a oh yeah on this release I solve all the problems that can close this issue here on GitHub and then the issue goes back again just because someone changed the model. That's the type of thing that it will happen like systematically.

32:51 Right. Right now let's go to chapter two. Pause for the pictures. Thank you. Chapter two which is talking about cost. Uh before I actually move forward to this like have you noticed that as I was implementing and solving systematically all the problems with context engineering I ended up using LLM more and more and more and more and more than just the actual call to answer the act the questions.

33:22 Right. So that that there is a prelude of what we're going to discuss here in this section right. So uh I did an analysis at the time right and I mean this whole this whole experiment like I said in the beginning it was to check whether the agent member server which is an open source product that we've developed would be suitable for use cases that would be stressed by human interaction right and it was right not virtually everything was kind of have to be fixed outside the registed memory server that was the

33:55 conclusion at a time but uh one of the realizations that I had was for a minimal experimentation, right? I basically was using me, my wife and my son at the house, right? The average number of queries that we would have per day would be 10 per user, right? So cost with the LLM was a flat cost with like $4.20. Okay. Right. For per family. Uh cost with Coher.

34:21 Coher is super cheap, right? So basically like 10 quirs per day for one user, right? Uh less than a dollar, right? But now let me remind you of a mistake that I did. You just charted in the beginning. I kind of planted that seed when I was explaining short-term memory. Can you remember where was when I was explaining that? Oh yeah, I used that built-in implementation with the uh uh from lang chain forj to keep the last 10 messages of the chat history and then I kind of oh it was having problems with two calling and all

34:56 of that. So I created my own that was simply a pass through that was a forever growing context right and the the experience was amazing because even like after up to five minutes if I actually uh ask any of those questions I would have very elaborated answers because we still remember the context but context the context growth exactly uh different from the LLM cost and the coher cost for reanking was not linear was exponential Right?

35:26 Because the context must keep growing for every single interaction. And believe me when I tell you like because when the LLM makes multiple invocation calls to your tools, all of those if forever appended to the JSON data structure that I was using, man, the context and the number of tokens being consumed was very very high. Right? Obviously, again, if you think, oh, Ricardo, you're so cheap.

35:49 We're talking about like a cent of dollar for one user. Yes. Yes. for a small family of one user or three members of a family. But what if it put at scale right at scale that would become a very very serious problem right and I I did some extrapolations here and like I said it would be like an exponential growth of the cost. So what we want is constant time right we want flat and constant time fixed cost right so um this type of question here that that's what I was explaining before you add a question data to the

36:25 context and then after some elaborated answer you keep adding all those answers to the context along with every garbage of uh not garbage baggage of information garbage is a very negative word uh baggage of information that comes with LLM2 calling and all of that right so what I actually did to solve this ish I'll explain this why ish at the time so there's a built-in implementation uh on link chain forj and link chain as well called token window shet memory right so how this implementation works you basically give

37:03 your first of all you have to kind of a set what is the maximum amount of tokens supported by that model right Um I think at the time I was using GPT3.5 at the time when I started developing this and if I remember correctly I think the month the maximum number of tokens was 4,96 so that was the number amount of tokens supported by that model on that time right so that was the number that was provided on that constant over there right and also you also have to kind of provide a very specific to the LLM implementation

37:38 about account estimator Right. So what is account estimator? It's kind of a teach the teach how the specific LLM counts tokens because believe it or not when I switch this model for OpenAI for example for entropic how it parses all the information about tools and human calls and AI calls is different. The data model the actual payload is different. Right?

38:03 So the strategy of counting varies for every specific model. Right? The point I'm trying to make right given this is that this token window basically like a create a cap out of this max tokens. So by the time you start like a getting closer to 4,96 it would cap, right? And you would behave similarly to the very naive implementation of keeping the last 10 or last any uh uh as a buffer, right?

38:31 But here it would like a trim all the old ones and keep only the recent ones but better on the amount of tokens not the amount of messages which is an improvement if you think about it right. So at least that way I would solve my cost problem which is keeping the costs flat right that would be a pragmatic way for you to actually solve the cost problem and turn the cost for exponential to flat good.

38:57 However, what is the problem that actually start happening with this? Can somebody guess? >> Losing context. Thank you. Exactly. Because the trimming that would cure it would basically have this kind of a built-in rule which is trim everything that is old. But what if that single message that is indeed is old but was very relevant to the conversation, right?

39:23 So we could not simply start deleting arbitrarily like this implementation. So that was my first uh attempt to fix this. Recently I use another technique from u also from context engineering called summerization right so basically I what I do I created my own implementation of this which is the chat memory but basically uses another LLM call to here's the context that I have so far extract everything that is relevant to this conversation that is either a fact or a preference or a strong statement and summarize it Right.

40:01 So it was kind of a I I kind of took a look in the code here that how this uh implementation does. I also use this notion of uh using the token estimator because I think it's pretty cool right in the end of the day I have to maintain the cost flat but also having a more intelligent implementation of okay it cannot simply delete messages arbitrarily right it has to delete messages consciously at very best right so this is how I fix this problem.

40:26 Uh the other problem that I would start having actually let me go back here for a second. Uh that this is one of the most like interesting things interesting and worth things that I ended up discovering. uh when you actually make calls to the LLM like uh oh uh who is coming for dinner tonight like both John and Susan are sometimes either you or some other member of the house would ask the same question right whereas the answer is exactly the same right perhaps they're going to ask this differently perhaps they're going

41:00 to like ask this in another session but the answer itself didn't change the the immutable vector entry lot there on that JSON payload start and rad is is essentially the same. You haven't changed. So why not why I have to incur into another call to the the same code. And for those of you that know how radis work and how radis became famous like throughout all the years what we do best at radis we cache data right.

41:27 So basically I start like using this this is another technique that became part of the umbrella of um context engineering that we call sementic caching. Right? So what is the difference between typical caching? Typical caching is a point in time very objective very direct hit miss hit or miss right given the key do a look up if it's there go to the back end populate the back populate the back end and retrieve the answer right the difference with semantic caching and that was the ugly part of the implementation was uh

42:06 you go to the cache same thing right but it evaluates the semantic meaning of your query Right. So in here for example, can we agree that those are the same questions just as it differently is the same question, right? That would return the same answer but they were asked it differently. But so that's why semantically they have the same meaning. So because simantically could have the same meaning they have to perform a cash hit, right?

42:32 They couldn't be a cash miss right. So the problem with this it solved the problem. So I I reduced most of the 75% of my redundant calls and more than half right which is good. That's the positive part. The negative part is that sometimes even for a small fraction of words or even my case because English is not my first language. I'm from Brazil originally.

42:58 Sometimes I would ask it with my this terrible English accent. It would do a cash hit on the semantic cache. Right? because of this kind of a very particular unique kind of a words. So he ended up like a asking from the cache instead of going to the LLM which is kind of annoying sometimes. So semantic caching is another area that you should pursue and the p on the implementation of context engineering.

43:24 But one of the danger of semantic caching exactly this you have to calibrate the just like the reranking model you have to calibrate the how this cache hit and miss is going to trigger all the time. So what we actually do on this implementation called link cache it's called link cache because it was originally to be said that the cache for link chain.

43:45 So that's why I originally call link link cache. So if you go to the link chain documentation, you're going to see that uh dimensions of this uh we have this parameter that basically calibrates how the the the the assertion goes to the vector similarity search that's going to be performed against the cache. Right? So that's what how we use it. Uh and this is a plain old what is how how many of you can tell me what is the name of this pattern here that I've implemented here?

44:17 start with cache and ends with aside cash aside right so basically it's a cash aside implementation no big deal so basically you go to the clash if it's not there go to the LLM right to wrap up because we have virtually four minutes am I right fourish minutes context engineering I would say that this is not just like a fancy thing that you would just put for declaration purposes in your implementation like a feature like oh it would it would be good to have this or not.

44:49 I would say that for any serious implementation that you're going to do with AI that has to provide humanlike experiences for users, I think it's become part of the architecture, right? So my my main lesson learned from all those implementations that I did with this Alexa skill device is that uh there was a bunch of things like you can see that the amount of things that I had to change around the implementation.

45:13 There was a bunch of things that I should have thought sooner in terms of architecture, in terms of design in order to actually start implementing things and not keep changing on the fly, right? I did the reactive mode, but I could have been proactive, right? And that's one of the reason why I came here today to share all these lessons with you. So you can also can be proactive, right?

45:34 So my impression is that when you get context engineering ingrained in your architecture, everything else comes much easier. I think it's more like it's more natural right wrap up I would like to give uh make two invitations for you the first one all this implementation that I have done you can find here on GitHub right so clone the rapple play with it I think ultimately at least if you want to learn how to develop Alexa skills is all there in Java at least right and then you can all the context engineering patterns

46:08 are also there right uh obviously like this is going to be like a minimum lish version what what I have built at home because what I have built at home also have like some very specific like I I kind of have to make sure that I read my blood pressure every day. So I kind of create a tool that actually res keeps storing all the blood pressure readings and I can say to my doctor when I have the appointment say oh this is the average that I have been through in the weeks right so this is that and related to that this is

46:34 my LinkedIn profile I would love to continue this conversation that specifically if you try this implementation yourself and then you kind of have your own discoveries and old design discoveries and uh yes thank you for being here today I wish that you enjoyed and the rest of the conference and I really hope that this presentation was useful for you because this is what matters right thank you everybody but but let me repeat this question because it was very very good uh so his question was when I do the summarization

47:09 right because in order to trim the messages I have to summarize so I can keep the the cost flat if I ever use a different LLM to get better answers uh I didn't I use ended up using open AI as well but honestly I haven't checked if they are able to provide better answers. So that's a good thing to check. Yeah, that's a interesting thing.