← All transcripts

Why Bigger Context Windows Won't Save Your Agent — Elizabeth Fuentes Leone, AWS Transcript, AI Summary & Key Points

AI Engineer · yesterday · Science & Technology · 15:54 · EN

Watch on YouTube

Answer

A bigger context window won't save your agent — past a point the model forgets the middle of the context. The fix is context engineering: externalize, select, compress and isolate, using memory pointers, conversation managers and capped tool calls.

AI Summary

Bigger context windows do not save agents: when tools return large amounts of data, the context window overflows and agents hallucinate, and models lose the middle of long context (the 'attention curve' — only the start and end are remembered). Context engineering — giving the model the information it needs when it needs it — fixes this, via four strategies: externalize (move big data to persistent storage with memory pointers), select (retrieve only what is relevant), compress (summarize and compact), and isolate (separate context across agents). Using the open-source, model-agnostic Strands Agents framework, the talk walks through sliding-window and summarizing conversation managers, short-term, long-term (vector database) and graph memory, memory pointers that keep large tool outputs out of context, sharing pointers across a swarm of agents via invocation state, capping tool calls to stop endless loops, and async handles for slow MCP tools and external APIs, ending with anti-patterns (context stuffing, naive summarization, context pollution) and a cheat sheet.

Key Points

  • Agents break when tools return large amounts of data: the context window overflows and the agent starts to hallucinate
  • The attention curve: when context is very large, the model remembers only the first and last parts and forgets the middle
  • Context engineering means giving the model the information it needs when it needs it; a poisoned context window can also carry injection problems
  • Four strategies: externalize (move big data to persistent storage via memory pointers), select (retrieve only what is relevant), compress (summarize and compact), and isolate (separate context across agents)
  • Strands Agents is open-source, free and model-agnostic; one line of code creates the agent loop, and swapping in OpenAI, Gemini or Anthropic models takes one added import line
  • Conversation managers in Strands: sliding window (keep only the most recent messages), summarization (summarize older messages while preserving the last ones — e.g. summarize 50% and preserve the last four messages), or a combination of both
  • Three memory types: short-term (recent conversation history), long-term (a vector database for past sessions), and graph memory (entity graphs and traversal of how things connect); the three can be combined
  • Memory pointers keep large tool outputs (like logs) out of context: one tool stores the data and returns an ID kept in agent state; another tool retrieves the logs by that ID only when needed

Tools & resources

8 items

ANo. 5315
AIAINotes.us Tool

Amazon API Gateway

aws.amazon.com/api-gateway/

Amazon API Gateway is a fully managed API management service from Amazon Web Services for building, publishing, maintaining, monitoring, managing, and securing HTTP, REST, and WebSocket APIs. In the cited agent architecture, it is an external service whose response-time limits motivate handling slow operations asynchronously.

Mentioned in
1 video
Kind
Other
ANo. 0096
AIAINotes.us AI product

Amazon Bedrock

aws.amazon.com

Amazon Bedrock is a managed AWS service for building and running production generative AI workloads with hosted foundation models and APIs. It provides access to Amazon's and third-party models with managed infrastructure and tooling; customer data remains within the customer's VPC, prompts are not sent to model providers, and most inference traffic runs on AWS Trainium chips.

Mentioned in
4 videos
Kind
AI
CNo. 2302
AIAINotes.us AI product

Claude

anthropic.com

Claude is a family of large language models and service developed and hosted by Anthropic, an AI safety and research company building reliable, interpretable, and steerable AI systems. Available through Anthropic's web chat products and a commercial API, Claude accepts text prompts and structured inputs or file attachments via the API and returns natural-language and code outputs for tasks such as coding, summarization, document Q&A, and general business and productivity workflows. It is used inside Anthropic products such as Claude Code, by third-party services via Anthropic's API and enterprise agreements, and as an alternative model provider for the model-agnostic Strands Agents framework.

Mentioned in
2 videos
Kind
AI
FNo. 4047
AIAINotes.us Tool

FastAPI

Open source · fastapi/fastapi

FastAPI is a Python web framework for building APIs using standard Python type hints. It is built on Starlette for web functionality and Pydantic for data handling, providing validation and conversion for request data from JSON, path and query parameters, cookies, headers, forms, and files. Type declarations also generate OpenAPI and JSON Schema documentation, including interactive Swagger UI and ReDoc interfaces. The video describes FastAPI routing requests from a mortgage application to five specialized agents.

Mentioned in
2 videos
Kind
Other
GNo. 0098
AIAINotes.us AI product

Gemini

gemini.google.com

Gemini is Google's AI model and assistant accessible through gemini.google.com. It processes mixed media including video, audio, images, and text, and can convert recordings into SOPs, summarize meetings, extract decisions, and draft follow-up emails. In research and development contexts, it serves as an alternative LLM backend for tasks such as adversarial audits of requirements and resume evaluation.

Mentioned in
34 videos
Kind
AI
NNo. 3296
AIAINotes.us Tool

Neo4j

neo4j.com

Neo4j is a native graph database. In the cited discussion, it is used as an example of a graph-database implementation, which Mike Stonebraker characterizes as less performant than relational representations of graphs.

Mentioned in
4 videos
Kind
Other
ONo. 0105
AIAINotes.us AI product

OpenAI

Open source · openai

OpenAI is an AI model provider whose language models are available through an API and can serve as one provider in multi-model routing architectures. The videos describe its models being used for AI replies and extraction, Amazon-review analysis and customer-service email generation, legal workflows, complex software development and debugging, and experimental or auxiliary agent tasks.

Mentioned in
21 videos
Kind
AI
SNo. 4385
AIAINotes.us AI product

Strands Agents

Open source · strands-agents/harness-sdk

Strands Agents is an open-source, model-agnostic AI agent SDK and toolkit maintained by AWS for Python and TypeScript. It runs the agent loop in the user's process without a hosted control plane and uses a model-driven architecture based on system prompts, tools, and a selected model. Its Strands harness provides a preconfigured agent with tools, a one-line agent loop, context management, sessions, and memory, while the Harness SDK lets developers assemble agents from their own models and tools. It supports providers including Amazon Bedrock, Anthropic, OpenAI, Gemini, and others; MCP connections; structured output; lifecycle controls; multi-agent patterns including delegation, swarms, graphs, handoffs, and workflows; conversation managers for sliding-window, summarizing, and combined histories; short-term, long-term, and graph memory; memory pointers that keep large tool outputs outside the active context; agent state for sharing pointers across multi-agent swarms; tool-call limits; asynchronous handles for slow MCP tools; streaming; guardrails; tracing; evaluations; hooks; telemetry; context management; and deployment support. Extensions include a command-line interface, sandboxed and virtual shells, evaluation utilities, research labs, samples, reusable agent instructions, lower-level SDKs for custom agent loops, and Strands Box, an open-source sandbox.

Mentioned in
6 videos
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Why Bigger Context Windows Won't Save Your Agent — Elizabeth Fuentes Leone, AWS — AI Engineer (15:54). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:12 Hi. Well, first I have to say I am Elizabeth. I am not Morgan. But the guy for the agenda they never change agenda. So we are almost the same but with different colors hair. Okay. So it's it's going to be similar. So today I going to talk about the infinite content window is a meet. Yes. Yes is because right now edge system breaks when tools returns a large amount of data.

00:44 Context window overflow and the engine start to hallucinate and start the recar is degrad. So right now context engineering help us to uh give us a better architecture for this agent application. So we have some techniques that allows us to improve the context the context window for this agentics application. So imagine that you have a agent that is responsible to uh supervise some application.

01:21 So this agent they have to retrieve log logs for that application. So every time this agent is doing the retrieve for the logs is putting data for the content window and every time this content window is going to grow. It's going to grow every time because when we have a one session and each session is going to retrieve the last session because it's going to remember everything.

01:47 That's the way that agents remember and they use the content window for that. In the past about probably one year ago before this agentic era we saw that just add more token that will fix it. Yeah, there's more bigger content window that is going to fix it. And today we know that is not true because we have the uh attention you curve. What is this? We have a lot of data, but if that data is super super big, the engine is going to forget the meter in the in the in the curve.

02:28 It's only going to remember the first part and the last part that we sent to that content window. And I'm talking here about context engineering. But I can move. But content engineer is giving the model the information it needs when it needed. Sometimes the agent doesn't need all the information inside that content window. Sometimes the content window is poison the engine because you can have a problem injection inside your content window and the content engineer help us to incorout the context window for optimal

03:08 outcomes and token consumption. So today I going to show you some techniques that can help you to improve the content window. And these techniques uh I manage for a strategy. We have we have the externalize because move this that this big data to a persistent storage. You can create a memory pointer to storage that data. I going to share the deck at the end.

03:37 So you need to you don't need to have to. Please pay me attention and because I going to share you all the information that I going to talk share. And the second one is select you can retrieve only what is relevant. We have compress that reduce token because you can summarize and you can compact that information. And the last one is isolate because you can separate context across agents.

04:04 and let's start. So the starting point conversation manager this strategy yeah we have some strategy that are a little complicated but we are going to start for the easy one that we can use inside the framework. So here in this strategy we are going to use a strand engine is a framework like we are a we maintain is completely open source yes it's free and is a yeah it's open source free hence it's completely model and so here with strands you can use three different strategy inside the framework so we have sliding

04:48 window that allows you to keep only the most recent messenger we have summarization inside the framework that allows you to summarize all their messenger and keep present one and the last one is a combination of two we have that you can combine summarize of older messenger with resent one so first let's see what is a strand a strand is this so you only need one line of code to create the agent loop so this line here create all the agent loop behind the scene.

05:23 You need to create the nose, the graph, nothing because the strands can do that for you and you have the system prom and the tools. And yes, this is simple because it's using Amazon better behind the scene. But if you want to use Lamba uh Gemini and tropics, you only need to add one line because you can do that. You import the the library for OpenAI for example and then you add the model line and that's it.

05:54 So let's return to our strategy. If you want to use summarization conversation manager that's the only line that you that you can add for this agent loop. So you have the summary radio that is in this case is 50% and you want to preserve preserve the last four messages. So every time that you invoke this agent in this session, this session is going to manage your content window in this way.

06:23 So it's going to keep the four last messages and it's going to summarize the other 50%. And yes, this is amazing. But what happen when you have a lot of memory and you don't want to lose you don't want to lose all the memory inside that content window. So you can use well always we have shortened memory because shortterm memory is when you have the recent conversation history.

06:50 So you can use this uh application in the shortterm memory. But what happen when you want to return and uh you for you uh you have a new session for this agent and you don't have short-term memory. You can have long-term memory and you can create a vector database for that long-term memory. You can have different types of memory inside that vector database with different prompts that can understand your conversation and say whatever you want.

07:21 And the third one is relation memory. You can use entity graph. You can go to the uh to the guy for neo forj that they can explain you more about that. You can use traversal baset or how things connect. So you can understand you can use a combination of three of those. You can have long short-term memory for the me the session that you have. Then you can retrieve the memory that is in the past the long longterm memory and if you have some relation thing that you want to keep in a different way you can add for that a

07:56 graph memory as well. Let's continue with the other one. So what happened with this application where a lot of logs that I told you you don't want to put that locks inside your content window every time that you invoke the agentic application that agent you can save that inside a memory pointer. So what how this work? You have the the tools. The tool returns you and a bunch of uh logs.

08:23 You can use another tools that is going to capable to took that logs and save that logs in a storage and you are going to retrieve an ID for that storage and for that using a strands you can have that using the agent state. In the agent state you are going to have all the information for the tools. Then you can use another tool that is going to capable to understand that logs and it's going to send um it's going to send to the agent hey the logs okay you don't need the logs to understand that but what happen is you

09:00 need the logs you can go say hey this is my ID come on retrieve the logs because I needed to know so this is something that you can build in your architecture and it's something like this you can have two different tools that's the way that we create tools inside the trans only with the tool decorator. So we have fetch application logs and if you can see there you have the tool context agent step state.

09:26 So we retrieve that information and we save that in a memory pointer. We create a memory pointer ID that is going to preserve inside our content window. But what happened is your application is super complicated and the tools is not enough. You can use ah if I forget that you can use a multi- aenting application. So you can create a a trans swam with multi- aents but you share the context window between agents is like add garbage to the other agents because no all the agents need the information between them.

10:09 So you can share the information between the invocation state pointer. So you can create an ID between that swam agent swam and you can share the memory pointer between the the invocation state between that aents and that's the way that we create the swamp in a strand agent and oh my god 9 minutes I going to be fast and the next one is where the next one is what happened when the engine uh want to stop never in sometimes we have this engine that stay in the loop like forever invoking the same tool the same tool again

10:55 and it's never ending so you can stop that with a minimum amount on invocation inside the loop so what is the problem I said the problem to sell more soul may be available so what does happen when you not when you you don't have a clear response when you evolve the tool So you can avoid that with the tool way with the clear response and with amounts of invocation for the tool and you do like this.

11:26 So you have a limited tool count. So you set max tool counts for this example is my tools search fly. You can you only can invoke that tool three times and the other tool. Yeah. Three time as well. So the agent never is going to do to do this uh internal loop. It's only going to invoke that tools three times. And the last one, yes, I almost there. The last one is uh is not something that help you to the context window.

12:00 is something that you can use to help your content window because sometimes we are invoking IMCP tools or IMCP or we have external API that never answer you at the time that you need it. Sometimes for example we have EP gateway AWS EP gateway it have a a maximum time that is going to receive the the answer. So if you don't have that time, it's going to uh like stop your application because the delay.

12:32 So you can avoid that using a a synchron. So when you use synchronous, your engine is going to wait forever for the answer for your NCP. But if you have an a sync handle, you can invoke your NCP, your external API. So then you can save that and you can retrieve the answer in the next invocation or when you need it. Yes, this is going to depend of the case that you are building and you can do that like this create MCP this is a function that use fast API and with fast API you only need to create the tools.

13:12 So I going to have two different tools. One is the star lob and another one is check job status and I did a lot of test using this uh a sync handle and the different is a lot. For example, if you have the different um the well the agent never stop to answer me. So because I wait then he can retrieve the information and to close this session six minutes we have the antipartis so why uh you have to avoid you have to avoid contact stuffing so you need you don't need to include everything and the kitchen you know you don't

13:57 have to include everything the content window because the engine doesn't need all the information there you can only share what a need to create the answer that your user or you're asking for. You can have uh native summarization. Sometimes when you do native summarization, the agent uh can uh you can uh forget what you need. If you're doing whatever superization what you may for probably in that summarization, you are not going to have everything that the engine is needed to answer you.

14:30 and context population. Again, don't pass everything to the agent. The agent doesn't need all that information. And what is a good scene here? So tool returns large data, logs, database, alert information, externalize, select and use a memory pointer. Many engines need the same large data. You can share invocation between stage and agent loops or tool hands.

14:59 You can have clearest st you know give a clear answer for the tool and limited to counts and a thin handle. Thank you. And of course if you want to learn more about stringent we have a huge boot there. We have a lot of swag. We have a little Legos for Amazon that you can have and see you there. We have the specialist and I am Elizabeth No Morgan remember. [laughter] Thank you so much. >> [music]