← All transcripts

The Four Types of Memory Every AI Agent Needs Transcript, AI Summary & Key Points

IBM Technology · May 26, 2026 · Education · 10:41 · EN

Watch on YouTube

AI Summary

AI agents use four distinct types of memory, mapped by the CoALA framework: working, semantic, procedural, and episodic memory. Working memory is the current context window and is fast but volatile and limited. Semantic memory stores persistent facts, rules, documentation, and conventions. Procedural memory stores skills and instructions for performing tasks, often using progressive disclosure to avoid overloading the context window. Episodic memory records distilled experience from past interactions and decisions, helping agents improve across sessions. Simple reflex agents may need only working memory, while coding agents may need all four. Persistent knowledge and accumulated experience are what distinguish an agent from a chatbot.

Key Points

  • The CoALA framework, Cognitive Architectures for Language Agents, maps four types of agent memory: working, semantic, procedural, and episodic.
  • Working memory is the agent's context window, containing the current conversation, system instructions, and loaded files or data; it is volatile, limited in size, and can degrade when overloaded.
  • Semantic memory stores persistent facts, rules, conventions, documentation, and other general knowledge the agent needs to avoid repeating mistakes.
  • Procedural memory stores instructions for how to perform skills, with progressive disclosure loading full instructions and supporting resources only when needed.
  • Episodic memory stores distilled and compressed experience from past interactions and decisions rather than necessarily retaining complete transcripts.
  • Forgetting is an engineering challenge because episodic memories can become obsolete, especially when projects or user circumstances change.
  • A simple reflex agent may need only working memory, a password-reset agent may need working and procedural memory, and a coding agent may need all four.
  • Persistent knowledge, accumulated experience, preferences, and remembered mistakes help distinguish an agent from a chatbot.

AI in practice

Agents

  • coding agent — Work on coding tasks using current context, project knowledge, reusable skills, and experience from previous sessions. 2 held 09:26

Tools & resources

2 items

ANo. 1130
AIAINotes.us AI product

Agent Skills

Open source · addyosmani/agent-skills

Agent Skills is an open-source collection of engineering workflows for AI coding agents, developed by Addy Osmani. It packages senior-engineering practices into skills covering the development lifecycle: defining specifications, planning, incremental implementation, testing and debugging, constraints, code review, web-performance auditing, code simplification, and shipping. Slash commands such as /spec, /plan, /build, /test, /review, and /ship activate the corresponding workflows, while skills can also activate automatically based on the work being performed. The /build workflow can generate a plan and implement its tasks in an approved pass, while retaining test-driven verification, individual commits, and pauses for failures or risky steps. It can be installed with the skills CLI or integrated into supported coding agents including Claude Code, Cursor, Codex, Copilot, and Cline.

Mentioned in
4 videos
Kind
AI
CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

Mentioned in
74 videos
Kind
AI
🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of The Four Types of Memory Every AI Agent Needs — IBM Technology (10:41). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 AI agents have different ways to remember stuff and each serves a different purpose. So let's take a look at the four main types of AI agent memory from some pretty foundational stuff to what I think are some quite interesting emerging areas. And I think it's really, first of all, worth considering how we do it. How does human memory actually. We can think of human memory as having first of all short-term memory.

00:35 So that's the stuff that is active in the brain right now like what I'm saying at this very moment. That's one type of memory but there's also a type of memory called factual knowledge. So this is things like the company security policies that you remember or it could be facts like Python is an interpreted language. Then there are learned skills. Like, I don't know, writing backwards on a sheet of glass, for example, which I am totally doing here.

01:12 There is absolutely no camera trickery involved. And then there is personal experience. Like the time I spent three hours debugging a Kubernetes cluster only to discover... I was pointing at the wrong cluster the entire time. Seriously, that was three hours of my time. Anyway, anyway, it turns out that well-designed AI agents, they also need these three types of memory or these four types of memories that I've got here.

01:44 And there's actually a well-known framework for this. And it's from a Princeton research team and they gave it the name of CoALA. That's Cognitive Architectures for Language Agents, and CoALA maps out four distinct types of memory that agents need. So let's walk through each one and see how they actually work in real agentic systems today. So type one, that is working memory.

02:14 This is the agent's context window. It's everything the agent can see right now, the current conversation, if there's any system instructions, they'll be in there. If there's any files or data that have been loaded into the prompt, that's where they'll be as well. So it's really kind of the scratch pad. And the analogy everybody uses for this is this is just like RAM, random access memory.

02:39 It's fast and immediately accessible, but it's volatile. When the session ends, it's gone. And it's also limited in size. I mean the- the biggest context windows available today are pretty big. I mean, it could be like one million tokens or even more than that, but that still has a ceiling and try to stuff too much in there and performance is gonna degrade as the model starts losing track of things that are kind of buried in the middle of the context window.

03:13 So every agent has working memory, but then so does every chat bot, it's just the context windows. So the question is... What else do agentic systems need? Well let me add to that list number two semantic memory and this is the agent's knowledge base, so semantic memory stores facts and rules and conventions, documentation and in the academic literature this often gets described in terms of things like vector databases or as knowledge graphs, and yeah, those are real implementations, but, in a lot of production agentic

03:50 systems today, semantic memory is something much simpler than that. It's just simply Markdown files, .md files. So take Claude code as an example of this. So it has one of these Markdown Files. It's one is called Claude.md and that sits in the root of a project. And that file contains the project architecture, the coding conventions, the build commands, what frameworks to use, and also what not to do.

04:25 And that far gets loaded into the context window at the start of every session. So semantic memory tells the agent what it needs to know in general. And without it, the agent is, well, it's kind of destined to make the same mistakes over and over again, because it has no persistent knowledge to draw from. Working memory, semantic memory, what else is there?

04:51 Well, number three, that is procedural memory. Now procedural memory is how the agent knows how to do things. And there's an open standard for this that's called agent skills. And it uses a file format called skill.md. A skill is just a folder with a markdown file that describes the skill and what that skill does and some step-by-step instructions for how to perform that skill and it could be anything from creating a PowerPoint presentation to running a structured code review.

05:29 Now skills use something called progressive disclosure so the agent doesn't load all of its skills into the context window or I guess I should say into the working memory at once because that can blow through the working memory budget if there are a lot of defined agent skills. So instead the agent just sees a lightweight index, which is just the name and the description of each available skill.

05:57 So maybe that's a hundred tokens per skill. Then when a task comes in that matches one of these skill descriptions, the agent loads the full instructions and if the skill references other stuff like other files or templates or scripts. Well, those only get pulled in when the agent actually needs them during execution. So the agent advertises what skills it has, it loads the instructions in when they're needed and then executes with the additional resources pulled in as they're needed as well.

06:27 And all that is quite different from semantic memory where the knowledge is always present in context. All right, number four is episodic memory. Episodic memory is the agent's record of what happened in past interactions and past decisions and what it learned from them. Now a naive implementation of this is just to save every conversation transcript and then just search through them as you need to.

06:58 And that technically counts as episodic memory but often that's not very useful. So what production systems actually do is a bit more distillation. So as the agent works across sessions it kind of accumulates notes for itself, but it doesn't save everything. It decides what's worth remembering based on whether that information would actually be useful in a future conversation.

07:26 So the result is distilled or compressed experience. So things like last time we debugged the auth module, the issue was in the middleware layer. That's something that's a lot more useful to remember than just a full transcript of a 45 minute debugging session. And this is where memory starts to kind of genuinely look like learning because the agent is gonna get better over time.

07:52 But episodic memory is also the hardest type of these to get right because what do you delete? When does information become obsolete? If a user changes jobs, do you keep the old project memories around? Or should we forget them? Well, humans are actually pretty good at forgetting. I do it all the time. And as frustrating as that can be, it can be quite useful.

08:17 But for agents, forgetting is an engineering problem. So four types of memory, working, semantic, procedural, episodic, but not every agent necessarily needs all four. Let me give you an example. So let's say, we're building a simple reflex agent. So that's something like a thermostat or it's just like a basic routing bot. Doesn't need all four. It might only need access to working memory, the context window, and that's basically it.

08:51 Now, if we take something a little bit more complicated like a customer support agent, but one that's still fairly simple and narrow, like an agent that resets passwords, for example. That will still have access to the working memory, of course, but it probably also needs access to procedural memory as well because it needs to recall the password reset skill.

09:18 But that might be it. Whereas if we take a look at something like a coding agent, it probably needs access to all four, so it certainly needs access to the working memory, the context window, but it also needs the product knowledge it would get from semantic and then the skill system from procedural and also the auto memory from episodic that learns across sessions.

09:49 So memory is really what separates a chatbot from an agent because a chat bot gives a response, but an agent can give a response shaped by persistent knowledge. Accumulated experience. It remembers the project. It remembers preferences and a good memory architecture also remembers the mistakes so we're not destined to repeat them which honestly would have been wonderful if an agent had told me about that Kubernetes cluster before hour three.

10:21 So four types of AI agent memory. Which of these are you using in your own agentic workflows? Thank you.