← All transcripts

How Many Credentials Should Your AI Agent Have? Zero. — Jim Clark, Docker Transcript, AI Summary & Key Points

AI Engineer · 3 days ago · Science & Technology · 18:13 · EN

Watch on YouTube

Answer

Zero. Credentials should never be inside the sandbox or harness; access should be managed through the MCP gateway and authorization grants.

AI Summary

AI agent safety comes from controlling the tools and context that flow into an agent rather than relying on the harness itself. Sandboxes should be modeled around task intent, with only the resources, tools, network access and work trees required for each task. MCP gateways provide one control point for tools, resources and prompts, keep harnesses interchangeable, and allow credentials to remain outside the sandbox. Cross App Access (XAA) can connect agent identity to existing authorization systems so agents receive authorization grants without storing credentials in the sandbox.

Key Points

  • Agents are running longer and with less supervision, increasing the need to limit what they can access and do.
  • An agent harness is a loop that receives context and produces tool calls, while MCPs provide additional context and tools for external work.
  • A sandbox should contain only the resources and tools required for the task, reducing the blast radius of incorrect agent behavior.
  • Newsroom example — Separate researcher, fact-checker and publisher roles into different sandboxes so unfiltered research is not combined with dangerous publishing tools.
  • Newsroom example — The researcher can access the web and write to a private temporary location, the fact-checker can access research and a fact database without network access, and the publisher can access filtered research and a publishing MCP without network access.
  • Coding-agent example — A coding agent may work for about an hour but need Git signing keys for only about 2 minutes, so signing keys should enter the sandbox only when the agent is committing.
  • Task-intent modeling — Agent capabilities should reflect the intent of the current task rather than the capabilities of the entire workflow.
  • MCP gateway — A single gateway endpoint can route all MCP traffic, resources and tool calls for each sandbox.

AI in practice

Used for

Agents

  • Complete long-running coding tasks and make commits along the way. 2 held 08:48

Tools & resources

8 items

CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

TypeScript
Stars
★ 149,337
Forks
25,438
DNo. 5112
AIAINotes.us AI product

Docker MCP Gateway

Open source · docker/mcp-gateway

Docker MCP Gateway is an MCP gateway and Docker CLI plugin for configuring, running, and connecting MCP servers to AI clients. It routes requests between clients and MCP servers running in isolated Docker containers through a unified interface, allowing tools, resources, and prompts from multiple servers to be managed through one gateway endpoint. Servers are grouped into profiles that can be connected to clients, exported, shared through OCI registries, and configured with tool allowlists. The gateway supports stdio and streaming transports, dynamic tool and resource discovery, server lifecycle management, logging and call tracing, secrets management, and OAuth flows. It can run with Docker Desktop or independently, and the repository is licensed under the MIT License.

Mentioned in
1 video
Kind
AI
MNo. 5110
AIAINotes.us Tool

Marp

marp.app

Marp, the Markdown Presentation Ecosystem, is an open-source toolset for creating slide decks from Markdown. Its format is based on CommonMark, with horizontal rules separating slides, plus directives and extended syntax for images, mathematical typesetting, auto-scaling, and theming. The ecosystem includes Marp Core, Marp CLI, and a VS Code extension; it can export decks to HTML, PDF, and PowerPoint through Google Chrome or Chromium. Marp is built on the pluggable Marpit framework, and the official tools and related libraries are MIT-licensed.

Mentioned in
1 video
Kind
Other
MNo. 5111
AIAINotes.us Tool

Mermaid

mermaid.js.org

Mermaid is an open-source diagramming and charting tool that creates diagrams and visualizations from text and code. It includes the Mermaid Live Editor and integrations with other applications, and is developed by an open-source community around the Mermaid.js project and its associated platform.

Mentioned in
1 video
Kind
Other
ONo. 5115
AIAINotes.us AI product

Okta Cross App Access

okta.com/solutions/cross-app-access/

Okta Cross App Access is a protocol for secure agent-to-app and app-to-app access. It enables independent software vendors to use an organization's existing identity provider and single sign-on systems to establish agent identity and obtain authorization grants, allowing agents, sandboxes, and gateways to access applications without embedding credentials in the sandbox.

Mentioned in
1 video
Kind
AI
ONo. 0214
AIAINotes.us AI product

OpenAI Codex

openai.com

OpenAI Codex is an AI coding agent from OpenAI available as a command-line tool (Codex CLI) that helps developers produce software. It can be used alongside Gemini for adversarial audits of software requirements and implementation plans, listed as a supported coding-agent or model option in several projects, and its logs can be joined with task and test evidence. The Codex CLI can also receive and answer requests from the Penako canvas.

Mentioned in
61 videos
Kind
AI
SNo. 5114
AIAINotes.us AI product

sbx

docker.com

sbx is Docker's command-line tool for building containers used as sandboxes for AI agents. It is presented as part of Docker's sandboxing approach for controlling what agents can access while they run, and can be installed with Homebrew.

Mentioned in
1 video
Kind
AI
SNo. 5117
AIAINotes.us AI product

sbx CLI reference

docs.docker.com/reference/cli/sbx/

The sbx CLI reference is Docker's documentation for the command-line tool used to build and run containerized sandboxes for AI agents. The tool supports sandboxing agents so their tools, context, credentials, and access to external systems can be separated according to a task's intent.

Mentioned in
1 video
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of How Many Credentials Should Your AI Agent Have? Zero. — Jim Clark, Docker — AI Engineer (18:13). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 [music] >> I guess I'm ready to start. Uh can everyone can I I actually I can hear myself. Yeah, so you guys can hear me. Uh my name's Jim. Uh I'm a engineer at Docker. Normally when I introduce myself in the slides these days, or at least what I used to do, was I would say, "Hey, I'm Jim. I uh like spaces, not tabs. I use Neovim, not Emacs." But, that's uh all irrelevant now.

00:41 So, uh I will say that I made this slide with uh I made this presentation with Claude Code. I used the Codex model. The MCPs that I that helped me write these slides were Marp and uh Mermaid. And I'm going to talk to you about MCP Gateways. But, I'm really going to be talking to you about AI safety. So, I'm going to kind of go through sandboxes, gateways, harnesses, and sort of motivate why we're talking about AI safety at Docker right now.

01:14 So, like probably the last 6 months has been similar for all of you as it's been for me. Agents are doing way more than I thought they were doing uh than I thought they were going to do. And I'm no longer going to sit in front of my laptop and say, "Yes. Yes. Yes. Yes." They're doing longer running things, and I like that. And as I get more and more accustomed to them doing longer running things, I sort of decide that my metric for me having designed a problem that an agent can sink its teeth into is that I don't

01:55 really need to give it a lot of supervision. But if I'm not going to supervise it, maybe I'm also a little bit worried about what it's going to get up to. So I think just to set context, this is what I think is so the sort of shape of the problem right now. We need to talk about agent harnesses, what the what the agent harnesses do for us. We need to talk about sandboxes.

02:20 Like when a agent harness delivers work, what is the what is the thing it delivers work into? And then of course we've got MCPs, which which everybody knows, and we'll talk a little bit more about what we're doing there. But let's start with the agent harness. So I I mean it's it would be an oversimplification to say that an agent harness is something which just takes in context and emits tool calls.

02:48 But there's also a little bit of truth to that. It's not the most complex part of it. It's a loop that takes in some context that you give it and asks you to do something on its behalf. But belaying that simplicity, we get we're all using a lot of different harnesses. And the inter the ability to why why do we pick up different harnesses all the time?

03:11 Like we love the new sort of interactive interactive possibilities that they bring us. So let's just admit right from the beginning, we're going to be using a lot of harnesses. They're going to evolve pretty quickly. We're going to be swapping in different harnesses for one another. But a harness by itself is just kind of in a little bit of a dead room.

03:30 Like uh hello, someone give me some context. Some somebody give me something to do. This is where this is of course where MCPs come in. MCPs allow us to do new things. They allow us to go outside of ourselves, pull in new context, pull in new tools, and and get actual work done. So, I think it's when you're trying to think about safety, I don't think it's an oversimplification to say that one of the things that you need to do is figure out what are the resources that you need to do your job.

04:08 Like anytime, even for us, one of the first things you'd analyze when you want to do a new task is what do I need to get this tasks done? What information do I need? What tools do I need to actually get this done? And if you knew that the sandbox that you built had the right resources, had the right tools, and only the right those resources and tools, it would make you feel safer.

04:32 It would make you feel like you limited the blast radius for what could go wrong when that agent is working. So, this is kind of where sandboxes come in. Sandbox is a place for an agent to do to do work. So, our typical sandboxes today, you're most of us are accustomed to the Yeah, we have code sandboxes. We give it We give it a bunch of code. We give it all the tools on our laptop.

04:58 Maybe like when you first started using some of these agents, the sandbox was your entire laptop. As over time, we're starting to learn how to get that a bit smaller. We're trying to make the sandbox boundaries small enough that they can actually just do the task that we want to do want them to execute. That helps us feel safer about letting the agent run for longer periods of time if this if the sandboxes is a little is is is smaller.

05:30 And you know, it's not that we're trying to sandbox the harness itself. The harness is kind of already sandboxed. Like it's a it's a pretty simple thing. We're trying to harness the tools and the context that are flowing into this harness. So, in order to kind of illustrate some of these concepts, I came up with two stories that I'm I'm just going to walk you through.

05:53 The first one is um a newsroom analogy. So, if you've got a newspaper reporter, they go out in the world, they find cool stories, find cool information, they bring it back. Maybe someone inside of that agent agency uh look it does some fact-checking on that, make sure the information is right, and then you write a story. It's totally natural that those roles are completely separate.

06:22 The actual reporter is not going to be allowed to like publish to the website or or or write write the actual article. We we break those up. That's actually how we already build complicated systems. We split them up into little pieces. So, on this slide, what you see is I've broken down three sandboxes in blue here. So, we've got a researcher, we've got a fact-checker, and we've got a reporter.

06:47 So, that researcher sandbox, I think I think about that as the reporter. Now, when we define that sandbox, let it access the web. Let it access as much of the as much of the internet as we want because we're not going to give it the MCPs that it can like write to our our corporate site. It can't publish anything. We're just going to let it do we're just going to let it go out, go crazy, do research.

07:14 So, our sandbox is like, "Yeah, not too many tools. I'm not going to give you too much write access, but definitely do lots of research. Just write your research out to a I don't know, to some temp directory, some private thing that is local to this agent." A fact-checker is going to need some MCPs, but it's not going to need any network. So, then let's start the fact-checker up.

07:35 Let's put that in a different sandbox. Let's say no network. Yeah, and you can read the research that's been done for you, and you can fact-check it. And you can have access to our our fact database. You know, filter that out. And then finally, we get to the little blue thing on the right. And this certainly doesn't need any any access to networks. In fact, it shouldn't have any, but we do want to give it our notion MCP or our publisher MCP because it's just looking at filtered research.

08:06 So, you know, look at this whole slide, draw a box around this entire slide, and that effective agent has a bunch of MCPs. Maybe it has an MCP for publishing, it has an MCP for fact-checking. At one point in time, it has access to the entire internet. But we break it up into logical sandboxes, so at each individual point, we have a don't have a dangerous concept a dangerous combination of both being able to read in crazy unfiltered context and do things with our tools.

08:40 So, let's look at example number two. And I'm just going to talk about uh building a Actually, this is how I I build my coding agent right now. I tend to send the the coding agent off to do pretty long-running tasks. And let's just say that they work for about an hour. But as the as the work is happening, I also like it if it just does commits along the way.

09:11 But the ratio of the amount of time that it's my agent is actually committing is almost nothing. In an In an hour, it might be it might need, say, my my my Git signing keys for like 2 minutes of that time. And when I'm committing, I need to see my my Git commit tree. I need to make that commit. I need some signing keys, but I don't need anything else.

09:33 So, I would like it if the actual sandbox that an agent is running a task in is modeled after the intent of what I'm doing. So, if I'm not going to make a commit, I don't want my commit signatures, my commit signing keys inside that sandbox. And I think this this you can map across a lot of different tasks. If you know the if your agents know the intent of what you're trying to do, then use that intent to model the capabilities that you give to that.

10:07 So, this is driving in in our our Docker sandbox product, this is defining a lot of how we're thinking about exposing things like an MCP gateway. So, when you build a sandbox, of course you have to put a harness in in that to do things for you. And then we're just putting a little gateway endpoint into that. And that gateway endpoint funnels all MCP traffic.

10:34 So, imagine it's something like MCP an internal an an internal URL mcp.gateway.docker.internal or something like this that every single agent harness has and all traffic, all resources that you need to pull in, all tool calls move through this one this one single gateway. So, what does this buy us? Well, we end up with a control point. So, the gateway manages which tools, resources, and prompts are given to each sandbox.

11:06 So, you're starting to see the beginnings of your ability to say, "Okay, I'm giving you a harness. I'm giving you this gateway endpoint. The gateway this the configuration of the sandbox controls what you can do inside of that sandbox with things like MCP." Another nice thing is suddenly all your harnesses are MCP agnostic. So, you don't have to go to codex and configure your MCPs one way and go to uh cloud code and configure your MCPs one way.

11:36 You just have one you almost make the harness a parameter of the sandbox. Stick a harness in, stick the MCPs you need in, give it some network rules, Bob's your uncle. So it's it feels the MCP MCP is is a good place to manage things like credentials. In fact, I would say that a maxim that you can use here is how many credentials should be in uh a sandbox.

12:02 Well, I mean the the the right answer is always zero. It just they should just never be in there. And the blast radius of an agent going doing something going going uh doing something incorrect, if there are absolutely no no credentials in there, is is reduced. So one of the things that um Docker, Entropic, Octo are all partnering on is a is a new thing called XAA, cross app application.

12:34 And I think this is a really great illustration of why sandboxes and um and managing credentials in this way is important. So today in your well you're you're working at at at your companies, you have some sort of identity management, you have SSO set up, and you've got all these resource servers that that manage that that that you work with for OAuth.

12:56 It could be Notion or it could be Atlassian or GitHub or Slack. They're you've already configured SSO for those. But it's not yet accessible to your agent. So with a cross app application scenario, your harness, your your your sandbox, and your gateway can define things like agent identity or who are you that this agent is behaving on behalf of? And by extension by exchanging identity claims with your identity provider, which is already there, you can get back a new thing, which is a new part of this spec that you

13:38 you'll now see in the latest version of MCP, called an ID jag, an authorization grant. And with that authorization grant, we're granting access to this act to the to this actor to this agent to the same authorization servers that you've been using. There's nothing new here. We're not We're leveraging all this existing investment that these corporate internet corporate networks already have, but suddenly, without any crazy consent screens, yeah, you can do that.

14:09 Yeah, you do that. Yeah, you can do that. We can centralize the administration of what an agent now does with all these existing resource servers. It's a great It's a huge simplification. So, the MCP gateways, I like to think are we're starting to talk about progressive disclosure, which is a term we learned from from skills. I like to think we're starting to use this now we're able to progressively disclose tools and resources and MCPs into these sandboxes using very, very similar principles.

14:42 And of course, it's amazing because we keep contact size down. Just because you might at some point in a workflow use 50 different MCPs, doesn't mean that each sandbox has to have all 50 of those MCPs. Give your agents room to solve problems, but let individual sandboxes represent the intent of the task. Let them represent what you're actually trying to do in this task.

15:08 So, I like this picture because suddenly we start to think about orchestrators as one of their jobs is to go, "Well, what am I doing? What's my task? And I need to put a task into a sandbox. So, what sandbox do I put it in? What capabilities does this need? And of course, it needs a harness. Like, what do you put in? Open code? Put in Claude? Put in Gemini?

15:30 Put in Codex? It's going to need some MCPs. Take a look at the task, decide what MCPs are allowed to be used in this. Um what networking rules should you apply? Does this need access to api.github.com? Does it need access to surf the web in in in random ways? Or does it just need nothing? What resources? What work trees? We get used to thinking about modeling capabilities, building the right sandbox for the task, putting the task in, and now we're starting to see that this individual loop is safer because it has less

16:08 access than the entire workflow effectively has. So, this is um this is about containers. This is about containerizing agents, which is why uh I guess this is why we're at I'm I'm at Docker. But this ties together as a kind of like role separation. What are the roles that you that your agents have? How do you map them and containerize them and and make sure that you're giving them sandboxes that express that intent.

16:37 Never have any any creds anywhere in any in any harnesses. Disclose capabilities progressively into agents as they need them. The sandboxes for you are are are what actually represent intent. And if we allow some of these things to run longer, but they're more sandboxed, we feel a little bit safer about that. That's this is where safety comes from. So, we have a demo.

17:04 Our our EVP of of engineering, Tushar Jain, is I got to talk upstairs today. I think it's at 4:00, and he's taking a lot of these ideas and just demo demo demo. How do you make move sandboxes to the cloud? How do you build orchestrators? He'll be really showing a lot of these things live. We also have a booth over here. Um it's a great opportunity to come over if you're interested in any of this stuff.

17:28 We can show you the command line that's available today. Anyone uh just brew install sbx. That Everything I've been talking about today is this tool sbx. This is This is how we build we build containers. Um so yeah, please please uh please come over and talk to us at the booth. We'd uh we'd love to hear what you're uh what you're doing with AI safety and how we might be able to help. But uh thank you very much. >> [applause] [music]