← All transcripts

Your Agents Are in Solitary Confinement: Why MCP & A2A Aren't Enough — Vlad Luzin, Band Transcript, AI Summary & Key Points

AI Engineer · yesterday · Science & Technology · 17:34 · EN

Watch on YouTube

Answer

MCP and A2A are not enough because they leave stateful communication, bidirectional interaction, message ordering, retries, persistence, discovery, runtime identity mapping, routing and governance for developers to build.

AI Summary

AI agents need an interaction layer that supports autonomous discovery, real-time communication, persistence, identity, routing, observability and governance across frameworks, languages and deployment environments. Manually routing messages between Claude Code and Codex sessions creates a multi-agent problem that already exists today. Loop engineering reduces manual copying but still leaves developers to build coordination logic. MCP provides stateless agent-as-tool calls, while A2A uses one-way client-server interactions and leaves discovery, queues, persistence and timeout management unresolved. Band connects agents through conversational abstractions such as rooms, channels and participants, and Gem provides onboarding, routing, context-overload reduction, cost management, attribution and human-in-the-loop visibility.

Key Points

  • AI-to-AI communication will span businesses, consumers and businesses, with autonomous agents written in different frameworks and languages and deployed in different environments.
  • Agents can create conversational spaces, discover other agents, add participants, exchange messages, solve tasks and report results to humans.
  • Adversarial agents — developers often run separate Claude Code sessions for implementation and review, then manually route information between stateful agents that cannot communicate directly.
  • Loop engineering — Python and TypeScript libraries can make agents prompt one another, but developers still have to create the coordination loops and abstraction layers.
  • MCP — calling an agent as a tool is stateless, making sticky stateful sessions difficult.
  • A2A — client-server communication requires agents that need bidirectional interaction to act as both clients and servers; chained calls also create REST API timeout-management problems.
  • Messaging platforms — connecting an agent requires 5 steps for Telegram, 7 steps for Discord, 8 steps for Slack and 11 steps for WhatsApp, and the result is an agent that can talk to a person rather than to other agents.
  • Multi-agent coordination is already a current problem because developers copy and paste information between sessions and connect agents as tools.

AI in practice

Used for

Agents

  • Codex agent — Discover and collaborate with another user's personal assistant by requesting a connection, inviting it into a conversation, and sending it messages. 2 held 10:23
  • Claude Code agent team — Work together as engineering manager, developer, and architect agents to review PRDs, SRS documents, and implementation. 2 held 15:08

Tools & resources

7 items

BNo. 4668
AIAINotes.us AI product

BAND

band.ai

BAND is an enterprise-grade platform for real-time collaboration between AI agents and humans across frameworks, languages, runtimes, and deployment environments. It provides a shared interaction layer with agent communication, discovery, delegation, conversations, rooms, participants, routing, ordered transport, persistence and rehydration, runtime binding, identity, observability, and governance, allowing agents built with systems including Codex, LangGraph, Claude Code, and personal assistants to discover one another and collaborate without point-to-point integrations. BAND supports shared rooms in which humans and agents collaborate while each agent retains its own tools, models, and memory. Its offerings include the BAND Platform for building multi-agent systems and BAND Desktop for controlling and observing coding agents and enabling their collaboration.

Mentioned in
2 videos
Kind
AI
CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

Mentioned in
74 videos
Kind
AI
DNo. 0521
AIAINotes.us Tool

Discord

discord.com

Discord is a communication platform developed by Discord, Inc. that provides voice, video, and text chat for communities, groups, and gamers. It is available as a web service and native apps for Windows, macOS, Linux, iOS, and Android and was first released in 2015.

Mentioned in
7 videos
Kind
Other
LNo. 0165
AIAINotes.us AI product

LangGraph

langchain.com

An agent-orchestration framework for Python that models complex, multilayered AI workflows as directed agent graphs. It manages context and tool calls, provides state checkpoints and persistence, and supports human-in-the-loop approval gates for production deployments. The framework is reported to allow existing agents to be exported/imported into IBM watsonx Orchestrate.

Mentioned in
5 videos
Kind
AI
ONo. 0214
AIAINotes.us AI product

OpenAI Codex

openai.com

OpenAI Codex is an AI coding agent from OpenAI available as a command-line tool (Codex CLI) that helps developers produce software. It can be used alongside Gemini for adversarial audits of software requirements and implementation plans, listed as a supported coding-agent or model option in several projects, and its logs can be joined with task and test evidence. The Codex CLI can also receive and answer requests from the Penako canvas.

Mentioned in
55 videos
Kind
AI
TNo. 0157
AIAINotes.us Tool

Telegram

telegram.org

Telegram is a cloud-based instant messaging service developed by Telegram Messenger LLP and founded by Pavel and Nikolai Durov. It offers text, voice and video messaging, group chats, channels and a bot API, with optional end-to-end encryption available in "secret chats."

Mentioned in
10 videos
Kind
Other
WNo. 0087
AIAINotes.us Tool

WhatsApp

whatsapp.com

WhatsApp is a cross-platform messaging and voice/video calling application owned by Meta Platforms. It provides text messaging, group chats, voice and video calls, file sharing, and end-to-end encryption for communications. The service is available on iOS, Android, Windows, macOS and via a web client.

Mentioned in
10 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Your Agents Are in Solitary Confinement: Why MCP & A2A Aren't Enough — Vlad Luzin, Band — AI Engineer (17:34). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 [music] Hi all, my name is Vlad. I'm co-ounder and the CTO at B. Now before we start, let's uh do a bit of trivia or C. Okay, if C, raise your hand. Okay, who believes both are correct? Okay, just a couple. Both are correct. One is from the Lord of the Rings. Another is from Warhammer game game workshop games. I have in my agenda four topics to cover.

00:46 Topic number one, I would like to tell you about our physics as a company, what we believe in. Then we will talk about the evolution from adversarial agents to loop engineering and beyond. Then we'll touch base on the technical challenges of tomorrow that we have to solve today. And then I would like to present the company what we do and what problems we solve.

01:10 So our thesis is future belongs to AI to AI communication within a business between businesses and between consumers and businesses. Agents will be everywhere. They will do work on our behalf and they will have to communicate between each other. They will be written in different frameworks, different languages and deployed in different environments.

01:30 And these agents will be fully autonomous and they will communicate with each other without human intervention. Now just to paint a picture before I start talking uh how this type of communication between the agents is going to look like. So what you see here is a conversational space that agents can create themselves. So they can be added. They will receive a task either from a human or from another system and they will be able to discover other agents, add them as a participants in this conversational space, have

02:02 back and forth communication, solve the task and report back to humans. This is how the future that you all read in the newspapers will look like. And when people hear this, they think either that this is too far in the future or that I'm crazy. Okay, that's it. And here is an example of a conversation I had with a CTO of10 billion company who said that before thinking about technical solutions to hypothetical problems like multi-agent coordination, I try to keep things simple and avoid the problem.

02:40 Now let's unpack. Is multi-agent coordination a hypothetical problem? Is it possible to keep things simple? and is it possible to avoid the problem? Now, let's talk a bit about the adversarial agents as a as a concept. Um, we all use this. I'm pretty sure everyone here is a developer. And as a developer, you have two cloudcoded sessions. One is doing the work, another is reviewing the work.

03:11 And you probably have not only two sessions, you probably have multiple tabs and multiple sessions working at the same time on multiple problems. You are basically acting as a router, a Cisco router or a switch between two stateful agents that do work on your behalf, but they have no ability to communicate. Hence, you need to prompt them. Let's talk about loop engineering.

03:40 You've heard the the guy before talked about Peter and Boris, right? So what they say to you? They basically say to you, stop being the router between your agents and let your agents prompt each other. Now what they're basically saying to you is instead of being the router and copy pasting stuff, start fighting with Python and TypeScript libraries and different abstraction layers that will invent how agents should prompt each other.

04:11 We also have protocols A2A, MCP, ACP and 10 other protocols cryptoreated. Simple solves all the problems, right? So this is how your mutated system looks like probably in production MCP maybe if you're super advanced it's A2A and sounds simple. You call agents as tools. Maybe you connect to agents through A2A and life is great. But MCP calling agent as tools.

04:39 It's completely stateless. If you want two agents to be stateful sticky sessions, good luck to you. A2A is client server. If I'm an agent, I can send a task to another agent. But if this agent also wants to send tasks to me, we both have to be a client and a server. Chaining multiple agent calls together involves rest API timeouts. Good luck managing that.

05:04 And obviously what no one says to you, you need cues with persistency and so on to keep track of the messages that are being sent. And of course discovery which is not even part of the A2A protocol. So you are basically doing the planning work. You're not creating a multi- aent system. You deal with planning. But we have messaging platforms. Slack Antropic released a wonderful agent.

05:31 You can talk to it in Slack. We have personal agents here and they connected to Telegram, right? Wonderful. To connect an agent to a Telegram, five steps. Discord, seven steps. Slam, eight steps. WhatsApp, 11 steps. Every step is manual that you have to do by hand and you have to read the documentation. Again, you're doing doing a lot of planning and manual work.

06:05 And this gives you only one thing and one thing only. An agent that can talk to a person. Usually, it's you. Your agent is still alone. They cannot see each other. They cannot communicate with each other. They in digital solitary confinement. So let's unpack the stuff that I heard from the CTO. Is multi-agent coordination a hypothetical problem? Clearly if you are copy pasting stuff between two sessions, this is the problem of today and it's not hypothetical.

06:39 Plugging in MCPS and A2A, this is today's problem. It's not a future problem. Is it possible to avoid the problem? Clearly it's not possible because otherwise we would have all worked with only one session and not two and we would have no need to call other agences tools and so on. But the question is, is it possible to keep things simple? And the answer is if we look at all the options that we just talked, not really.

07:12 But how hard can it be to connect two sessions, two processes on my laptop together? How hard can it be to connect my cloud to Landruff or Salesforce agent to the datab bricks agent to SAP agent to my my my codeex actually it's pretty hard think about it these agents even your two sessions are processes that have to talk to each other through a network so it's a distributed systems problem and distributed systems are hard even before you have introduced agents on top of it and multi- aent system where every agent is

07:54 remote is basically a distributed system of microservices where each microser is nondeterministic. So it is hard. What do you need to solve all of that? So it will become easy. You need to solve the transport layer order message delivery, realtime message delivery, retries and so on. You need to solve continuity microservices, your agents, their software, ports, dockers, crash and so on.

08:20 So you need to have a persistency hydration. You also need to do the runtime binding between different agentic frameworks. You have thread ids, conversation ids, execution ids and so on. And someone needs to map all of these IDs together. So your agents can actually interact. But it's also not enough. Your agents cannot communicate at an IP port level.

08:43 They cannot communicate at the URL level. They cannot even communicate at the popsup level because it's still a lot of planning that you need to do as an organization. So for this wonderful future of agents talking to each other, we need to raise abstraction of a technical stack to conversation and talk about rooms, channels, participants and figure out the deterministic routing of messages within a channel and also across different channels.

09:09 And even if you solved all of that, it's still not enough. You need to solve the governance layer, the identity audit and etc. So I would like to introduce band. This is exactly what we solved so you don't have to. We connect every agent together and we add global collaboration layer for all the agents any framework whenever they deployed inside. It's not just the communication.

09:36 We implemented every primitive that is required for your agents to talk to each other. Let's see demo. What you will see are two different users connected to the collaboration layer. Vlad with a personal assistant and Mike that has no agents. On the right in a terminal, I'm going to spin up different agents and they will be on boarded into the platform.

10:01 So let's see how fast it is to on board a new agent. We just spun up a new agent. Programmatic registration and on boarding an agent card appears. This is a Codex agent. Next we spin up a langraph agent. This langraph also appears in the platform. From this moment they know that they exist and they can talk to each other. Now we will ask Codex agent to send a connection request to my personal assistant.

10:25 Keep in mind different users different registries. There will be a connection request contact request sent to my personal assistant. It will require bilateral consent. It will arrive here in a second. And from the moment that I approve this, Codex will know and see my personal assistant will be able to invite the personal assistant into the conversation and send messages to this personal assistant.

11:02 So we asking the codex to invite Andy. And he got invited, received the message, reported back. But let's talk about problems of today. We are all developers. We use multiple sessions, probably a lot of sessions. We do routing. We don't like it. And when we go to uh have a snack, we come back and we have no idea what our agents have done. And you as a manager have no idea uh how much it cost the distribution uh of ticket and the tokens and so on and so forth and you have no idea even how long your human was involved in

11:40 in the work. I would like to introduce gem. Gem is internal product. This is how we develop the software and we've built it on top of bank. Jam is a desktop application that simplifies the onboarding of your local agents to the platform and uh solves uh problems that we just mentioned. So we solve routing, we solve context overload, we solve cost management and attribution and we allow multi- aent and multihuman collaboration together.

12:19 So what you see here is local agents and remote agents working together and we capture all the tasks generated by cloud and codeex that uh they generate for themselves when they do the work and we present this to you so you can track the work that these agents are doing and this can be your local sessions or your local session with a session with your friend because trying to understand what your agents are doing.

12:48 1 million tokens multiplied by three, that's a lot. You need a completely different way to understand what your agent team is doing. Moreover, we provide a way for agents to describe the layout of the software architecture that they are working on. This is what you see on your right side. And you can see in real time where your agents are working right now, what piece of component they're touching in real time.

13:13 And once they're done, the market is done. If there is a human in the loop involvement, you'll get pinged. Now, you don't have to use this desktop application. You still can open your uh terminal and you can work from your terminal. But because every communication goes through the network, we can monitor all that stuff and we can surface a lot of other very useful information and we can enable my agent join your agent.

13:42 We can enable our team member who works remotely to join the session together with the agents and the humans. So he can help solve us at some problems. If we have a security guy who maintains skills for the security uh agent, I don't need to copy his skills. I can just ping his agent to join this conversation and solve uh this uh for me. Now I'd like to show you the platform itself.

14:14 So since we all managers, we'll start with graphs. Uh so the moment you open application, you can see all the stats of all the traffic that happened between your local agents and also your remote agents. You can see for instance over here I have a full stack developer uh $2,000 in tokens and this is a local cloud session. You can see an architect over here $600.

14:36 This is a local codeex session and the rest are different a agents running in different environments. But how do they work together? Right? Do we have a bunch of Python code triggering the agents so they can look together? We do not because all models right now they trained on a lot of data. So they understand very very good how to communicate through messaging platforms.

15:08 So here I have real work that I've done this morning engineering manager developer and architect different instances of cloud code working together reviewing PRDS and SRS reviewing implementation. There is no need to handcode all the loops. They know how to do it natively. And for the managers, we have full statistics. Let me kill off the man the engineering manager who sent too many messages.

15:37 So you can see the attribution. You can see the [clears throat] full usage and cost. So if you ask your question if my developer is actually involved in the code that he pushes as a PR or it's all AI slope, you can see it here. Not only locally, right, but also through the remote agents. You can see attribution by developer or by teams of your agents and developers and how they work.

16:02 Everything is gets updated in real time. You can see all agents and every agent is basically a session, right? So if I click here, I can see all the sessions of my cloud and codex instances that are running and they're connected to the platform and they're connected to a global platform. So if I want I can connect any of you to any of my agents in 30 seconds.

16:28 Obviously you can set permissions. You can see all the rooms. You can see all the apps and so on. And also you can see work what work was done, what is pending, what is in progress and everything works in real time. Thank you very much. If you want to know more about the future software development that does not involve you pulling another 50 packages and if you want to enable your Salesforce and Slack and data bricks and cloud and codex to work together and collaborate come to our boog7 QR codes for the band the

17:01 infrastructure layer and for the gem the uh desktop application to allow you to be part of this future that everyone talks about but has never seen. Thank you. >> [applause]