← All transcripts

Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto Transcript, AI Summary & Key Points

AI Engineer · yesterday · Science & Technology · 16:40 · EN

Watch on YouTube

Answer

Reducto deleted its first MCP server because one tool per API endpoint produced confidently wrong, unexplainable workflows; the rebuild used fewer tools that carry the structure of real work, made uncertainty visible, and added instrumentation that turned it into a self-improving system the sales team adopted voluntarily.

AI Summary

Reducto rebuilt its internal MCP server after a first version — one tool per API endpoint — broke as soon as the whole team used it. Abhi Arya frames agent-first software as an architecture question that reduces to: who pays when the agent is wrong? He lays out three approaches — auto mode, user-first, and agent experience in service of the user — and argues the third wins. The original server let the agent confidently build entire workflows nobody could explain, so he deleted it in one commit and rebuilt with fewer tools that carry the structure of real work, plus a snapshot tool for live pipeline state. Uncertainty was made visible through confidence scores, bounding boxes and evals that fail the agent for shaky assumptions, and an ask-user tool instead of guessing. Instrumentation of every MCP session created a self-improving loop, and the sales team adopted the system on their own, using it to build customer pipelines before sales calls. The closing definition of agent-first: when the agent is confidently wrong, who finds out and how fast?

Key Points

  • Reducto has processed over 3 billion documents in 3 years across customers like Harvey, Scale AI, and top five global tech companies and hedge funds.
  • Agent-first software is an architecture problem, and its structure comes down to one question: who pays for the agent being wrong?
  • Auto mode optimizes for autonomy; when the agent messes up, nobody notices because it is confident and the user pays for the errors.
  • User-first optimizes the human interface and leaves the agent capable of little — a chatbot experience that caps both the user's understanding and the agent's capability.
  • The winning bucket is agent experience in service of the user: rich, scoped capabilities that aid the user while the human can still verify the work.
  • Users, especially the sales team, were spending all their time in configuration instead of outcomes, so the fix was an MCP server letting the agent do the heavy lifting.
  • Version 1 mapped every API endpoint to an MCP tool in one massive file, built as an MVP with Claude Code — the curated demo worked, but it broke immediately when given to the whole team.
  • The agent built entire workflows end-to-end confidently, told users everything was right, and neither user nor agent could explain what the workflow actually did.

AI in practice

Used for

Agents

  • Build end-to-end document pipelines (parse, classify, split, extract) from a half-baked customer description. 2 held 05:26

Tools & resources

4 items

CNo. 0421
AIAINotes.us AI product

Circleback

circleback.ai

Circleback is an AI-powered meeting assistant that generates meeting notes, action items, automations, and searchable conversation records. It integrates with Zoom, Google Meet, Microsoft Teams, Slack huddles, and can capture in-person conversations.

Mentioned in
3 videos
Kind
AI
CNo. 0021
AIAINotes.us AI product

Claude Code

Open source · anthropics/claude-code

Claude Code is Anthropic's agentic coding tool for the terminal, IDEs, and GitHub. It uses natural-language commands to understand a codebase, create and read files, execute commands, run tests, explain code, manage Git workflows, and handle routine development tasks. It can also load persistent project context, run custom slash commands, use plugins with custom commands and agents, and operate with configurable autonomy while leaving actions such as final pull-request merging to a human. The official repository documents installation for macOS, Linux, and Windows, and identifies npm installation as deprecated.

TypeScript
Stars
★ 149,337
Forks
25,438
RNo. 5419
AIAINotes.us AI product

Reducto

reducto.ai

Reducto is an agentic document platform for AI teams that turns unstructured documents into data for tools and AI agents. It parses complex layouts, tables, formatting, and bounding-box citations; extracts schema-defined fields; splits multi-document files and long forms into usable units; and edits detected blanks, tables, and checkboxes without predefined templates. Its r-1 parsing model is presented as a single-model alternative to multi-tool document pipelines, while outputs can be reviewed and corrected in real time. Reducto also provides an MCP server and supports document-processing workflows involving parsing, classification, splitting, and extraction.

Mentioned in
1 video
Kind
AI
RNo. 5420
AIAINotes.us AI product

Reducto MCP Server

docs.reducto.ai/mcp-server

Reducto MCP Server connects AI agents to Reducto's document-processing APIs through the Model Context Protocol. MCP clients such as Claude Desktop, Claude Code, Codex, Cursor, VS Code, and Windsurf can use it to upload, classify, parse, extract, split, and edit documents, chaining those operations in an agent reasoning loop without custom integration code. It is available as a hosted HTTP server at mcp.reducto.ai/mcp or as a local server run with uvx. The hosted option accepts public URLs, while the local option can read local file paths and public URLs and uses Reducto credentials. The server is described in the accompanying talk as an agent-oriented interface with tools that represent document pipeline work, a snapshot of live pipeline state, confidence scores, and an ask-user mechanism intended to avoid silent guesses.

Mentioned in
1 video
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Why We Deleted Our MCP Server and Rebuilt It — Abhi Arya, Reducto — AI Engineer (16:40). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 [music] >> So, I'm Avi, and I work on product here at Reducto. And before I dive in right away, I wanted to give you guys some more context on what Reducto is. Uh we're an agentic document platform, which basically means that we help teams turn messy, unstructured documents into data that your tools and your AI agents can understand. I'll jump into the agents themselves in a second, but I'd like to talk about where we're at first really quick.

00:39 In the last 3 years, we've processed over 3 billion documents across customers like Harvey, Scale AI, and top five global tech companies and hedge funds. Running at that scale means that we've seen basically every way a document can break in production, and we've iterated our product and our models to handle basically anything that you can throw at them.

01:04 And getting AI to work on a clean document with a frontier model is not super hard anymore. We basically like you can throw something in the cloud, and you might get like 60% of the way there. But the hard part is production. Scanned documents, rotated pages, etc. And the model starts hallucinating because a large majority of enterprise data is unstructured.

01:24 And so Reducto is a layer that sits between your model and that data. We handle all the complex layouts and post-processing, and that way we end up working with and building a lot of agents ourselves while designing these pipelines. So, now that you know a little more about what we do, I'm going to dive into our main topic with something seemingly unrelated.

01:46 Um I visited my girlfriend this past weekend, and she gave me a sticker from the movie Mean Girls. And basically, it was an agentified version of it. So, it says, "Get in, loser. We're building agent-first software." And in true software engineer fashion, I had to put a sticker like that on my laptop. And she got this from work, but she asked me like, "What is agent-first software?"

02:08 cuz she doesn't work in AI. And as as I was like doing so or as I was thinking about it, I realized I have a lot of different answers. Building a really good harness, building a really good API, writing or iterating on prompts. All of these don't really um land on the agent being put first. They're more so just optimizing the model. Agent-first software is inherently an architecture problem.

02:35 And so, I think that structure, that architecture, really comes down to one question, and that is who pays for the agent being wrong. And to that, I think there are like three buckets that you can put this into. So, you have auto mode to start, where you optimize for autonomy, you write a prompt, and you let the agent do whatever it wants. And when it messes up, which it will, you won't have any idea because it's really confident about what it's doing, and you're the one that's paying for it being wrong to your

03:03 customers, your users, etc. Then you have user-first. This is where you optimize the human interface, give the user all the knobs, and then you build on an agent that can't really do much, which is basically just a chatbot experience. You've both capped the user's understanding of the product because now they can't prompt it, and the agent's capability within as well.

03:26 And lastly, the bucket that I think wins here is agent experience in service of the user. This is where the agent's capabilities are structured to give the model um rich and scoped capabilities where it can do things that the user really wants, aiding and enabling them, and even where the human can still verify the work that it's doing. This framing is what turns agent experience into less of a question of the model and more of a question of context, capability, and overall self-improvement.

03:58 And here's how I discovered these buckets and came to that last point. At Reducto, we've been building tools that basically automate end-to-end document work through pipelines. Think that we first parse a document, then we classify it, or maybe split it into sections, and then extract from it. Um and that's in use cases like invoicing, contract management, and all these things that require human input.

04:20 One thing that I noticed early on was that any of the users, especially our sales team at Reducto, was spending all their time in configuration instead of outcomes. And that's a larger problem across the board because legacy interfaces were designed for kind of building that pipeline. Meanwhile, in this kind of agentic work world, we're increasingly focused on the final deliverable and how can we get the user to that point as fast as possible.

04:49 So, the clear move to get here for us was an MCP server. Let the agent do the heavy lifting and only burden the user with thinking about their use case and writing a prompt. To get some validation for the direction, I threw together an MVP with Claude code, and every API endpoint that drove our app was given an MCP tool in one massive file. I gave it a really pointed prompt.

05:11 I put it in the Slack, and honestly, it was pretty nice. Um to some extent, it should have worked in that setting. I knew exactly what I was doing. I was building a really curated demo, and that's why the standalone demo told me nothing at all. The real test happened when I gave the MCP to everyone on the team, and immediately it broke. And here's how.

05:35 Someone would feed it a half-baked description from a customer call or maybe a Notion page, and the agent would just start going. It'd build the entire workflow end-to-end, but very confidently, it would tell the user that everything it did was right, but neither the user nor the agent could explain what the workflow actually did. And that is the root awakening of auto mode in production.

05:58 The agent never really said that it's not sure. It handled with every tool that it had, and there was nowhere for a human to really validate or think about what was going on. We built something that could communicate with our system, but we hadn't built something that could actually work with it. And you'd think that the fix here is more detail on the system prompt or more model guidance, but realistically it's architecture.

06:23 The whole file that we used to talk to the MCP, I deleted the entire thing in one commit. And what replaced it wasn't more tools. It was mostly fewer tools, but they carried the structure of actual work. So for example, instead of one create step of this workflow to uh call over and over, and just praying that the agent assembled right in the pipeline, I built tools that understood how our internal dashboard or our product works.

06:53 For example, a common use case with Reducto is to route our classify endpoint that classifies documents into an extract node. And what that does is basically extract information downstream. And instead of basically letting the model guess this or figure it out through the prompt, I turned it into a single classify to extract tool. And after each step, the agent gets contextual guidance on what actually comes next, with certain tools even chaining automatically like this one.

07:21 Rather than the agent having to guess or think about what is right versus being told what's right. Now the agent doesn't get to freelance the parts that we know are true. And since agents are stateless, we even added a snapshot tool. So this is one call that basically hands back the live state of our pipeline from the context of the entire flow to the errors that were surfaced with like really specific error codes.

07:48 And that way they can reference what's actually there instead of guessing variable names, hallucinating, etc. By moving context out of the prompt and into structure, the agent stopped guessing and started referencing. And because the tools only let pipelines take the shape that we as the engineers on the product wanted, the human could follow exactly what happened and we could explain the agent's outputs to downstream users of this MCP tool.

08:18 So, to avoid auto mode, you don't hand your agent an API, you hand it tools that carry the overall expertise such as the right patterns, the state, and the validators so the correct path also becomes the simplest path that the MCP can take. So, we can build good pipelines now. We're done, right? Not really. Thinking back to auto mode, the agent was confidently wrong and no one could see it.

08:45 Um that what we just talked about fixed that for building. The tools made the structure really legible and the human could follow what the agent did. But now the problem showed up at runtime. When the outcome itself that the agent is trying to build goes wrong, who can tell? For example, if we had a really bad schema or maybe the user didn't really specify the actual use case, how can we tell that the agent did something wrong?

09:09 And with agents it goes wrong quietly. So, instead of a seg fault or some actual error, um the pipeline runs across 100 documents and maybe it misses scan on page 30 or maybe the user didn't specify what types of documents were going into the pipeline. And basically both of these things are the same failure. The agent itself is very confidently wrong and the uncertainty that the agent has is completely invisible.

09:36 So, everything that we did to fix this basically came down to one uh change, which was making uncertainty really legible. In our system or with Reducto, uh failures carry information. So, with Extract, we output confidence scores, bounding boxes, all of this information about what went wrong, and we were able to build tools that feed that information back into the model, which actually helped optimize the extract schema over time, and get more accurate document outputs.

10:08 But, that's just the outcome uncertainty. The agent also has to surface its own uncertainty. When it hits something ambiguous, um two ways to read a field or a document type that it hasn't seen, it shouldn't guess, and it shouldn't just pause exactly where it is. And rather, it should ask, "I can read this two ways," and tell the user, "What did you mean?"

10:29 And so, basically in our evals, we started failing the agent for building any shaky assumptions, and we gave our harness uh prompt guidance to use the ask user question tools, etc., that Cloud Code and other like frontier harnesses offer, instead of guessing. So, we rewarded asking instead of guessing. And judgment doesn't live in the agent alone. This entire pipeline is still overseen by the user, and the confidence and the reasoning of the agent was then outputted as basically prompt guidance in the MCP.

11:02 So, now every time you make something new, it explains to the user what it did and why, so the user can then review it, inherently acting in service of the user. The agent basically satisfies what you wrote, and not the um outcome that you wanted. So, what happens if you make the uncertainty basically a bit more legible, and you keep the final judgment where it can actually be accountable in the hands of the human, then that one clarifying question does the work of all three, giving the model context, capability, and

11:38 also adding a human in the loop to actually build that judgment and verify things as well. So, going back to kind of all of these failure modes overall, we had initially an auto mode that nothing nobody could inspect. We had wrong answers that basically arrived in silence. And then an agent that passed the test by reward hacking or hiding what it did.

11:59 And underneath, they're basically one shape where the agent had more autonomy and more access than the human could verify. And that intelligence lives in the construction overall. So, building the pipeline, generating the schema, and pulling in the right context is where the agent is more most powerful and where a mistake is the cheapest because the human is right there.

12:25 By building for that interface, we were able to also have user satisfaction go up with the MCP because they could really see what was going on. And by using Reducto for the document layer, we can guarantee accurate parses and extractions for parts of the pipeline that we can't trust agents with. Where the agent operates, it earns trust by being verifiable, correct, and building trust in service of the user.

12:53 And the whole conversation we had about legibility overall is what inherently pays off. Agents are just intelligence in a box and the harness that you surround them with is what provides them with the ability to deliver these outcomes downstream. And so, once we reach this point, we focus on instrumentation where every MCP session, the prompts that were sent from the model were logged, and we even asked like our sales team and other people on the team to basically export their cloud code sessions to understand how how

13:22 was behaving. Where humans overrode the agent, where confidence of the model ran low, and where it asked questions. And because we added all this instrumentation, um it flowed back into the long-term guidance that the MCP had overall. So, the system got better at building pipelines by building more pipelines without us rewriting by hand. This then became a self-improving loop per org or per user that was building these pipelines.

13:51 The foundation became overall the same decision where we kept the model very guided and very inspectable, and the information around it became everything else that actually guided the model. And so, one thing that kind of told me that we were doing something right with these improvements is I never like explicitly told the sales team to use the agent to build pipelines, but rather I just exposed it in our Cloud Enterprise as product tools.

14:19 And the sales team and various other people on the team started basically wiring it into their workflows. So, before they had a customer call, they would use CircleBack or Slack and basically learn about the customer and then use the MCP to basically ask them questions about the customer and then build a pipeline for that customer downstream. So, when an agent's experience was built really in service of the users that were using it, we weren't pushing people directly toward it and rather they reached for it and

14:49 actually started using it as a company-wide resource. And that's where they also started to trust the output enough to put this in front of high-value customers because the demo didn't just work, but the human actually said, "Hey, I built this. Come check it out." And it actually paid off in these sales calls. And I hope the Slack reactions can also help gauge that a little bit.

15:13 So, overall, a few months ago I would have told you that agent first is means that you have the best model, the best inference speed or intelligence of the model. And with models getting better from closed source to open source, from expensive to cheap, I think it now means almost the opposite of what it sounds. To some extent, the agent doesn't go first.

15:34 And to build for agent experience means you have to consider the agent being in the loop with the people who have to answer for the work downstream. So, it's not just what does the agent connect to? Everyone's agent connects to everything nowadays, and Claude can make a new connector for you immediately. But, regard Beyond that, it's when your agent is really confidently wrong, who finds out and how fast did they find out?

16:02 If you build for that, then you can build for something that people will actually trust with real work. And yeah, that's about it. Thank you, guys. And also, we're at Reducto's at a booth P8, and we are hiring for a lot of positions. So, please come check it out at reductoai.notcurious. Thank you. >> [music]