← All transcripts

Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex Transcript, AI Summary & Key Points

AI Engineer · 2 days ago · Science & Technology · 23:22 · EN

Watch on YouTube

AI Summary

Agentic document search over company corpora needs a hybrid approach: Claude Code skipped embeddings and pre-indexing because code is small, semi-structured, full of breadcrumbs and already on disk, but company data — PDFs, images, PowerPoints, schematics — is huge, messy, permissioned and not organized like code. Models now have 1M-token context windows, but pushing binary documents through ingestion and search is expensive, and sub-agents add token-heavy, latency-heavy environments. The answer is a harness that lets the agent choose its own primitives: hybrid keyword-plus-vector retrieval as a compass direction, then listing files, metadata filtering, grep and reading to dig in and ground answers, plus pre-parsed markdown text and persisted page screenshots for multimodal documents. Production requires a three-step sync pipeline (ingest, parse, index), disk-persisted storage with a memory layer (turbopuffer in the demo), permission metadata pushed into the vector store, and attention to freshness. A live demo pre-indexed 135 Alphabet financial filings (HTM and PDF) and had an agent build a cash-flow table for 2021–2025 by searching, reviewing screenshots, grepping and reading remote files without downloading the corpus.

Key Points

  • Claude Code launched with no pre-indexing, no RAG and no vector search — grep, read and glob over the local file system — because code is small, semi-structured, leaves breadcrumbs (imports, tests) in a natural hierarchy, and is already on disk
  • Anthropic experimented with a vector database for Claude Code and dropped it for practical scaling and maintenance reasons: keeping a pre-indexed vectorized index in sync as code updates is hard
  • Company data breaks the grep approach: it mixes images, PDFs and PowerPoints, has no organized folder to grep, is not format-standardized, needs security and permissioning, and is too large to run locally
  • Scale decides the approach: 100–1,000 files can stay in local file search since the agent manages them in context; at thousands to millions of files you need pre-indexed retrieval as a rough compass rather than burning tokens reading everything
  • 1M-token context windows (up from 200k, 100k) make shoving documents through the model expensive; sub-agents that isolate context are still token-heavy and latency-heavy because each needs full context
  • The recommended harness gives the agent five tools: hybrid search, listing files, metadata filtering, grep, and reading complex documents via pre-parsed text plus page screenshots
  • Hybrid retrieval — combining keyword/BM25 with vector search, then re-ranking the top results with an LLM — beats either method alone in benchmarks
  • Complex documents (charts, scans, schematics) break traditional chunking; parsing into markdown with spatial reasoning preserved is needed, via LlamaParse or the open-sourced LiteParse

AI in practice

Used for

Agents

  • Null — Search and traverse a pre-indexed corpus of 135 Alphabet financial filings to answer a question about cash flow and build a table from it. 2 held 20:52

Tools & resources

6 items

ANo. 5267
AIAINotes.us AI product

AWS Textract

In the AINotes directory

AWS Textract is a traditional OCR pipeline used as a baseline for measuring document parsing quality, particularly table extraction and layout fidelity.

Mentioned in
1 video
Kind
AI
LNo. 5265
AIAINotes.us AI product

LiteParse

Open source · run-llama/liteparse

LiteParse is an open-source, standalone document parser from LlamaIndex focused on fast, local parsing of PDFs and other documents without proprietary LLM features or cloud dependencies. Its Rust-based pipeline converts Office files and images to PDF when needed, extracts native text with PDFium, optionally applies Tesseract or an HTTP OCR service, merges the results, and reconstructs spatial layout using text positions and bounding boxes. It can output Markdown, JSON, or layout-preserved plain text, generate page screenshots, detect documents that need OCR or heavier processing, and expose classified layout blocks such as headings, paragraphs, lists, tables, code, and figures; the JSON representation can include per-cell table coordinates and other opt-in PDF metadata. LiteParse is distributed through CLI, Rust, Python, Node.js/TypeScript, and browser WebAssembly packages, and is licensed under Apache 2.0.

Mentioned in
1 video
Kind
AI
LNo. 0164
AIAINotes.us AI product

LlamaIndex

llamaindex.ai

LlamaIndex is an open-source data framework and developer platform (originally GPT-Index) that provides connectors, indexing structures, and retrieval utilities to connect documents and external data to large language models. It is maintained by the LlamaIndex team and is used to build LLM-powered pipelines, agents, and document workflows.

Mentioned in
2 videos
Kind
AI
LNo. 5263
AIAINotes.us AI product

LlamaParse

developers.llamaindex.ai/llamaparse/

LlamaParse is LlamaIndex's document-processing platform for converting complex documents such as PDFs, scans, Office files, tables, charts and schematics into structured Markdown or JSON for indexing and retrieval by applications and AI agents. Its API also provides schema-based JSON extraction, document classification, splitting and indexing capabilities, with SDKs for Python, TypeScript, Go and Java.

Mentioned in
1 video
Kind
AI
TNo. 3086
AIAINotes.us Tool

Tesseract OCR

tesseract-ocr.github.io

Tesseract OCR is software for performing offline optical character recognition on captured document regions. Its official documentation includes a user manual and source-code documentation.

Mentioned in
2 videos
Kind
Other
TNo. 5266
AIAINotes.us AI product

turbopuffer

In the AINotes directory

turbopuffer is a vector storage service used in document-ingestion systems to persist indexed data on disk with a memory layer. In the described ingest–parse–index pipeline, it supports permissioning and metadata filtering for multi-tenant data.

Mentioned in
1 video
Kind
AI

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Grep or Embeddings? Agentic Search Over Company Documents — George He, LlamaIndex — AI Engineer (23:22). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:01 [music] We're going to get going. Nice to meet you everyone. Uh my name is George. Uh today uh we'll be going over a talk on search document search and uh how we get better results here at Llama Index. Um my name is George. I'm the head of engineering here at Llama Index. We are a series A startup. We focus quite a bit on document parsing, orchestration and uh knowledge management.

00:34 Um today the topic of the talk will center around how we've helped enterprise customers scale up search GP and generally manage context using aentic orchestration and just general infrastructure tooling. Uh the plan today is going to be going through a discussion or a debate around uh how to best orchestrate and perform document search uh how to get the best results and we'll iterate through their uh technical challenges that we've seen all the way to a live demo at the end if we do have time.

01:11 All right. So to get started um the question that I like to pose to everyone is when and why specifically do agents need vector context in some discussions or just context period. Um and the major takeaway here is over the last few years models have gotten much better at orchestration and instruction following. They're able to traverse local file systems.

01:35 They're able to work with skills MCPS in order to retrieve information in ways that are better than before. But you still have your local search and your local context problems, right? Um in order to manage large corpuses, in order to manage complex documents, easiest way is still to manage it locally. Um and so there are kind of two evolving approaches to orchestration uh between GP or what I like to call kind of local file search and file execution or embeddings or pre-indexing pre- retrieval.

02:04 Um we'll kind of talk through what the different approaches offer and when you might prefer one over the other or when it's appropriate to do a hybrid of both. Uh to get started in understanding where this problem kind of popped up in industry u when cloud code launched uh you know there was a lot of discussion around the internal mechanics of how cloud code works.

02:25 It's actually a really interesting problem to think about how you're going to traverse and manage your local file system if you're going to go through coding right as a solution or a problem. Um, and Anthropic's taken the approach of using your local file system prep read using glob management basically to manage your local file system and it's perfect for your local execution needs because code is easy to manage locally.

02:51 It's concise and it's already in text format. But you'll also notice that um this tweet by Boris uh cloud code creator kind of highlights that uh cloud code experimented with a vector database didn't end up choosing it and the reasons aren't just for performance reasons but because of practical scaling and maintenance reasons. It's hard to keep a pre-indexed vectorized index in sync especially as code keeps on updating.

03:20 And so specifically the solution that cloud code took which was no pre-indexing, no rag, no vector search, uh works really well but only because of a few characteristics of code. Um they did not care about pre-indexing information because the code in your repo is usually small. It's usually semistructured. Usually your code leaves breadcrumbs for what else is being imported, what else is being tested.

03:41 Um and it leaves a natural hierarchy for how an agent might move about your file system. Um and the largest part there is code is already on disk and local uh especially if your agent has context. So why complicate it with a pre pre- retrieval system? And the reality is when you switch over to larger company data sets uh where you have images, you have PDFs, you have powerpoints mixed in uh it's harder because there's not really a folder to grip.

04:12 Uh you can't really grip the folders across a company data set because it's not organized like code, right? You have to think about how much effort has been put into curating the code that you write, right? It must be syntactically sound. It must have proper references. And then you see your company database or you know whatever is in your data lake.

04:29 Um and you realize it's really hard to keep things in sync, right? The scale is at a completely different level. Uh you care quite a bit more about security. Uh who can access what information. Uh the format is not standardized. you might have really complex PDFs, schematics, manufacturing diagrams mixed in and the scale just makes it so that you can't run this locally anymore.

04:48 So, how do how do you keep a system in sync? Right? The industry has gone through many different solutions. Um, you can use MCP server, you can try to manage this using Anthropics CloudMD. So, you in addition to having just this blob of a data source available somewhere, you keep track and let your agent know, hey, you can traverse the data this way, right?

05:10 So MCP skills, claude, claude, uh agent files, instruction files, these are all options for keeping track of some type of hierarchy. But the major takeaway here is actually it really comes down to the scale of the data that you have and how much curation effort you're going to put in in order to maintain that retrieval system or that data management system.

05:29 That makes sense, right? People are talking about context engineering, about harness, you know, making a harness, but at the end of the day, it's all about keeping the right primitives in place. Um but in order to do that you have to understand what's the scale that you're dealing with right so if you have 100 to you know 100 or you know 100 to a thousand files uh it may be appropriate to just keep down to local file search because your agent can manage this in context.

05:54 It won't burn too many tokens just iterating through files or managing the literal file names as it traverses. Um but as you scale out right to thousands or millions of files that's not possible. You don't want to burn those tokens. uh you don't want to read the files every single time. You want to have some needle in the haststack sense of roughly where you're looking for at least if it's not the needle in the haststack.

06:14 It's a rough compass on the direction you should be going, right? Um and so we'll talk a bit about kind of what the right approach is uh between hybrid based search, semantic based search as well as uh your old school file traversal here. Uh the other reason for why you might care about this is obviously as each model increases in its memory capabilities, right?

06:35 We got 1 million token context length now. Uh before it used to be 200k, before that 100k and even smaller before that. Um the important thing to think about here is actually as the model capabilities increase what you're willing to pay as you're going through this process, right? Do you want to shove those tokens through the model's ingestion pipeline?

06:53 Do you want to uh send binary data over to your model for it to try to render in maybe its own engine and then extract data out of uh for your search? um it just becomes extremely expensive, right? In order to use your model context for ingestion and search. Um and you know there have been solutions around how you can use sub agents right have one larger agent issue subcomands so that a smaller agent can deal with the memory context the memory bloat and just give a response back.

07:21 Um but the issue there is because each sub aent needs full context because sub aents uh even though they might use cheaper or simpler models uh if you've noticed in cloud code sometimes your you know main session may issue commands to weaker model uh sub sessions in order to execute the jobs. um the environments are actually rather tokenheavy and latencyheavy, right?

07:43 So in order to speed up your process and make sure that you're not burning tokens unnecessarily reprocessing documents each time, uh you may actually think about building out a pre-indexing system. Um and so we'll kind of talk about that here. Um in kind of building out a harness and letting it agent choose what's going on, right? So uh specifically we don't really want our system or any system that can scale out to be manually deciding uh should I use a hybrid rag approach or should I just individually search through

08:19 files and the cleanest way that we found that in this era of engineering is uh managing that with a harness. So give your agent the tools needed to either search through specific files using your old school GP, your old school cat, file name filtering or metadata filtering in addition to semantic filtering. Um, and that way you can get kind of the best of both worlds going through traversal.

08:41 uh so you can get the rough direction of which files you want to search and use semantic search to down scope your files and then using your agent's built-in file traversal capabilities to dig into the meat of the details so that you can actually ground your answers instead of hallucinating them. So the five tools that we'll talk through today um in enabling kind of document search or uh complex document search uh as you might see in like insurance, finance, healthcare industries, right?

09:11 Where you just have a ton of unatategorized or semilablled data is what are the tools that we're going to use? Right? So the question first is where am I even going to look? What's the compass direction we're going to go? Um and then you have three tools that are going to be there and you know you have a lot more of these in your local file system to figure out what actually is in this subset or this subcorpus of documents that we've identified with our compass direction.

09:37 Right? So listing files, traversing files by metadata, uh gpping specific files to see if there are multiple keywords available that make a specific document more desirable. These are all things that you could run in a pre-indexing pipeline or a live pipeline. Um, and the last one is kind of, you know, what does the document actually say? And specifically for multimmodal documents, you need a way to pre-render the documents.

09:59 And this is one kind of unique approach uh to multi-page or long schematics, which is in your pre-indexing process, one of the most important parts is to chunk or break your document down into digestible portions, right? But that applies to not just the textual side, but also the rendering side. uh because one of the hardest parts of free form document rendering is you actually have to rasterize or render the document into a human compatible view.

10:23 Uh these use cases are usually where you have an actual person that would have done this uh instead of AI, right? So these schematics, these PDFs, these PowerPoints, they're not really built for pure machine consumption. Um so giving the agent a pre-built screenshot index to search through is also very useful as you're performing rag. All right. So, we'll take a look at some of the details of the tool itself.

10:48 Um, so in terms of retrieval itself, right, which is the compass direction. There are tons of approaches you can take here. There's graph rag, there's traditional rag, there's kind of your old school keyword- based BM25 search. Um, we we usually recommend a hybrid for enterprise use cases. And that's just mostly because uh keyword-based search has its place for developers.

11:09 Keyword based search also has its place for AI agents that are looking for specific pre-indexed words in a corpus, right? Um at the same time, semantic meaning and similarity is really important. Uh combining the two through a hybrid based search usually gives to much gives you much better results. Um you can see in various benchmarking data sets that you know just raw keyword or vector will get you to a baseline.

11:31 But when you put in a fusion, you know, you combine and weight uh keyword as well as vector search, you get better results uh because you aren't blinded by, you know, specifics of your search uh method. And re-ranking the final results, you know, you might take the top 100, re-rank them using an LLM or some other deeper process uh ultimately gets you much better retrieval results.

11:51 U but that's just one part of the pipeline, right? That's the document search. Uh we care about results. Um and so tools two through four we'll talk through and kind of show later which is um once you have the compass direction of what are the what is the general group of files that might matter right what's the metadata that might associate most closely with your query uh because your agent at the start might not know all the metadata or the extracted content of your system um it can then traverse through the file

12:19 hierarchy uh search for similar files inside of your corpus that weren't even returned in that initial search in order to expand out and then further narrow down its results. Um, so giving your agent the primitives of listing out the files GP and read that you have locally is really important in a shared and distributed file system, especially if it's one that you don't want to download onto your agents context every single time, which can get expensive.

12:45 And um, diving a bit deeper into the content itself, right? Um this is where uh complex documents like you know stuff with charts, stuff that's scanned in, stuff that has uh you know schematics for example kind of breaks apart for your traditional document indexing and chunking workflow because it is not a human read line of text right it's a separate representation.

13:07 Um and so thinking about how you do your pre-indexing flow and parsing flow uh so that you don't have to reread the documents every single time using your agent is also important. you don't want to burn those uh you know binary encoded tokens every single time you want to read a new document. So we usually recommend pre-indexing uh or pre-parssing your documents using uh structured parsing flow.

13:28 Um we have a product called Llama parse. We've also open sourced a tool called light parse which handles structural content extraction. So this is like if you have tables, charts or segmented uh information, you want to encode that as either markdown or just something that your agent really can understand the spatial reasoning of. uh because when you have multimmodal data uh spatial reasoning is important.

13:51 Last one there is of course uh in your pre-parssing flow you usually want to generate those screenshots as you actually render the page right. So whatever you're doing for pre-parssing you probably also want to store and persist screenshots as they're being generated from document rendering. Uh it's not always clear if something is text or is you know got metadata associated with it.

14:08 You probably want to combine the two together. Um and so parsing is an important part but it's not the only part of a pipeline. Uh we usually you know advocate for people to think about how you can distribute parsing um and think about scaling out to improve your parsing pipeline right um there are some metrics that I would advocate for people to think about you know if they're relevant to your pipeline and uh kind of how you can optimize or uh improve your pipeline for each of those dimensions.

14:37 So if you dump your data into a traditional OCR pipeline like Textract uh or you use a OCR engine like Tesseract um does it actually capture the tables the table representation? How faithful is it to act the actual content layout? Right? Because if you get those wrong your LLM on its semantic reading side later down the line is going to pay for it. Uh it might be fine at the start but if your LM doesn't understand the location and spatial reasoning it falls apart um at some point.

15:08 All right, so before we get to the demo, uh there are just a few caveats to think about with building out an actual production pipeline, right? Which is when you have actual customers that care about using it and scaling it out. There are a few things that will always be brought up. Um there's like multi-tenency, you know, can one person see another person's data?

15:23 How does permissioning metadata get pulled in? Right? if I'm using Google Drive or if I'm using SharePoint, how do I pull in and yank that permission metadata without causing the whole system to grind to a halt or without reindexing and repparssing the pipelines each time? Um, and that's related to the last one which is freshness, right? This is kind of why Anthropic, a lot of other companies have kind of given up on trying to pre-index data uh because data freshness is really hard to achieve if your datas are complex,

15:48 your data is complex, and you also have metadata that needs to be up to date. All right. So, uh in order to break it down more easily, um we like to think about the synchronization process as three steps, right? One is pure ingestion. You know, can we get all your documents into a synchronized system? Uh the second step is going to be parsing, which is how do we actually convert your document to a standardized format.

16:10 Uh for our use cases today, we'll talk through converting all documents into a markdown or a text format. But there are also other formats you can use if you only care about multimodal representations. Um and last stage is obviously uh rebungging that data and then exporting it into a pre-index state um in order to make it available for all of the tools and steps that we talked about before.

16:39 All right. So uh I'm going to skip through this uh for the sake of time. Um, one major thing that we want to highlight here is, you know, storing your data is actually quite important. Um, in order to distribute your data out, you usually want some type of layered storage system. Um, a lot of vector storage, uh, services store everything on memory. Um, usually you're playing with some trade-off of do I want it on disk and then loaded into a faster memory store or do I want everything just in the memory store.

17:05 Um, for us, you know, we've dealt with this problem at the multi-tenant level to, you know, millions, tens of millions of documents. Our recommendation here is to try to persist your data onto disk and then have a shielded memory layer in order to retrieve the data more efficiently. Uh for our use cases today, we'll be looking at turbopuffer. Um and turbopuffer as well as most vector stores will also handle uh permissioning, right?

17:29 So what's the metadata that you have? How do you pull the data in? Uh is actually quite an important part of your system success because when you have too many documents, you need to filter by metadata. Um, so we push that into the vector storage layer. You know, as the data is being ingested, let's just make sure we tag that into the vector stores.

17:47 All right. So, uh, we're going to jump through a demo. Um, and for the demo today, the the process to think about is, you know, how is the agent going to use the subcomponents in order to actually give you a response that's grounded that actually can traverse multiple documents and actually gets you gets you a result that you're pretty happy with. Um, all right.

18:07 So, we'll go through there here. Uh, let's see. All right. So, in this demo, um, we will have a pre-index set of documents um that is available from a data source that we've defined ahead of time. This is going to be 135 files uh that are derived from Alphabet's financial reports, right? So, we have all these uh htm files, they're mixed in with raw PDF files and if you actually look at them, uh we've pulled in the metadata for where this information can be found.

18:50 So if you have file metadata available along with uh you know later on you can tag this with uh the year that it came from or the specific company it came from that gives your agents the capabilities to filter down uh the data. Um and what we'll see here is uh in your hybrid retrieval flow right uh you'll have the option of if you want to be more keyword based or more semantic based in your retrieval flow.

19:15 Um let's just do a even balance for now and we're going to search for uh just any cash flow statements and we'll talk through kind of the tools that your agent would want to use along this process. Right? So um what you've seen here on the right is the set of chunks that are returned from our query as well as page screenshots that correspond to uh what is essentially uh the page that it came from.

19:39 Right? So these two together gives your agents the capability to uh filter down to a specific chunk and request context around that specific chunk in a much cleaner way than before. Um you then might want to think about how you would GP for a specific uh piece of text in the file, right? So what happens if you have a reax pattern that you're looking for uh you know now you're looking specifically for cash flow um and you can find that easily.

20:05 uh in order to find that in a specific file, all you need to do is issue a GP command and then read from that specific command kind of what that offset is and then uh kind of what the max length is that you want. So, uh this is a different file. So, I'll just give an example. You know, if we offset 500 characters and pull in the next thousand, we'll end up with this text being available, right?

20:25 So, with these two tools, this suite of tools together, your agent doesn't really need the files on a local file state anymore. um you can interact with what is effectively a remote file storage without needing to spin up or download all the data that would need to be streamed to your system. Um and that uh gives the agent a much more capable uh set of capabilities without actually downloading all the information and burning the tokens.

20:49 So we will uh go to a chat example and I'm going to pull it up here. So uh give me the cash flow from as a table from 2021 to 2025 income as a table. All right. So give me a second here. Run this again. Give it a few seconds. Uh, while that's running, I'm going to show you guys the end results. Uh, let's go to the chat. So, this is going to be in the chat flow timeline.

21:36 Um, let's see. Still thinking there. So, as the as an agent is running over the over the process of, you know, retrieving your answers from the corpus, it's going to look through uh your corpus by issuing search commands, right? And as it's issuing the search commands, you want your agent to be able to review the screenshots, review the context, and then determine what is actually relevant for its search in a much more holistic way, right?

22:00 So you get to think through the process, you get to read specific files, you get to search through files again and rescope your queries. And in the end, that gives you a much cleaner result, right? So uh in this agent response uh you can see the cash flow summary uh by you know and you can see that it's been pulled in from multiple files as part of its thinking process.

22:22 So when we do go into the traversal flow um you can usually find and ground your answers in uh in a specific file and that gives you much more confidence that your agents found the right component or the right chunk. Um, you know, that's what we're looking for at the end of the day for agent-based retrieval. All right, so we're going to jump back here.

22:45 We'll take a look at uh that session once it finishes streaming. Uh, but you know, the idea here and the takeaway is that, you know, search helps you find information and the harness gives you the results that you can trust. So, okay, if you're interested, um, feel free to, uh, visit the QR code here. Thank you. >> [music]