AI agents use APIs to retrieve specialized, real-time data, let backend services process complex measurements, and give the resulting structured data to LLMs for reasoning and human-readable answers.
Searchable transcript of How AI Agents, LLMs & APIs Use Real-Time Data at the US Open — IBM Technology (09:46). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 At the US Open this year, there's an AI-infused system that uses cameras to watch each tennis serve, and it gives it a score. Now, to do that, we need what's called an AI model that also can be wrapped within agents here. Now, this AI model, it would also connect down into what we call an API, which is one of the most important elements that we'll talk about in a minute.
00:24 But I want to consider, why do we even need this? Why use an API at all? How far can an AI model get on a tennis serve judging job all by itself? Well, let's say I prompt the AI model with, does Jessica Pagula have a good serve? Now the model can take a good stab at answering this kind of a question, but a large language model, it's really read pretty much everything ever written about tennis in what we call its training data.
00:52 And as we feed this training data in together, then we can build this foundation model. That can connect to this API. And it could tell me things like how Pagula rebuilt her serve, let's say around the spin of the ball and the placement of theball on the court. But all this data comes from even before the model's training run has ended, meaning that it really can't accurately answer how well a given player is serving in the given match that they're playing in at this very moment.
01:22 So if we could evaluate that, one way to do that would be to build and retrieve a player's current match stats, which we have here. And these stats could be pasted into a context window and sent into the AI model itself. So something around stats could be that Pagula's last serve was hit at let's say 97 miles per hour and was an ace. So now we've given the model a grounding and something that's happened in the current match as it prepares its responses that uses all this data that we have so far.
01:59 But stats like serve speed, aces, they're really outcome-based stats. They describe what the serve did, but not really a lot about how the player moved to actually produce it. So a serve maybe with an ugly mechanics can still be an ace, and trust me, I know firsthand about ugly aces here. So to get a better response, why not show the AI model a video clip?
02:24 And this video clip can be fed in along with all the other data and it kind of works. Now the model will report that the toss, it looked a bit high and the knees could be bent, maybe a little too much, but again, it's really limited to what a non-specialized model can pick out and some of the frames within this video. So without these APIs, the model's general knowledge, it really becomes very stale here.
02:51 The pasted stats within the context window only describe outcomes. And watching just some of this video gives a broad impression, but not a really precise measurement that we could use. So what the AI model is missing is a connection to a system that measures all of this information and context at the same time. And we can call that an API. So let's talk about what this API exposes.
03:16 So we have serve quality, which was introduced and the 2026 U.S. Open across all 254 singles matches. Now at the event, courtside cameras track the ball and the racket and 21 joints on the player's body as they serve. Now over that tournament, that adds up to one billion, yes folks, that's one billion data points, which is, well, quite a lot of information.
03:43 So no person is going to read all of that and actually neither is a language model. Instead, specialized services, they process the stream into two different scores. And the first one is gonna be efficiency. Now this score is the biomechanics. It's about how the knees are flexed, how the hips and shoulders are separated, and other mechanical tennisy things like that.
04:05 And then let's go to the second one, which would be effectiveness. This score is The Outcome, which are the stats that we pasted in earlier, such as we'll look at speed, placement. Whether the ball had found a zone that's really hard to return. Now, if we put both of these different elements together, we then in turn would have the quality score here.
04:30 Now this was developed using what we call IBM Bob to identify and weigh the metrics that are most impactful to the serve, and was grounded in the biomechanics and the kinetic chain research that we did to figure out which movements contributed the most to the effect of serve. Now this is all pretty clever stuff, but why does all this scoring need to happen within this API or a specialized service?
04:55 Couldn't we just. The data straight to the AI model and let it work out the scores itself? Well, if we think about it, we have 21 different track points, that each have three coordinates a piece, and we sample this at 50 times per second. Now, if I do the math here, this is over 3,000 different numbers for every second of play, and millions of position values over that duration of the match.
05:20 That's big context window. And it holds a million or so tokens or even more. So the raw data for a single match, it might fit, but even if it did, a large language model is really this next token predictor. Now, bulk number crunching, like driving the joint angles from a stream of coordinates, it's not really its strong suit. So the backend services behind all of this and this API is essentially, it's taking what we call the camera spot, and then we go down into joints, which then translates into the biomechanic
05:58 aspects of the data, and eventually we get what we call scores. Now, by the time the data reaches our large language model, it's really just a few kilobytes of structured data, which is easily small enough to fit into the context window and meaningful enough to reason about. And that's where this LLM comes into play with this data that we have here.
06:29 And something like a question I could ask and prompt the model is, Pagula's serve quality, it climbed this set more because it helped to give a more consistent toss of the ball. Now what happens is the model ends up doing the part that it's actually really good at, which is turning all of these different measurements into a human readable sentence that really addresses a user's prompt.
06:48 So that's one very specific pipeline. The API call happens in this US Open example when a fan asked a tournament app a question. Now, that question could be how Pagula is serving today. So an agent is then in turn used here to answer that type of question. And what happens in the next step is the model is handed a list of tools. Now these tools, they have a definition within its context.
07:23 So each tool, it has a name, a description, and the parameters of which it takes, and the ServeQuality API will then appear as one of these tools. So when the model decides that it needs data, it outputs a structured request that comes from what we call the harness here. And this harness would then eventually make a call down into the API, which then could be this quality serve element.
07:55 And then once it makes the call, the API then can return back the data into its context so that the model can then reason about all of the information that it has. Now, this can run as a continuous loop. So the model can reason about what it's missing. Maybe request another tool, then read the results and then decide, do I have enough information to answer yet?
08:21 If not, it'll then go around the agent loop yet again. But maybe it'll call a different API this time and it could even pick a different tool. Now, when it does have enough, it writes that answer within that loop. So a model with the tools and the goal is pretty much the working definition of this agent here. So I've talked a lot about tennis. But the same workflow of combining the AI agents and the API access, it shows up all over the place.
08:50 It's this pattern, like how an agent investigating a production outage it might pull from monitoring APIs as well as log APIs. Now specialized system, it does all of the heavy lifting and the processing. And the API is where the handoff to the model really happens. So why bother with an API at all? Well, It's all about dividing up the labor. So the specialized services, it measures and it computes the information that a large language model would struggle to really process on its own.
09:23 And the model it reasons over the results and turns them into something that a person can really act on, whether that person is a tennis fan or they could be an on-call engineer.