← All transcripts

Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI Transcript, AI Summary & Key Points

AI Engineer · 14 hours ago · Science & Technology · 18:08 · EN

Watch on YouTube

AI Summary

Kimchi is an open-source coding-agent harness built by Cast AI to make token usage effectively unlimited and inexpensive by optimizing cost per completed task rather than cost per token. The harness selects proprietary or open models based on task outcomes, and its model preferences can change as new models become more effective. After three months with 300 employees, Kimchi recorded 2.5 times savings while token usage increased by 1.5 times and coding costs decreased by 1.5 times. Ferment runs autonomous coding tasks for more than two hours, breaks them into milestones, scores the output, retries insufficient work, and deploys completed changes to staging. Teleport keeps coding sessions running in secure remote sandboxes after a laptop is closed, while Studio gives teams a shared board for reviewing and managing agent sessions.

Key Points

  • Companies have responded to rising token costs by limiting developer usage, but Kimchi aims to let developers use coding agents for as long and as extensively as needed.
  • A company in India spent $500 million on Anthropic in one month, and Laurent Gil cites an Uber CTO who used the entire annual Anthropic budget in four months as examples of enterprise token costs becoming unmanageable.
  • Cost per task is more meaningful than cost per token when comparing models because models can require different amounts of work to achieve the same quality.
  • The cited comparison gives Gemini 3 Flash a cost of $3.5 per million tokens and a cost per completed task of $75, while Minimax 2.7 costs 1.5 per token unit and $148 per task.
  • Kimchi's automated harness selects the model for each task based on the outcome, rather than relying on people to test and select every new model manually.
  • Kimchi used the harness for three months with 300 employees, approximately two-thirds of whom are developers, and recorded 2.5 times savings over the period.
  • Token usage increased by 1.5 times while coding costs decreased by 1.5 times, with the same outcome.
  • The harness changed its model selection over time: Kim 2.6 was most popular on June 3, while Minimax 3 was most popular on June 21.

AI in practice

Used for

Agents

  • Ferment — Autonomously implement a coding task over multiple hours and produce a verified artifact suitable for staging deployment. 2 held 08:10

Tools & resources

5 items

CNo. 4842
AIAINotes.us Tool

Cast AI

cast.ai

Cast AI is a Kubernetes optimization platform for cloud-native teams. It uses workload, infrastructure, cost, and SLO signals to automate actions such as pod rightsizing, node scaling, GPU and Spot optimization, workload replica and resource tuning, bin packing, and operational remediation with approval workflows. The platform connects to EKS, AKS, GKE, and on-premises clusters without requiring cluster changes. The videos also identify Cast AI as the company that created Kimchi, an open-source coding harness with tools for autonomous coding tasks, persistent agent sessions, and team collaboration.

Mentioned in
1 video
Kind
Other
KNo. 4838
AIAINotes.us AI product

Kimchi

Open source · getkimchi/kimchi

Kimchi is an open-source coding-agent platform and harness developed by Cast AI. Its terminal coding-agent CLI, built on the pi-mono coding-agent SDK, connects to Kimchi's LLM infrastructure and can operate in single-model or multi-model mode. In multi-model mode, an orchestrator classifies work into roles such as planning, building, reviewing, codebase exploration, and research, then delegates tasks to configured models or model pools using tier and task metadata; requests are tagged by phase and user-defined labels for usage and cost attribution. Kimchi includes Ferment, a persistent project-management mode that stores goals, phases, steps, decisions, memories, and append-only transition logs so work can resume across sessions. Ferment V2 provides an experimental objective controller with token budgets and independent completion evaluation. The CLI also provides Language Server Protocol tools for diagnostics, navigation, references, renaming, and type information, plus remote sessions that synchronize a working tree to a cloud sandbox and keep an agent running after the local terminal closes. Kimchi supports MCP and native Pi packages, migration of configurations from Claude Code, OpenCode, and Cursor, configurable command hooks, and installation through Homebrew, shell or PowerShell installers, or standalone binaries. The repository is licensed under Apache License 2.0.

Mentioned in
1 video
Kind
AI
KNo. 4840
AIAINotes.us AI product

Kimchi Studio

Open source · getkimchi/kimchi

Kimchi Studio is a team collaboration interface for sharing AI-agent sessions, plans, tasks, reviews, backlog items, and work in progress. The video describes it as a shared Kanban board for agent sessions and presents it alongside Kimchi Teleport, which keeps agent sessions running remotely. Kimchi's platform applies an enforcement layer to AI coding requests, with model routing, spending limits, spend attribution, and deployment options for cloud or on-premises environments.

Mentioned in
1 video
Kind
AI
KNo. 4839
AIAINotes.us AI product

Kimchi Teleport

kimchi.dev

Kimchi Teleport is a remote-sandbox feature of Kimchi for running coding-agent sessions away from a user's local computer. Sessions continue running after the user's laptop is closed and can be accessed from laptops or mobile devices. Kimchi is developed by Cast AI as a governed AI coding platform with a CLI coding agent, model-routing layer, spend controls, and open-source model support.

Mentioned in
1 video
Kind
AI
KNo. 0464
AIAINotes.us Tool

Kubernetes

Open source · kubernetes/kubernetes

Kubernetes (K8s) is an open-source container orchestration system hosted by the Cloud Native Computing Foundation. It manages containerized applications across multiple hosts, providing mechanisms for deploying, maintaining, scaling, and scheduling workloads. Applications can be configured with YAML and deployed across hybrid-cloud environments; the platform is also used for production operations involving storage, databases, networking, and security hardening.

Mentioned in
11 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI — AI Engineer (18:08). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by AI Engineer. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:13 All right, guys. Very nice to meet you all. My name is Laura. I'm the co-founder and president of Castai, who is the inventor of Kimchi, which I'm going to show you in a few seconds. Um the talk is going to be about preparatory models for edge case but use open source for the rest and you're going to see some great result we have with that with the coding agent called Kimchi.

00:34 I have the pleasure to be here with Zilinas. >> Hey guys, I am Julianas. I am leading engineering for Kimchi coding platform. >> So you have Zilinas who run the entire engineering team for Kimchi. I'm so excited to be here with you. All right, let's go. So um you you may have seen this in the news uh two two piece of information that came up in the last few weeks.

00:56 The first one is there's a company in India that spent $500 million on anthropic alone in one months which is a kind of a good example of tokens that are completely out of control. The other example is a famous tweet from Uber, the CTO of Uber, who say that in the last four months he actually used the entire budget of anthropic for the entire year. And this is how much the token um the token maxing or the token cost is a thing for the enterprises.

01:24 And you're going to see how we fix this. Companies responded this by limiting the use of token. You may have heard this many many times with our friends developers and as a developer when you limit the amount of token it feels like this. It's like hey you can use your laptop you have an hour battery and you can only charge once once once a day. That's how it feels to limit the amount of token for a developer.

01:57 We at Kimchi think this is wrong. Our jobs as manager is to do the opposite is to ensure that we want to make the tokens unlimited and inexpensive. Our job is not to prevent the developer to use coding agent. Our job is to make sure they can use it as much as they want for as long as they want in a completely unlimited fashion. This is why we started Kimchi.

02:32 At Kimchi, we have 300 developers. Oh, sorry. 300 employees. Twothird of those are developers. The Zeven has run that. And we've been using this coding agent for the last three months. I'm going to show you what we built, how we built it, and the result of it. And then Zilvenat will show you a few more things that make the coding really, really cool.

03:00 It's coming now. We built an entire token optimization based on this idea. What you see on screen is the true cost of LLM model for the same task. It was made by the famous white paper in a university. And here's what they say. On one side, you see the cost per token. On the other side, you see the cost per task for the same task with the same quality.

03:30 That is the important thing. Otherwise, you cannot compare Apple with Apple. And here's what is written here. First, they are all different, but also look at Gemini 3 Flash, our friends at Google. What it seems to be a relatively inexpensive model at $3.5 per million token. It's a blended average at the time this study was made, but actually the task to complete with Geminina 3 was 305, sorry, $75.

04:03 And look at the others. Minimax 2.7 is one of my favorite. The cost per token is 1.5 and the task the t the cost for the task is $148. It means the models are not the same. They don't cost the same, but they also not the same per task. And that's why we built our coding agent. Would it be nice that you have an automated fashion, automated engine, automated harness that will select the right model for the task at the right time based on the outcome of the task.

04:41 That is what we did with Kimy Coding. That is what Zilmanas did with his team. And here's the results. We've been using it for three months. 2.5 times in the amount of savings we have overloaded for the period. This is a graph of what would have been our cloud bill by day. The purple one on top is what would have been cloud cost over time. The blue line is what we paid.

05:17 same outcome. But look at this. The number of tokens increased by 1.5 times. Okay, we all experienced that. The cost to code decreased by 1.5 times. This is our job as managers. We provide to our team a coding agent that is essentially unlimited because of that coding agent is growing 3x less 2.5 time less than the gross of token for the task. There's another result the multimodel I showed you before.

05:55 I'm going to show you the evolution of the selection of the model over time by the harness. And here it looks like you see there is a lot of orange early June and there's a lot of yellow late June. There's a complete shift meaning the harness is selecting a very different model. It happened the 12th of June. And this is what happened the 12th of June.

06:27 The 3rd of June Kim Kim 2.6 is winning by far the most popular of the harness. 21st of June the one winning is Minimat 3. Us as human probably we can find and understand and test and be comfortable with a new model coming in. Do you think we look at this as human? Absolutely not. But an autonomous harness, an automated coding agent that is obsessed with cost is going to do just that.

07:03 That is what I mean when I say our responsibility is to think about providing our developers with the right cost per token for the same outcome. It is our responsibility to do it. Zinas is going to tell you a little bit more of how it works because whatever you're going to see it's open source. Go Zas. >> Okay. So one more thing to mention. So Laura mentioned that token maxing is no longer an option for everyone and uh as a manager at the company as most of you guys probably heard the having unlimited amount of tokens

07:44 to spend is no longer looking to be like a case and uh we had to build this hardness for ourselves. So meaning like we ended up spending too much on cloud code like half a year ago. We saw this uh cloud bill skyrocketing and we started working on our own hardness. behind the scenes it's using like pono SSDK but uh the goal of that is like to take the open source experience still contribute back to open source by providing our own harnesses open source building constructs like ferment which fits into the kimchi theme

08:14 which is meant for longunning tasks where on average we estimate the human envelope for these tasks is uh typically happening only like every two hours every three hours and etc. So the goal is uh you use Kim to harness you run a firmment process. It asks you a bunch of questions and autonomously runs a coding task for like more than two hours. The goal of that is that because it breaks it into milestones.

08:38 It knows how to implement it and at the end of the cycle you would get scoring for your output. So you would actually know what's happening. And going back to previously discussed like the amount of models used with uh Kimchi harness the scoring is there to prove you that whatever model we will pick for you you will always have like good enough scoring for the output on the task you're implementing.

09:01 The way whole Kimchi platform works the goal is like uh we are not just providing you the coding harness. We want to provide you the full software development life cycle solution. So you do a change it runs a build for you. It checks if something broke. If something broke, it will go back, do a change again, rebuild again. And that's uh how this fermentation process looks.

09:21 If everything is complete, uh there is no extra limitation needs to be done. The changes will be deployed to the staging environment. And you see what what I like with the way explain it is the harness select the model fment is constantly checking the quality of the code and going back meaning there is a a higher order model that is checking that quality and thinking that quality is not good enough please do it again this is what I was saying earlier by it looks at the outcome and it's optimizing the token based on the

09:58 quality of the outcome. That is what is super cool with FET. >> Yep. And artifact is only considered complete once like the output of the scoring like at least B as a score. And even for that like if you don't like the score you can ask agent to go back fix it up and like it will try to achieve the A as a scoring. The intention of stuffing at the staging is currently by design.

10:23 We still believe there's like some little human that needs to be in the loop before going to the production. At least at this point we all know that the ultimate f future is to avoid that. You go to production, you look at the metrics, you look at the observability data, the solo metrics you have in your environment. If you're running on Kubernetes, you want to check if the pots are not crashing, if application health checks are passing.

10:46 So all of these things need to be done before it ship it shipped to production automatically. So this is like one of the things that are coming. So the current limitation we stop at staging but uh we'll go to production and and further on that what we believe is that reading code is not enough anymore. So like uh probably most of you noticed like across this year like uh the more vibe engineering or agentic engineering we do in our coding environments the more code changes we get in our chain sets and it kind of

11:16 becomes unbearable to read especially as you keep bringing like a little bit of a less technical personnel into your pipelines. So imagine like product manager doing the coding or my dear colleague Laura starting to do some coding and suddenly he comes up with a PR which is like 2,000 lines long and you need to actually review not just the differences but the actual intent which he wanted to achieve and the specifications which he provided to the coding agent.

11:42 So yeah moving on uh one thing we built more on top of that is the thing we called kimchi teleport. So Kimchi teleport allows you to close your laptop. It spins up sandbox in the remote environment. It's a secure sandbox where you don't need to care about anything. We can do that both on our SAS platform. We can ship it to your on premise deployments if you're like a financial technology company and uh things like that.

12:08 So we will manage this agent sessions for you. We will make sure they are running. We are accessible both from your laptop, from mobile device if we are running in the background. So it goes back and ties to the story where I mentioned that the human in the loop is only appearing like every two or three hours. So teleport emphasizes that. You run your process, you teleported to the remote environment.

12:29 The y spins up in the background. You don't need to care about it. We will make sure it's the right size, it's the right security requirements and everything else. And it will keep on running until the task is done or until the human in the loop needs to be involved. Get notifications for some clarification of specs and things like that. >> Guys, I really like this idea.

12:48 Look at the this this is a screenshot of teleport. I am puzzled that sometimes innovations are always the simplest one. Like this was born out of the idea that one of us was in the plane coding and then Wi-Fi breaks. And what do you do when Wi-Fi breaks? It breaks the coding because it lives inside your thing. No, I think it was not the Wi-Fi. It was the battery that broke, not the Wi-Fi.

13:20 So you lost. So he That's it. The guy was stopped there. What teleport does it it takes all your environment. It synchronizes with a container that live inside one of the hyperscaler. In our case, it's Google and locally on your laptop. It feels exactly the same. It just doesn't run inside your laptop. It feels like it. It runs somewhere else in your environment in a container in a Kubernetes cluster inside the place at Google that will continue running even if you close your laptop.

13:55 how simple that is. And yet our favorite tool is this one. I think the the last one I I I looked at it, 62% of our engineers are only using teleport when they are coding now for that problem. They go home, it continues. They're in transit. It continues. They on holiday, they don't look at it anymore. It continues. That's so cool, huh? >> Yep. So very smart to that.

14:23 So we built another tool called Kimchi studio. If that teleport is something typically individual uses, studio is kind of like a teleport for teams and enterprises. The goal behind the studio is uh that you have as kind of like board on your environment. It again runs in the teleport sandboxes in the secure environment. The thing here is like you're supposed to share your tasks and work in progress and plans which are made by the Kimchi coding harness with open source or proprietary models with your teams.

14:54 So for example imagine you as an engineer you are like working on a task uh you started implementing asking for a plan the plan can be easily reviewed by your colleagues like by your pe engineer by your product manager or someone else who cares what's happening and this is how you kind of work with your whole team on collaborating on this agentic engineering journey instead of like saying oh I'm going to type this prompt I'm going to try to execute it to the production only later on to learn that the specification and

15:20 requirements were incorrect the output is not what was actually wanted by the task and things like that. >> And this is yet another thing we built for ourself, we find it so cool, we decided to make it a product. How do you work with your colleague? If what you have is a C is a is a terminal, a CLI, how do you share that with the others? Well, first he has to live somewhere else.

15:48 So teleport is the base of what we call studio and studio is a very nice way to visualize all the sessions that are running including those that are asking for a review and then someone in the teams can say I take it and I answer the uh I I answer the question that is for review. Then you answer it. It goes back to in progress until the next one goes into review.

16:14 And you can have many many of those in a team in a pizza team, a team of five to 10 guys. Think about it. Yet yet a very simple innovation that is so essential in how our engineers are coding. They can now share in a very simple way, right? It's a it's a browser. It's graphical. They can share and they can take over the task according to the comban style of what is in review, what is in progress.

16:48 They can start new ones with backlog. And you know what? If they close the laptop, >> it's continues. >> This is so cool. >> Yep. >> This is really cool. Thank you. >> Yeah, we forgot to show. >> Oh, this is this is the one that we took over, right? That's the one, >> guys. You can see this at our booth if you want to uh to see more of it. We are in U4.

17:13 It's exactly this way across a few streets. But have a look at it. It's really good. It is open source. Most of it at least the harness is entirely open source. Uh studio and teleport. You need a Google account. It's a little bit more com uh more complicated, but it's a a five minute install. And then you get your teleport setup as many sessions as you want inside your Google account and then you can share with share it with your colleagues with Kimchi Studio.

17:40 So come see us. We'll be delighted to speak to you. Thank you guys. >> Thank you guys.