← All transcripts

An ex-OpenAI researcher just deleted language from the LLM... Transcript, AI Summary & Key Points

Fireship · 9 days ago · Science & Technology · 05:27 · EN

Watch on YouTube

AI Summary

Jev is a System 1 AI model that removes natural-language generation and returns typed outputs: a choice, a score, or a yes/no value. Its stated advantages are 200 times faster, 400 times cheaper, free output tokens, and zero hallucinations, although type-safe output does not guarantee correctness and identical inputs can produce different results. Typesafe AI raised $40 million. Jev provides calibrated confidence values through reinforcement learning for calibrated decisions, while its architecture remains undisclosed. OpenJV reproduces Jev's interface with a frozen Quen 4B model in one forward pass, requiring no new training and running on a 3090.

Key Points

  • Jev cannot talk, write code, or write college essays because it removes language generation from the model.
  • Jev returns one of three typed shapes: a choice, a score, or a yes/no value; schema matching is guaranteed and a type error is mathematically impossible.
  • Jev is designed for fast, low-cost decisions such as checking whether content matches a category or moderating an account.
  • Jev is described as 200 times faster and 400 times cheaper, with free output tokens and zero hallucinations; another comparison puts it at 440 times cheaper than a big-brand model.
  • Jev can support real-time applications, including NPC behavior in video games and a real-time AI calculator.
  • Type safety does not guarantee a correct answer, and Jev is not deterministic: the same question and context can produce different results.
  • Jev returns a calibrated confidence value through RLCD, or reinforcement learning for calibrated decisions; a 60% confidence value means the result is right 60% of the time.
  • The Jev architecture remains undisclosed, with a paper possibly coming later.

AI in practice

Used for

Tools & resources

2 items

Links mentioned

  • Mux https://mux.com/fireship
🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of An ex-OpenAI researcher just deleted language from the LLM... — Fireship (05:27). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by Fireship. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

Last week, everything about AI changed forever. I realize everything about AI changes forever almost every week. But this time, AI changed forever more than usual because one of the OpenAI researchers behind the instruction following work that eventually became Chat GPT, who just released the next big frontier model after 2 years in stealth development, a model that can't talk, a model that can't write code, a model that can't write your college essays, and a model that will never tell you you're absolutely right.

Its name is Jev. My name is Jeff. >> And this is a huge deal because large language models have one fatal flaw that they won't shut the hell up. You give Fable or Astra a simple instruction like return true or false and it'll discover a third option after thinking for 4,000 tokens and then charge your credit card 11 cents. Jeb fixed this problem with a radical solution.

It deleted language from the large language model. And the result is a new type of classifier that's 200 times faster, 400 times cheaper with free output tokens and zero hallucinations. This sounds too good to be true. So, in today's video, we'll take a look at Jev's code. It's Trust Me Bro Benchmarks and the dude who says he built an open- source Jev over a year ago.

It is September 21st, 2026, and you're watching the code report. The big AI duopoly is literally shaking right now because Jev is a cheaper, faster way to solve basically any AI problem that requires a quick gut instinct decision. >> It's afraid. But the first thing you need to know is that Jev was created by an exopai researcher, Diego Almeida, and his company, Typesafe AI, which just raised $40 million.

But the company name is the first clue to what Jev really is. Like a regular large language model, you send it a question and some context, like a bunch of unstructured text. However, it differs because it behaves more like a type- safe programming language like TypeScript. The question you send to the model is a strongly typed question that must return a specific shape.

One of three shapes actually, a choice, a score, and a new, which is basically just a yes or no. It schema matching is guaranteed, and a type error would be mathematically impossible to produce. They call Jev a system one model, which is a name that comes from Daniel Conorman's thinking fast and slow. A system one model is fast and goes from gut instinct, while a system 2 model is slow and deliberate, like these old antique reasoning models like GPT6 and Claude Fable that burn 40,000 tokens to name a variable.

But the difference is huge for app developers like myself who want to integrate fast cheap AI into their applications. Like on horse tinder, we recently had an issue of some donkeys trying to use the app, which is strictly forbidden in the terms of service. Thanks to Jeff, we implemented an AI moderation step that will instaban any account that is not a horse, which is accomplished by returning a new response to is this a horse.

Not only is it extremely fast if we are to believe these TMBBBs, but more importantly, it's off the charts cheap, like 440 times cheaper than one of the big brand models. In fact, it's so fast and cheap that you can even use it for real-time applications. Like developers are already using it to implement NPC behavior in video games. And this guy even used it to build the world's first real-time AI calculator.

But just because the output is type- safe, that doesn't mean it's always correct. And it's not even deterministic. Like you could send it the exact same question in the exact same context and get different results just like any regular large language model. But to get an idea of the response quality, it returns something called the calibrated confidence number.

Chap models are trained to please human raiders, and humans love confidence, which is how we got models that are wrong with the confidence of Kanye. Jeb gained its confidence through a technique called RLCD or reinforcement learning for calibrated decisions. This means every response provides a confidence value like say 60% which means 60% of the time it's right every time.

But the big question is how does Jev actually work? Well, nobody knows for sure because the CEO says the architecture is staying close to the chest with a paper possibly coming in the future maybe. But Jeb also has some doubters as some people say it's no different than zeroot classifiers of the past. But the company gives no credit to the original pioneers of this technique like Jiny Yang who were building zeroot classifiers over a decade ago.

In addition, this guy claims his paper he released a year ago is the exact same thing as Jev. And another developer already built OpenJV, which reproduces the entire interface by reading option probabilities off a frozen Quen 4B model in a single forward pass. It requires no new training and can run on a 3090. And there's even a web GPU demo you can run in your browser right now.

It's an awesome time to be a developer, which is why you need to check out MX, the sponsor of today's video. Their highly customizable API is by far the easiest way to add video features to your application without getting jump scared by FFmpeg. We've used it for years to handle all the hosting and streaming for our courses. But it does a lot more than just infrastructure.

When you upload a video to MX, you automatically get transcripts, storyboards, thumbnails, and clips along with structured data about what's actually in the video. that powers MX Robots, which is their AI hosted workflows that can translate your audio into other languages, moderate content, and lots more without you needing to host a model or maintain a pipeline.

You can automate all this with directives where you define a workflow once, and it runs on every new upload. And you only pay for the jobs that actually run. Perplexity, Patreon, and many other prestigious companies all trust MX. And their free plan includes 10 videos and 100,000 delivery minutes per month with no credit card required. And you can get an extra $50 credit at the link below. This has been the Code Report. Thanks for watching and I will see you in the next one.