Jev is a System 1 AI model that makes fast decisions by selecting from predefined options and returning a probability for each option instead of generating text.
Searchable transcript of What Is Jev? The AI Model That Doesn't Generate Text — IBM Technology (15:03). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by IBM Technology. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 All right, I want to talk about Jev. Jev's a new AI model from a startup called TypeSafe. And it's a bit unusual because it's an AI model that doesn't actually write any text. It actually answers questions by picking from a list and then giving a probability for each one of those answers. Now type safe causes a new type of AI model. It's a system one model, which is a name borrowed from Daniel Kahneman in his book, one of my Favorites, Thinking Fast and Slow.
00:35 Kahneman describes two modes of thinking, human thinking, that is, and the first of those modes is called system one thinking. And system one thinking is fast and automatic. So what's two times two? You instantly know the answer is four without having to really work it out in your head. That system one thinking. The other type of thinking is classified as system two thinking.
01:07 And that's slow and deliberate. So if I ask you what's 17 times 24? Well, you're probably going to need to work through that one step by step to figure it out. Or just look at the answer on your teleprompter, which is 408. There we go. Got it. Now, most of the AI models we use are built to write stuff, and a chatbot produces its answers one token at a time in a forward pass.
01:43 Now a reasoning model goes a bit further than that, and it actually writes out its chain of thought as we're going through this. So you can actually kind of see what it was thinking, which is basically as close to a system two model as AI gets. But, but, a lot of decisions inside software are not really system to they're just kind of quick judgment calls.
02:06 So let me give you an example of that. Let's say that we have got a email from customer support. And that email says that I was charged twice this month. Please fix it. All right. So how do we respond to this email? Well, the software that's handling this email probably has to look at three types of question. One of the questions is is this a refund request yes or no?
02:43 Another question might be which team needs to process this email, maybe the billing team. Or maybe it's the technical support team, or maybe it's the sales team. And then also how urgent is this request? How quickly do we need to respond? Now a person just skimming an email like that, they would instantly know that refund. Yeah, it's yes, this is a refund request.
03:10 That's kind of a system. One bit of work. And it's that kind of work that Jev is built for. So a common way to automate a support email problem like this today, using an LLM would be to send this email to the limb and ask for the answers back. The answers to these things in a structured formats. Maybe we would go through this large language model outputs and tokens.
03:37 And these tokens would represent a format maybe in JSON. Now that mostly works, but it doesn't come back with a reliable sense of how sure the model is. How confident is a large language model that it got the answer to these questions right? Well, we could ask the model, but whatever it responds with wouldn't necessarily actually match the real probabilities.
04:00 The model used to come up with its answer. But Jev well, give does come up with probabilities. So if we send these questions to Jev, we're going to get a very different looking response. So it takes in two things. It takes in what's called state which is the data the decision is about. So that might include our support email. That might include some recent customer charges as well.
04:31 All of that is input into the Jev model. And it also takes in each of these questions as well into the model as input. And are three questions are actually three different types of questions. So refund that is a a yes or no question. Which type save calls a null a team. This team question that really comes down to a choice as to which team is this. So picking one option from a list.
05:05 And then that really comes down to a score on a scale from low to critical. Now all three of these questions go in in a single request. And whereas an would write it Json answer out one token at a time, Jev doesn't generate any text. So all three answers, they all come back at once. And that's a big part of why this is so much faster and cheaper than an LLM on these types of questions.
05:39 Now we provide questions as input. What is coming back? I said Jev doesn't give you text. What it gives you coming back are probabilities as the output. So for refund it's going to be a single number. Let's say it's going to come back with 0.9, meaning there is a 90% chance that yes, this is a refund request. For the team, it's a probability for each one of the options that were part of this choice.
06:11 So for example, the billing choice which is one of the three. May come back with a probability score of 0.85. And by the way, the answer to choice can only be the options that we give it. So if a team is always going to be billing or technical or sales and then urgency, well that's going to come back. If we kind of put this on a scale pretty high to critical here.
06:37 So this is quite important. Now the model can still pick the wrong options or score any of these questions incorrectly of course, but the probability tells us how likely that is. So how do you train a model to give accurate probabilities like this? Well, let me go back to training traditional LLMs for a moment. Most language models they start with a stage known as pre-training.
07:08 So you pre-train a model first. That means the model reads a huge amount of text and it learns to predict the next token. Now that creates a model that knows a lot of stuff, but it isn't much used to talk to yet. So if we want to use this as a chatbot, then we need to go to a second stage, which is called post-training. Now, a lot of post-training uses a technique called reinforcement learning.
07:41 The model produces an answer. Something scores that answer, and the model kind of gets nudged towards the answers that score well. So the question is what's doing the scoring? Well, one way to do that is with a technique called RLHF reinforcement learning from human feedback. So people are doing the scoring. Their human raters look at two responses from the model and they pick the one they like best.
08:11 And those choices train a reward model. And then that chatbot is going to be trained to give answers that the reward model scores highly. People tend to like answers that sound confident. So an unfortunate side effect of RLHF is that a model can learn to sound like really sure of itself, even when it's wrong. But that's just one form of reinforcement learning.
08:34 Another form is RLVR. That's reinforcement learning with verifiable rewards. And the scoring here is automatic. Did the model get the right answer to the math problem. Does the code it wrote pass its unit tests those sort of things. And that's a big reason that reasoning models today have got so good at math and coding. But the reward here only checks whether the final answer was right.
09:06 It doesn't reward the model for knowing how sure it should be. And those models spend a while working through their chain of thought before they come up with an answer, and that can make them a little bit slow and expensive to run. Now Jev is trained with a different form of reinforcement learning. It's called RLCD reinforcement learning for calibrated decisions.
09:31 Now type safe haven't published much about Jev's architecture, so I stick to what they have said. Well, we know that Jev gives a probability for each option, and it's rewarded when those probabilities turn out to be right. So that's what the C is all about. It's saying that the model is calibrated. And a model that's calibrated kind of basically looks like this.
09:59 So along the bottom axis here if we map probability and then along the side axis here we can map whether the model is correct. So a well calibrated model is going to have kind of a diagonal line much like this. So when we say that there is a 80% probability that 80% probability is going to be correct about 80% of the time, that's calibration. So back to the the refund question.
10:40 That was the first of the three questions we posed to Jev about that support email. So if we have probabilities we can use them to inform code using a threshold like this. So we've got a probability threshold here. One is certainty zero is definitely not kind of thing. So if a probability score is above let's say point nine. So maybe our score comes back here somewhere.
11:14 That means we can treat this support email as a legit refund request. And we can send that straight to the refund queue for processing. If, however, the score comes back anywhere kind of around between here and let's say point one, well, there's some uncertainty here. So the code that received a number like this would know to actually get a human involved.
11:48 So now we need to bring in a person to actually check that because we're not sure. And if a score comes in below 0.1. Well I think we can say that's not a refund request. So keep calm and carry on. We don't need to do anything with that one because the numbers are calibrated. The threshold also tells us roughly how often the automatic path will be wrong.
12:16 So the more a mistake is going to cost, the higher we would set the threshold. Now this works for more than support emails. Of course, Jev can act as a guardrail that can be used around all sorts of things, like around a chatbot that's checking messages, going in and out for things like, well, jailbreak attempts. But there are limits. Jev only takes text in as of today.
12:45 It's not very good at math or even counting, so you're going to want to leave that stuff to something else. And like any AI model, it can be tricked by instructions hidden in the data is reading. So is it time to throw away large language models and replace them completely with system one models? No, hardly. But the two can work together. So if you think about a workflow, maybe we have Jev in the workflow that can make quick calls like sorting that customer support email to see whether it's a refund or not.
13:19 And then in the workflow, we could go to a large language model to do the slower work, like replying to the customer. And then we could go back to use give to classify any response that came back from the customer, which is pretty much how Daniel Kahneman said, we work to most of what we do runs on those fast automatic system one processes. So that's these two stages here, and then the slow and careful system two, which is the large language model in this case that only kicks in when something needs real thought.
14:02 Oh well, I wanted to to share a thing about the name of Jev as well. Jev is named after William Stanley Jevons, who was an economist who pointed out in 1865 that as steam engines got more efficient, Britain actually used more coal. That's known as Jevon's paradox. So perhaps the same thing will happen with system one models. When a judgment call is super fast and super cheap.
14:31 Like with Jev, well, it can go into places where an LLM would be too slow or too expensive today, like in every row of a database or an every line of a log file. Who knows? All right, that's Jev. Where would a system one model fit into the software that you work on? Let me know in the comments.