AutoGluon is an open-source Python machine-learning library for automated modeling of tabular data, time series, and other data types. It trains end-to-end predictive models by selecting and combining suitable algorithms, including classic machine-learning models, ensembles, and foundation models; its tabular API can fit a predictor from a training file and generate test-set predictions with a few lines of code. The library supports Python 3.10–3.13 on Linux, macOS, and Windows, with optional GPU support, and is licensed under Apache 2.0.
mitra-finetune is a Python package for fine-tuning and running inference with second-generation Mitra v2 tabular foundation-model checkpoints. It supports classification and regression through the `MitraFinetune` API, including `fit`, `predict`, `predict_proba`, and distributional regression outputs. Each fit performs a 50-step full fine-tuning run through AutoGluon's eight-fold bagging wrapper. The package uses capped, optionally class-balanced in-context support for large tables, narrows wide tables with top-K feature selection or truncated-SVD projection, and uses hierarchical label decomposition for classification problems exceeding the checkpoint's native 10-class head. Regression is represented as classification over 1,000 target bins, with predictions returned as point estimates or predictive distributions. The package requires Python 3.11–3.13, a CUDA GPU, AutoGluon with the Mitra extra, and the TabArena execution wrapper. It uses checkpoints hosted on Hugging Face, including `autogluon/mitra-classifier-2` and `autogluon/mitra-regressor-2`; FlashAttention is optional. The project is distributed from its Hugging Face repository and can be installed from a cloned checkout.
Mitra-v2 Classifier is a tabular foundation model from AutoGluon and Amazon for classification tasks. It is pretrained entirely on synthetic datasets sampled from random classifier priors, including a Hybrid SCM prior, rather than on real-world datasets. The model uses a 12-layer 2D Transformer with attention across both rows and columns and incorporates in-context learning. Its recommended recipe fine-tunes 50 steps on a user's table and bags eight model copies for prediction; it can be run through the mitra-finetune package or used as a Mitra model in AutoGluon Tabular. The repository describes it as the second generation of Mitra, with a longer context, more features, and an improved optimizer than Mitra-v1. The weights are distributed on Hugging Face under the Apache-2.0 license and require a CUDA GPU for the documented fine-tuning and eight-fold bagging recipe.
Mitra-v2 Regressor is a tabular foundation model from AutoGluon for regression tasks. It is pretrained entirely on synthetic datasets sampled from mixed random-regressor priors, including a Hybrid SCM prior, and uses in-context learning with a 12-layer 2D Transformer that attends across rows and columns. Regression is formulated as classification over 1,000 target bins: the model predicts a distribution for each row, whose mean provides the point prediction and whose probabilities can also produce quantiles and probabilistic metrics such as CRPS. Using the accompanying mitra-finetune package, the model is fine-tuned for 50 steps and bagged across eight copies on a CUDA GPU. It accepts tabular features and regression targets and returns point predictions or full predictive distributions. The weights are distributed through Hugging Face under the Apache-2.0 License; the package is required because standard AutoGluon expects the scalar head used by Mitra-v1 rather than Mitra-v2's distributional head.
Searchable transcript of Mitra-v2 vs. TabFM: Amazon’s New Model is 20x Smaller and Faster — AI with Surya (10:25). Search for a phrase, then click its timestamp to jump straight to that moment in the video.
Captions sourced from the original video on YouTube, published by AI with Surya. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.
00:00 So, Amazon released a model called Mitra V2. This is a tabular foundation model, meaning it is specifically designed for tables and spreadsheets. And [music] to be honest, this is the third one I've covered after Google's Times FM for forecasting and Google's Tab FM for tables. It almost feels like a new wave of these models are coming for everyday machine learning tasks.
00:22 [music] I will explain how Mitra is different from Google's Tab FM, but what I cannot get over with is this model was trained on completely synthetic data. It has never seen >> [music] >> a real table. So, is this the end of data science? We will find out by the end of this video. >> [music] >> So, for 22 years, the first thing I told every customer when I met them was, "If you want to a prediction out of your data, someone has to actually train a model on your data."
00:49 A few weeks ago, when I talked about Google's Tab FM, that was the first time I doubted my own conviction, right? And a lot of you watched that video. [music] Mitra V2 is the second time I'm feeling the same. Both are foundation models for tables, but they work differently. And the whole app that you will see in the demo live is actually built [music] around that difference.
01:10 Tab FM takes your example rows as input and then predicts in one pass and nothing inside it changes. Mitra V2 by default first fine-tunes itself on your table for 50 steps, which is about a minute on it. Then runs eight copies of itself on different slices of your data and then averages their answers. So, its weights actually change for your table. It came out of Amazon's AutoGluon team.
01:39 [music] It is about 20 times smaller than Tab FM and on more than 300 real data sets, it lands in the same spot. And it is open source Apache 2.0, so you can download it and use it. The reason I care is the old way. Somebody has to clean the data and argue about which columns matter, then try three or four models and tune the best one. And a week [music] later, you have a baseline.
02:06 With Mitra V2, you have that baseline in literally a few minutes. So, as always, I wanted to test this out myself, so I have built the app >> [music] >> and I will be doing a live testing. But before we go there, a quick disclaimer, all opinions are my own and do not belong to my employer. With [music] that, let's get into it. All right, so what you're seeing on the screen is an app that I built in order to explain how the model works.
02:31 Every call to the model is live and the data is also real, and we will see all of that as part of the demo. You are able to also build something similar using antigravity or cloud code or any agents of your choice because the API and the code is available on GitHub. So, I'll share the link for that repository as well, which the Amazon team has directly made it available.
02:52 Now, there are two different use cases that I'm trying to demonstrate. [music] We will run at least one of those, and second one it's going to be very similar. The first one is really predicting [music] the price of the house. So, we will have 2,000 records of data, and then the second one is really predicting machine failure. So, both of these are very common kind of a use case when it comes to the world of machine learning.
03:14 So, let's look at the data now. So, when you look at this for the house prediction, we've got 2,000 rows out of that 1,800 are known and 200 are hidden. That means that is the one which we have to predict. And there are like nine columns over here, you can see the actual data over here, right? So, this is like as real as it could get. And you can see the distribution of the columns as well.
03:33 So, that is one, and then this one is the the other one other use case. You can see the data set over here. I've got 5,000 rows of data, 4,000 is known and 1,000 we have to predict, right? So, 1,000 are hidden. So, that is what the data set is all about. So, I'm going to select the the house price prediction data set in order to actually explain you the demo.
03:53 Now, this is an important one because this kind of helps us understand and [music] get us grounded on the difference between traditional machine learning steps and the MITRE foundation model. So, if anybody is a data scientist watching this or machine learning expert, you will very quickly understand that these are the eight different steps that any machine learning project will have.
04:13 You'll have to collect the data. Let's take that we're trying to predict a credit card fraud, right? So, you're going to collect the credit card statements data, and transaction data, and then you will be doing a lot of cleansing, and this is where a lot of time goes. So, you'll be fixing missing values, taking care of outliers, and stuff like that.
04:28 So, that is the second step, and then you will be creating features. So, that is where feature engineering comes into place. You will look into two different maybe variables, two or three variables, and combine them or create new variables altogether. And then, after that, once this sort of like data cleansing part is done, then you will pick up the right models or the algorithms, and then you will do the tuning, right?
04:49 So, that is called hyperparameter tuning. Then, you will do the training of that model on your data. And then, once the model is trained, you've got some of the weights of your variables, then you will go ahead and validate, that means check if the model's performance is good or not. If the performance is bad, then you repeat this, and if the model's performance is good, then you go ahead and deploy the model, right?
05:08 So, that is traditionally how it is done. Now, when you look at MITRE, this is very interesting because this is a pre-trained model, which is and in other words, it has already prac- ticed based on 29 million synthetic tables. So, it's a foundation model, right? So, like it doesn't have to get trained, that is the whole point. And then, you point it to your own table.
05:29 So, what it does, and this is where the big difference between this and Tab FM is that it will do a fine-tuning step. And we will see this live in the demo as well, but here it will look it it will run for 50 times and look into your table, and then it will do a quick fine-tuning of the model. And then, on top of it and it also run eight different copies of the output and then it it'll do a voting and then create an average and that will become the answer.
05:54 So, it's a very interesting one. It takes a little bit more time, but then the idea is this is a smaller model based on completely synthetic data, okay? So, now let's look into this actual flow. So, as I was explaining, step one is really it has already done the pre-training. So, this is something which we don't have to worry about. Step two is actually reading the data and that is very interesting.
06:16 So, if I run the sweep here, you will see that I'm trying to showcase how it is reading both columns and rows together. So, it is going to complete the sweeping very quickly. So, it has done that and then the next step is the fine-tuning. This one usually takes a bit of time. So, you can see that the model is being loaded and now you're seeing like each step where the fine-tuning is happening.
06:36 So, as part of this, what happens is it first loads the 77 million weights into the graphics card. Then it turns our table into numbers the model can actually read, right? So, embeddings. Then comes the part that matters, which is 50 small adjustment on 1,800 homes it already knows the price of checking each one against this slice that it is holding back, right?
06:58 So, you can see the line it goes back and up and it does it as often as the adjustment it needs to make. And at the end, what you would see is it will lock in the weight. So, we are already progressing. Approximately takes around like a minute to do that and you can see the error rate also being shown, right? So, now you've seen that step 50 steps are complete.
07:17 So, we're going to stop the fine-tuning and then we will go ahead and do the voting. All right. So, now when I come to step four, when I run the vote for this particular row and remember I could have done this for any other row as well, what you're going to see is it's going to create votes for eight different rows and it is going to automatically find the average, which is 238,569.
07:39 Now, we can also look at the final result over here where you can see that if I reveal the actual, this is going to be off by 15,000. Now, this is for this particular row. Now, what we can do is we can look at the entire result over here. So, you can also download this and here you can see the root mean square error, R square, mean, all of those interesting things, but then this is the main one.
08:01 You can see the model prediction. Some of the model prediction, like this is error is only $2, right? Which is very good. So, you can see how well it was able to produce the answer, right? So, pretty much like how you would be doing with traditional model and you are able to see this as well. So, if you compare it with the standard baseline model, so this has got 16.5% lower error and close to 5,378 saved per home, which is pretty fantastic so for a data set like this.
08:32 So, that is what I wanted to show you and then you are able to download this as well. But, this is very interesting. You can do the same thing for this particular use case as well. So, I ran this prior to running this demo and here also like you could see this was a confusion matrix, right? The previous one was regression, this one was a classification to you like confusion matrix and all of those kind of things and it is pretty good accuracy, precision, recall and F1 score and AUC as well.
08:58 So, this again like doing really well for both classification and regression use cases. This model, at least for the kind of data that I have given it, seems to be performing really well. Okay? So, that was the demo that I wanted to show. So, let's bring it back home. A model that has never seen a real table just priced 200 homes and if I ran the second use case, it would have flagged the machines about to fail on my laptop with just a handful of lines.
09:22 So, the big question we started with is did Amazon kill data science? And I think the answer is pretty clear. I obviously think that it did not. I feel that it has reduced some of the steps, right? For example, cleaning and tuning you do before you get a number. So, you can try this week if you have a GPU and a table with some sort of a question that you want to ask and the answers are already filled in.
09:47 Now, whether the table can be trusted, what a wrong answer will cost to you, that is still your job and I don't see a model doing that anytime soon. So, I don't think it will replace data science or kill data science. At the max, it is going to shorten the cycle and help you do things pretty quickly. So, do let me know in the comment section what you felt about this video.
10:07 I usually cover Google stuff and this was the first time I was covering Amazon stuff. So, it'll be interesting to hear your perspective and obviously if you have questions about the model as well. If you're new here, please feel free to like and subscribe the video as well. Thank you for your time. Thank you for watching and I will see you in the next one.