← All transcripts

Ex-YouTube Engineer Rebuilds YouTube in 45 Mins Transcript, AI Summary & Key Points

ByteByteGo · 6 days ago · Science & Technology · 43:50 · EN-US

Watch on YouTube

Answer

A YouTube-style application can be rebuilt in under 45 minutes with AI agents, but only as a small approximation with mock data and several missing production features.

AI Summary

AI coding agents can build a small YouTube-style application in under 45 minutes without manually writing code. The resulting CatTube application includes a homepage, watch page, dynamic PostgreSQL-backed data, an admin dashboard, AI-generated thumbnails and videos, multimodal search, related videos, authentication, subscriptions, dark mode, and deployment through GitHub and Vercel. It remains a limited approximation: video uploads are absent, the content is mock data, and the videos are AI-generated.

Key Points

  • The project focuses on YouTube's homepage and watch page rather than reproducing the full YouTube product.
  • A former YouTube engineer says modern YouTube.com took over three years to build, while the simplified version can be built in under 45 minutes without manually writing code.
  • 00:39 Prerequisites — Familiarity with an IDE, terminal, and some web development helps, but software engineering and coding experience are not required.
  • 00:39 Prerequisites — Agentic development services, Cloudflare hosting, and FAL.AI may require payment; everything else used in the project is free.
  • 01:51 Problem definition — The homepage includes a top bar, search, avatar, sidebar, video and Shorts shelves, while the watch page includes a video player, metadata, comments, and related or recommended videos.
  • 01:51 Problem definition — Breaking the pages into small components lets coding agents implement them incrementally.
  • 04:51 Project setup — A new application is scaffolded in Cursor with an agent using a default model, while Codex, Opus, and Gemini are available alternatives.
  • 06:23 Generating the UI — GPT image 2.5 generates interface images, which can be passed directly to an agent or converted into code through Figma or Pande.

AI in practice

Used for

Agents

  • Cursor agent — Create the initial application scaffolding. 2 held 05:04
  • Cursor agent — Fix visual mistakes in the generated homepage. 2 held
  • Cursor agent — Connect the application to a Supabase-hosted PostgreSQL database and migrate static site data. 2 held 15:21
  • Cursor agent — Generate and populate the site's mock content pipeline. 2 held 22:53
  • Cursor agent — Implement related-video retrieval and multimodal search. 2 held 32:43
  • Cursor agent — Add authentication and clean up the application. 2 held 34:33
  • Cursor agent — Initialize the local project in GitHub and connect it to the desired repository. 2 held
  • Cursor agent — Fix the failed Vercel deployment. 2 held 41:17

Tools & resources

13 items

BNo. 4495
AIAINotes.us Tool

Better Auth

In the AINotes directory

Better Auth is an authentication library used to add sign-in and protect an application's administrative interface.

Mentioned in
1 video
Kind
Other
CNo. 0953
AIAINotes.us Tool

Cloudflare, Inc.

Open source · cloudflare

Cloudflare is an internet infrastructure and security platform that provides hosting and delivery infrastructure for applications, agents, websites, and data. Its network combines compute, connectivity, and security at edge locations, with Workers running code close to users and backend data; the company states that its network operates in more than 335 cities and reaches 95% of the world's Internet-connected population within 50 milliseconds. Cloudflare also provides AI crawler controls that let site owners inspect, allow, or block crawler activity, along with pay-per-crawl and a monetization gateway for pages, data sets, APIs, MCP tool calls, files, and search indexes. The described x402 flow uses an HTTP 402 response to request payment, after which an agent pays, retries with proof, and Cloudflare verifies the payment at the edge.

Mentioned in
6 videos
Kind
Other
CNo. 0348
AIAINotes.us AI product

Cursor

cursor.com

Cursor is an AI-powered coding agent and integrated development environment designed to accelerate software development by handing off coding tasks to AI. It evolved from an email client into a multimodel development tool, supporting broader developer workflows. The platform also offers MCP-connected capabilities for tasks like searching and editing notes.

Mentioned in
33 videos
Kind
AI
DNo. 3048
AIAINotes.us Tool

Drizzle ORM

orm.drizzle.team

Drizzle ORM is a lightweight TypeScript ORM for defining database schemas and querying relational databases. Its migration tooling can generate, apply, push, pull, export, and check schema migrations; the documentation covers PostgreSQL, MySQL, SQLite, SingleStore, MSSQL, CockroachDB, and related database services. In the cited context, Drizzle owns the PostgreSQL schema and migration workflow for effect-mq.

Mentioned in
2 videos
Kind
Other
FNo. 3557
AIAINotes.us AI product

fal.ai

fal.ai

fal is a generative media platform for developers that provides a unified API and SDKs for running image, video, audio, 3D, and code-generation models, including open models, user-provided LoRAs, and custom model endpoints. It offers a gallery of production-ready models and runs inference on a globally distributed serverless GPU infrastructure that can scale deployments from zero to large numbers of GPUs. fal also provides on-demand GPUs, compute for training and fine-tuning, private model endpoints, dedicated clusters for custom workloads, and observability tools for monitoring deployments and usage. The platform develops and post-trains open-weight models such as MiniMax H3 to improve generation speed, cost, and control for real-time video applications.

Mentioned in
3 videos
Kind
AI
HNo. 4105
AIAINotes.us AI product

H3 Max

In the AINotes directory

H3 Max is fal’s post-trained video-generation model, based on MiniMax’s open-weight video model. It combines model post-training with systems and hardware optimization to reduce generation time while maintaining video quality, supporting fast generation, interactive streams, and professional video workflows. Its continuous-video experiments can retain previous scene context and respond to new directions while a stream is running, with development focused on controllable camera movement, lighting, characters, motion, and lip sync.

Mentioned in
2 videos
Kind
AI
NNo. 0895
AIAINotes.us Tool

Next.js

Open source · vercel/next.js

Next.js is an open-source React framework developed by Vercel for building full-stack web applications. It extends React with server rendering, static-site generation, hybrid rendering, components, and routing, and integrates Rust-based JavaScript tooling for builds. The project provides documentation, a learning course, a showcase, community discussions, and contribution guidelines.

JavaScript
Stars
★ 142,757
Forks
33,232
ONo. 0214
AIAINotes.us AI product

OpenAI Codex

openai.com

OpenAI Codex is an AI coding agent from OpenAI available as a command-line tool (Codex CLI) that helps developers produce software. It can be used alongside Gemini for adversarial audits of software requirements and implementation plans, listed as a supported coding-agent or model option in several projects, and its logs can be joined with task and test evidence. The Codex CLI can also receive and answer requests from the Penako canvas.

Mentioned in
55 videos
Kind
AI
ONo. 0368
AIAINotes.us AI product

OpenRouter

openrouter.ai

OpenRouter is a platform that provides a unified interface and API for accessing and comparing multiple AI models and their pricing. It aggregates model endpoints so developers can route prompts to different providers from a single place.

Mentioned in
17 videos
Kind
AI
PNo. 2016
AIAINotes.us Tool

pgvector

Open source · pgvector/pgvector

pgvector is an open-source PostgreSQL extension for storing vectors alongside relational data and querying exact or approximate nearest neighbors. It supports single-precision, half-precision, binary, and sparse vectors, with L2, inner-product, cosine, L1, Hamming, and Jaccard distance operators. The extension uses standard PostgreSQL tables and queries, including inserts, updates, bulk loading with COPY, filtering, joins, and indexed nearest-neighbor searches. It retains PostgreSQL capabilities such as ACID compliance and point-in-time recovery, and provides quantization options for scaling vector collections. It supports PostgreSQL 13 and later and can be compiled on Linux, macOS, and Windows or installed through Docker, package managers, PGXN, and hosted PostgreSQL providers.

Mentioned in
2 videos
Kind
Other
PNo. 2017
AIAINotes.us Tool

PostgreSQL

postgresql.org

PostgreSQL is an open-source relational database management system maintained by the PostgreSQL Global Development Group. It provides SQL-compliant, ACID transactions, MVCC concurrency, and extensibility through custom types, functions and extensions, and is used for general-purpose and enterprise database workloads.

Mentioned in
8 videos
Kind
Other
RNo. 4496
AIAINotes.us AI product

Rebuild YouTube with AI

live.bytebytego.com/courses/youtube-ai

Rebuild YouTube with AI is a live, hands-on cohort course from ByteByteGo about building and deploying a YouTube-like video site with AI coding tools. It covers decomposing the product into core surfaces, generating interface mockups, implementing React and Next.js pages, modeling data in Postgres, adding authentication and uploads, generating media, storing video with Cloudflare, and deploying with Vercel. The course also implements search and related videos with multimodal embeddings and pgvector, then discusses recommender-system limitations, watch-time tracking, AI-driven Playwright testing, and production workflows. Sessions are recorded, and enrollment includes course materials and access to the cohort recordings after the live course.

Mentioned in
1 video
Kind
AI
VNo. 0598
AIAINotes.us Tool

Vercel

Open source · vercel

Vercel is an application and agentic infrastructure platform for deploying and hosting web applications, marketing sites, platform products, and AI agents. Its infrastructure includes global delivery, deployment environments, serverless functions, fluid compute, a web application firewall, durable orchestration, sandboxed environments, and an AI model gateway. For hosted platforms, Vercel provides tenant isolation, domain management, custom SSL certificates, and preview URLs.

Mentioned in
6 videos
Kind
Other

Links mentioned

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free

Transcript

Searchable transcript of Ex-YouTube Engineer Rebuilds YouTube in 45 Mins — ByteByteGo (43:50). Search for a phrase, then click its timestamp to jump straight to that moment in the video.

Captions sourced from the original video on YouTube, published by ByteByteGo. The video, its captions and all related intellectual property remain the property of their respective owners; AINotes claims no ownership. Provided for research, accessibility and search — see the Transcript Notice and Copyright Policy.

00:00 Today we have a special guest, Mike, a former YouTube engineer. He will show you how to rebuild YouTube with AI from scratch in 45 minutes. It took me over three years to build modern YouTube.com as a staff software engineer at Alphabet. I'm going to teach you how to do the same under 45 minutes without writing a single line of code. My name is Mike and let's get to building.

00:22 Let me introduce myself properly. For 8 years, I worked at Alphabet on YouTube and other web properties. I then moved to way more where for four years I've been helping to launch it in San Francisco and beyond. I then started my own company released pony diffusion and now running fictional. I will be a guide for today. What kind of technical skills do you need to follow along?

00:42 If you are familiar with an IDE, can work in a terminal and have some amount of web development skills that would help you understand what's going on under the hoop. But you don't need to be a software engineer. You don't need to know how to code. As a matter of fact, we will be using our ID very lightly and only to set up a few things. So follow along even if you have none of the skills and you're going to have a lot of fun.

01:08 We will be using a small set of third-party services in our development and some of them do require payment. So this is obviously the next important question. If you don't have a subscription for an agentic development tool like cursor or codex, it's probably a good idea to get one. We'll also be using Cloudflare to host images and videos, and it has a monthly subscription, but I personally think it's totally worth it because of how much time it's going to save us.

01:35 And we will be using FI for image and video generation because we'll be creating a lot of mock data for our website and it's going to be looking really good. And that's pretty much it. Everything else is free. So let's figure out what can we reasonably ask from agents today. Obviously the dream is to tell them to do everything. Build the product, populate it with data, deploy it, promote it, get some real users, maybe even run some ads on it.

02:06 But we are while rapidly approaching it, still probably a year or two away. So we will be this interesting and slightly weird combination of a product manager and technical lead instructing our agents to build specific things. And for this exercise we'll be focusing on the homepage which is the face of youtube.com and the watch page which contains the most important surface which is the video player.

02:31 So let's look at the pages and try to figure out roughly what components we see. If we look at the homepage, it also has the top bar which contains a input box for the search, your avatar if you sign in, the sidebar which has endpoints to important other services of YouTube and multiple shelves of either videos or shorts that actually lead to the watch page.

02:56 And if we want to check the watch page, some of the elements remain the same. Top bar is there. Sidebar has disappeared. And we can because we wanted to maximize the amount of surface dedicated to the video, but it can still be shown. We have the video player, the most important part of the watch page. We have the metadata, the description of the video, its title, the social aspect with comments from other users.

03:19 And in a critical and very important component, the sidebar on the right, which contains the recommended videos and the related videos and autoplay videos. So pretty much a very limited set of building blocks that have a lot of technology backing them up. But we can split the pages into small pieces and implement these pieces one by one. And for that you could use codex, you can use cursor pretty much any editor which can use agents or even just a gentic CLI is sufficient.

03:54 I like cursor. I'm more used to using something that looks like IDE and has an ability to involve patients with it. But you're free to do whatever you want. Now, something interesting about the original YouTube.com is that at a time we were building it a few years ago, we focused on using what was considered state-of-the-art. We were looking actively at web components and specifically a framework called polymer which allowed us to use web components within the website.

04:24 Now the times have changed. Polymer is no longer around. Web components didn't survive the test of time. So we're going to be using slightly different set of tools. And my recommendation for this project is to use JS which is a very solid framework. It's going to help us avoid a lot of headache and generally just really well established TypeScript react based frameworks that very popular nowadays.

04:52 So let's get to work. We are going to move to cursor now and start implementing our very basic scaffolding using the netj before I get to building something cool. You can see that I just open a completely new folder. There are no files on the right. This is a completely new project and I will open up a new chat in this project. Cursor allows you to use multiple models.

05:16 It supports planning and agentic modes that are useful in many cases, but right now we're just going to go with default. It's perfectly fine for whatever we're going to try to achieve right now. It will default to the Gro model. You have a few options. You can go with Opus. You can go with Gemini. And the prompt that you saw before actually not a chalk.

05:39 We're going to be using that prompt exactly like I typed it in the slide. So let's fire away and let's see what the agent can cook for us. It took us a few minutes of waiting and I had to proof a few commands that the agent wanted to run for us and wasn't sure if that's safe to do. But we have a number of files generated by our agent. Our agent has reported that all of the operations has been completed and we have an exjest application running in the browser inside of curs.

06:12 That's not very impressive at this point, but it's our starting foundation. And now we can get to building something that resembles youtube.com. What we want to do is we want to capture the interface of the application that we're building. In our case, obviously the front page of youtube.com. And there are a number of text to image models that can help us to do this.

06:35 The modern GPT image 2.5 is very very good at building UI interfaces. So this is the model that I personally going to be using and the model that I recommend. But there are many other models. Nana Banana by Google available through Gemini is another great model and even open- source models like Quen image would be reasonably good at implementing this task.

06:57 Now when we have the image of the interface that we want to build, we actually want to convert the interface into code and we have a number of different strategies that we can use. A very popular way of doing this is using Figma. Figma is a service that you will see being widely used in real businesses in real companies allows collaboration between designers and engineers.

07:20 So it's a very popular strategy for a lot of real products. An alternative to Figma would be pande which is a local version of Figma. It allows you to connect to your own LLMs. It's another great tool and if you don't want to be stuck with Figma, you can use it alternatively. The third option is giving the image directly to the agent bypassing any kind of external application services.

07:45 It doesn't work in all cases but in our case would work beautifully. And this is actually the strategy that I'm going to be working with by taking the image generated directly by the GPT image inside of Codex and giving it to the agent. Or you can just wing it and in many cases you can just tell agent what to implement. In case of services like YouTube which are really wellnown and all of the large language models are well aware of what YouTube is that would work really well.

08:15 If you're building something of your own, some a complex UI in some cases text description may not necessarily be the best and image would be more beneficial. I have now switched to Codex by OpenAI and this is where I'm going to give it the command to generate something that resembles a homepage of YouTube.com. The prompt that I'm using is once again extremely simple.

08:43 Partially because prompts nowadays can be pretty simple, but also because every large language model is well aware of what real youtube.com is. Just one minute after I issued the command, we have a version of YouTube. And if we look at it, it's pretty reasonable. It does look like real YouTube. It has a number of icons that are all about cats. It generally follows the colors and the structure of YouTube.

09:12 I'm pretty happy with this. And this is something we can now move to our cursor. Again, back to cursor. I'm going to drag the generated image inside of the cursor and ask cursor to implement this version of cat tube. It will take us a few minutes. The agent will have to perform quite a lot of work, but we will be back in probably about 5 minutes and see the progress that the agent was able to do at the time.

09:55 All right, so we had to wait a few minutes and we generally got what we wanted. It's a version of our cat 2. There are a few mistakes here and there, but we've got all of the thumbnails 3D. We've got even functional selection of categories on top. Obviously, you can click on anything, but we only asked it to do the homepage so far. So, what we can do now is we can ask it to fix a few mistakes that are seeing on the page.

10:24 But one interesting thing that I wanted to confirm is if I go through different resolutions on the screen, I'll actually see that not only it implemented the homepage, it did implement it in different layouts including mobile layout. So let's fix the mistake that I see right now by asking agent one more time and see what it will be able to do. With just one command, we're able to actually fix our application.

10:49 It looks much better now. And in the process of building it, our agent actually did a lot of important things. They implemented the general structure. They created some logic to make the top categories functional. They also took our original image that contains thumbnails for avatars and the videos and it cut the part of our original thumbnail so it can insert them in our more or less functional version of the homepage.

11:19 So now that we have this page, we're going to be moving to the watch page. Returning to products, we already have a session. So we're just going to continue the session asking chat GPT to implement the watch page now. All right, let's check the image that was generated for us. Feels very reasonable. The layout here is a little bit weird in some places, but nothing that we can't really fix.

11:45 So let's move this to cursor and implement the watch page. text. We're going to pretty much replicate what we did with the homepage. I'm going to drag the image that we just generated into cursor and just ask to implement this page. This once again may take a few minutes, but because the agent was already able to implement a significant part of our layout and there are a lot of pizzas that can be reused, this time it's going to be significantly faster process.

12:20 Now, as you can see, a few minutes later, we actually got a functioning watch page implemented. The sidebar is exactly like it's on real YouTube, hidden by default. We have a video player. Uh, while this doesn't do anything right now, even the time progresses, but the rest of the page seems also very reasonable. All of the elements are here. We can actually click through the pages and see how we even have somewhat fake thumbnails show up.

12:54 But this is a very good start. Obviously, this is not real YouTube, but we're getting there. We got the homepage, got the watch page. So, let's continue and develop this further. Now, we'll have to make a detour and actually think about how to make our application dynamic. Right now we just have for the most part a static website. Yes, you can go through a few of the video pages.

13:21 It has content that resembles real one, but this is all fake. And in order for us to make it somewhat more realistic and dynamic, we need to connect it to database. There are numerous options of databases available to us, but I'm a big proponent of just using posgress for pretty much everything. It's a great database. We're not going to go into much detail on why you should use posgress, but I personally love it.

13:45 So, I'm going to use it. And in order to make our life even simpler, instead of installing it locally or you know working with Docker, we are going to be using a service called Superbase. It is free to try. You can host a database there. It's very easy, convenient, and provides us with a lot of tools. And in order to connect to the database that we'll be hosting on Superbase, we will going to be using Drizzle.

14:11 You don't need to use an RM. You can actually just use something lower level as Sloanic or even write everything by hand, which I wouldn't recommend. I like using RMS in places where it's useful. So, we're just going to start setting our project. And the first thing that we need to do, we're going to create the database and register on Superbase and connect it to our project.

14:39 So I went ahead and I've created a new database within Superbase. And what we want to do now is to connect this database to our application and instruct it to use this connection to store the data. If I go to my superbase uh account and if I click on the connect it will provide me with a list of different ways to connect this database. We are using ORM.

15:03 So I'm going to click on it select dri and it will give me an database URL but it also gives me a whole prompt so I can just instruct my coding agent to use this particular setup. Now I've pasted the exact prompt that was given to me by the superbase interface into my agent and I'll ask it to implement it. Perfect. So now we have created the necessary databases.

15:36 We have connected to the database. Everything works. Now we can actually migrate the data that we have an application to use this database and also create an admin interface which will allow us to edit the data. This task seems to be pretty large to me. There going to be a lot of moving pieces. So instead of just using an agent strategy, I'm going to switch into the planning mode and I will let the agent first try to figure out what steps are necessary to actually get this pretty complex task done.

16:11 As you can see, the agent is asking us some questions. First question is if we should be using some kind of authentification in the real project. Obviously, we're going to be having something. But we're going to be working on a notification a little bit later in this project. Perfect. So, the agent has created the plan for us. We're going now to quickly review this plan.

16:34 And as you can see, it created the schema for us. It went through all of the critical parts of this database. It actually went through quite a significant amount of reasoning and basically presenting us with a very reasonable plan that seems to be covering everything that I wanted. So, I'll go and move forward with letting the agent build this because everything here makes sense to me.

16:57 Everything that I requested is being covered. We are going to be creating real database objects. We're going to be populating them with the data that we currently have statically included in the website. And we're going to be building an interface to go and browse through this data and edit and add more if we desire to. The agent has finished processing the data and we have two things to check.

17:24 First of all, we have the same website. It's operational. We can continue clicking through it and seems like all the data is still there. Obviously, not a big change from what we had before. But if we now navigate to the special new URL that was created by the agent, we can see that we have a dashboard. And if we go through the data that dashboard represents, we see all of our videos.

17:52 We can go and edit these videos. We can even create a new video. Let's go back to comments. And all the comments are still here. So all of the data that was static before is now dynamic and available for us to be edited. But let's double check that this data actually exists in our database. And switch back to superbase and check what is inside of our database.

18:23 And indeed we can see that if we navigate to our table editor inside of superbase and open our comments table for example all of the data is here. Let's go to the videos. All of the data is available here. And the agent even created a few other tables. All the video related objects are represented here. Categories are represented here. All of the channels are also here.

18:50 So, we basically moved all of the data that was statically allocated before into a database. It doesn't help us immediately, but this is an important step to doing something actually much cooler. Now, because we live in the era of large language models and generative AI, we can use a lot of tools available to us to bootstrap the content that we have on our website and get some really cool and amazing videos of cat.

19:17 But before we get to that, we're going to connect Cloudflare. It's a third party service and unfortunately this is one of the services where you kind of have to pay for it. But it solves two very important problems that we have to deal in our project. First of one being hosted images. This is technically not required. You can do it yourself. But having a third party with a very good content delivery network serving your images, supporting image resizing, being able to just serve them very fast and efficient is really

19:49 cool. It's a very popular service. Majority of large companies are using it for image hosting. But for us, a much bigger problem is video hosting and video streaming. The real YouTube is a giant machine on the back end that knows how to process videos, how to re-encode them, split them, extract metadata, and generally be able to deliver them in a way that's efficient across many different platforms.

20:17 This is something we definitely can build, but it's going to be a project significantly larger than what we're building right now. So, we're going to rely on Cloudflare to serve videos for us in the CloudFlare player to play this videos because this is going to significantly simplify our application and help us to get to the finish line faster. We are going to connect them to our application in the next steps, but we actually don't need to do much.

20:44 I've already configured Cloudflare. I already imported the necessary permissions and tokens into my project. So we're just going to be asking agents to use Cloudflare as it will populate it with data. And this is the next interesting process. We're going to be generating this data. We're going to be creating cat videos in real time. How we're going to be creating these cat videos.

21:06 We need three different things. I'll just going to walk you through something that I personally like and found very useful. We need three components. a large language model that would help us to generate the name of the video, the general idea of it, maybe some metadata. Basically, it will elucinate a possible variation of a cat video that exist. Then we're going to be using an image model to generate the thumbnail and this thumbnail will be based on the LLM produced description of the video.

21:37 And then we will be using this thumbnail to generate a video. Now, technically, we can do it without the thumbnail, but I like the thumbnail approach because it gives us the thumbnail that we'll need anyway, and it gives us a lot more control over what we're going to be seeing. In order to do this, I will be using FAL.AI, which is a great aggregator of many different models, both text and video.

22:01 And uh inside of FAL, we're going to be using Open Router. This is a little bit convoluted because Open Router actually is an independent project. You can use open rotor on your own. But the reason why I'm using it through f I can just use one API key and have access everywhere. So we're going to be using open rotor and we're going to be using GPT image 2.5 model to generate thumbnails.

22:30 And for video generation, we will be using H3, specifically H3 Max variation which has been developed by FAI. And it has a turbo version which allows us to generate videos up to 15 seconds really fast and very cheap because video generation is still expensive. But this way we can get really amazing cat videos for really cheap. Let's move on to implementing our video creation pipeline.

22:56 We need to wipe the existing data that we have right now in the database. Generate mock users, channels, comments, and also generate videos. And for videos, it's a little bit more work because we both need metadata, the title of the video, the description of the thumbnail, and the description of the video that we will use to generate the thumbnails and the videos.

23:15 Then we need to generate the images, the thumbnails. And then we should generate the actual video using the H3 model, upload them to Cloudflare and make sure that this videos and these images are now used on the website. This is definitely a case where going through the plan mode would be great because a lot of action is going to happen. I would like to review them before the agent gets to building anything.

23:40 you may have noticed is that I'm using still the same chat session and in this particular case it makes sense majority of work that we're doing is just one continuous project we're building our cat tube while it's not critical a lot of the context of what happened before will be useful to the agent to understand where we're going we have still plenty of context available and there is no real reason to not continue using the same chart moving forward all right let's view what the plan is currently.

24:12 I really like this diagram because it actually reflects what I was talking about the whole process of creation of the video. Obviously, unsurprisingly, but it's good to see it visualized. We've confirmed that we want to have 12 videos. It's exactly the amount of videos we have right now. But as the image and video, especially video generation, can be pretty expensive.

24:30 It makes sense to start with just a few images. I see that the models that agent picked up are very reasonable. I have no concerns here. It picked up on the Cloudflare hosting. That's all great. The schema will be updated because we need to make the videos play. Makes total sense. It will create a script that also makes a sense. The script will populate the data.

24:58 Yep. Perfect. Everything here makes sense. So, we can get to building. We are back and the thumbnails has changed. If we go into one of them and see what we had generated, we get a completely AI generated videos and everything that you now see on the screen is another AI generated video which pretty cool. They are limited to 5 seconds so far, but we can work on that and we can improve the resolution.

25:32 But for now, for this test run, that's pretty good. I think that looks very, very reasonable. And it's starting to look like a real website at this point in time. While we're still using mock data and we're not letting users upload anything, if we deploy this right now, this will be a fully functional version of our catu. But there are a few really interesting things that YouTube does that we are currently completely missing.

26:01 And I want to focus on two of the things. First of all is search and the other thing would be related videos. So how we going to implement search and recommended videos? Well by cheating. You see in a real service like YouTube there are massive teams handling both of these problems. and both search and related or recommended videos in many cases are life or death for many services because being able to recommend exactly the videos that you want to watch is probably one of the most important things that a service like

26:36 YouTube can do. We do not have capacity to implement it at a full scale. So, we're going to do something small and we have to think about it. What are we actually doing? When you search for a video, you try to find something that is similar to your text query. And then something if we just think about this on a high level is a video. But what is a video?

26:57 We can think about video as the title of the video. Maybe we can look at video as a thumbnail or the longer script of the video, maybe the subtitles, maybe the video itself. So there are a lot of very different components to the video that we can use. And a very simple way of searching would be for example just using the title. If you typed cats in the search box, we will find all of the titles with the cat inside of them.

27:24 But that will be a significant simplification because what if you typed cat then we will never find the video which has tiger in it or the park. So we want to have something smarter. And then again, what are we using to figure out what are the video? The way I'm going to handle this in this project is I will think about two pieces of content. I will think about text and I will think about the thumbnail of the video itself.

27:51 Thumbnail is obviously not the whole video, but it gives us a very very good representation of what that video is all about. Now we have to talk a little bit about theory because this is required to understand how to calculate similarity. We will talk about something which is called an embedding. And embedding by itself if you look at it being represented on your computer is just a long list of numbers.

28:20 And these numbers have very specific meaning. But this is a meaning that we as humans can't really figure out. So you can just think about an embedding as a long list of numbers. Every title will have an embedding. Every video will have an embedding. The question becomes, well, how do you create these numbers and how do you use the numbers and why do we even need them in the first place?

28:45 Well, you see if we just think about just test itself, the way the embeddings are created is if you take words that similar and make sense to us as humans to be close to each other, they will be very close to each other. While the terms that are very different will have some amount of distance. We will be using both text and images. And this is where we can use models which are called multimodel models.

29:16 And they create embeddings for text and for images in the same embedding space means that the image of a lion is next to a text which says cat and the image of a cat is close to the word tiger. So this is extremely useful and this allows us to say these things are together while these things are not together and far apart. On a technical level in order to handle this similarity we use something called cosine similarity because this long list of numbers are actually vectors.

29:52 So figuring out the angle between these vectors allows us to understand if they're close or if they're not close. So if we create embeddings for our texts, if we create embeddings for our thumbnails, we will be able to understand where things are similar to each other and we can sort by that similarity. So if you have a video of a cat, we'll be able to find all of the other videos that are very similar to that particular video based on similarity of the thumbnails alone.

30:24 Simplification for sure, but it's going to work really well. Now we have two problems to solve. We need to figure out where and how to store this embeddings and obviously how to calculate them. So when it comes to storage for a significant amount of time, storing them in regular normal databases was discouraged because while you can definitely do this being able to efficiently look up similarity requires non-trivial math and traditional databases couldn't do it.

30:54 People used external special databases like Pine or VAVA. Lucky for us in 2026 posgress caught up and has a special extension called PG vector which works absolutely amazing. So we can store our embeddings in our regular database. We don't need to introduce any other component and we're back to basically just use Postgress for everything. I asked our agent to generate a few more videos just so it's more noticeable when we have videos that are related to each other.

31:31 I've opened a superhero cat video and as you can see the most relevant one is another superhero. If I will go to a sport video, the two top videos are both similar in terms of content and in terms of visual style. So, it's a uh black and white noir like video. If I go into something more fancy, the fashion show. We've got another anime video. Well, this one is interesting because K here has a magnifying glass and the most relevant video is another magnifying glass video.

32:12 But what I would like to do now is to implement search because this is only related videos. And with search, this is no longer about similarity between video and a video or in our case thumbnail and thumbnail, but between text and the thumbnail. And the complexity that comes here are from the fact that we are no longer prefilling this data on the back end.

32:33 We actually need to take the user query, we need to process it into embedding very quickly and then we can do our lookup. So that's the next task we're going to be handling. Now let's ask our agent to now implement the search by taking the query that the user provided, converting it to clip embedding, and then finding the most similar videos to that clip embedding.

33:02 All right, so it took our agent a few minutes to implement everything and let's see how it works. I'm going to try searching for noir. Perfect. Seems like, you know, we got a bunch of black and white images. Really like this one. Yeah, I would say this looks very noir. What about superhero? All right. Once again, this looks pretty superhero to me. What about sports?

33:42 Once again, seems like the search is now is fully implemented. So, we've got related videos, we've got search working, we've got videos being played, we've got thumbnails, we have mob data. It seems already like a pretty reasonable approximation of YouTube. We are still missing a few pieces. So, what I would like to do now is we have our admin interface which is just kind of opened to everybody.

34:05 We have faked our own user account. You can't really subscribe to anything. So, I think we're going to be doing two things next. First of all, we're just going to polish everything a little bit. And second of all, we're going to introduce a very simple authentification. We're going to be using a service called better o, but it's also a set of libraries.

34:26 So, you can use the service, you can use just the libraries, completely independent. And what we're going to be doing, we just need the libraries themselves. So let's ask our agent now to clean up a few things and get the authentication running. We are using extremely simple prompt. We are just going to ask our LLM our agent to implement exactly what I mentioned and there is a lot of work here.

34:52 So maybe going through the plan mode is reasonable. But I'm just going to use the gentic mode because well there a lot of work. It seems pretty straightforward and there are no weird caveats here. Well, looks like we were able to implement the outification successfully. A few things changed on the page. So, I now see that I need to sign in. And if I try navigating to the admin interface manually, it will present me with a uh notification form.

35:28 So, let's use an account that our agent created for us specifically to test things and confirm that we can sign in. And indeed, we can sign in. We can access our dashboard. Let's go to the website. Let's click through a few videos. Try the subscribe button. It's now working. Let's go to our subscriptions. And indeed, we now have videos from a few channels we've previously subscribed to.

35:59 So, this is becoming more and more like a regular website. But there is one thing that is missing, and this is something that's really makes me annoyed. If you use regular YouTube, you will notice that it actually has the light mode and the dark mode. And we only have the light mode. That's definitely a big oversight. So, let's just do a little bit more of a polish and add the dark mode because I just want our website to look really cool.

36:33 It looks like with just one prompt asking our agent to add the dark mode, we were able to fully implement it. Let's toggle it a few times. Looks very reasonable to me. I didn't have to fix anything. The watch page looks great. Everything here makes sense. I think even this cool little tidbit here is also changing colors. That's pretty cool. So, I really wish we had it back in the day because it took us months to actually make sure that YouTube looks good in dark mode and literally took me 5 minutes now.

37:10 So this is all great and I think we are now ready to show the world what we've been building because well a lot of pieces are still missing. We don't have video upload. All of our videos are still AI generated. I think this is pretty much a reasonable approximation of version one of what YouTube would look like. So I want to put it online. And in order to do this we have again many different options.

37:35 But because we're building on top of NexJS, we have a very simple way of hosting this by using Verscell. Versel is a company behind Nex.js. So it makes total sense that they provide very simple ways to deploy a NexJS project. But we're missing one thing is that the typical way of deploying a Versal NetJ application is by using GitHub. And we never created a project.

38:00 We never used GitHub here. It was literally a local folder on our computer. So let's put our code into GitHub and then connect it to Versel. So I moved to GitHub and I've created a repository for us. But here's the problem. I have never initialize this folder from this repository. And actually I'm not that good of a Git user to understand what to do here.

38:26 So why don't we ask our agent to do everything for us? Because well the agents are not just for coding. They are for doing pretty much everything. I asked my agent to put the code inside of the GitHub and let me tell you it took it a while because I have a complex setup on this machine. Multiple GitHub accounts, multiple versions of cursor talking to different repositories.

38:57 So, it had to actually go and figure out a lot of things just to be able to connect this particular version of my project into that GitHub account that I wanted to use. And I'm pretty sure that if it was me trying to figure it out, I would have been reading documentation for a few hours, but took it about 5 to 7 minutes. And now I can go to GitHub and check that files are actually there.

39:23 Yep, all the files are in our GitHub. We are ready to actually start connecting this project to Verscell. So I'm going to switch to Verscell right now. And it already recognized my GitHub repository because I connected my account previously by logging through GitHub. So I'm going to just go and imported it. And in order to fully set up our application, I had to do only two things.

39:53 I went to the application preset and selected NexJS application there and I went into the environment variables and I imported the file that I already had in my project because it's not part of what's being committed to GitHub. It contains your secret information. So you don't want to have it in your Git repository, but you still need it to fully deploy the website.

40:13 So it knows where to connect, what database to use, and all of the things that we've basically added through the development process. So now I'll move ahead and deploy this application. So looks like after deploying we encountered a problem. I can see that the preview of my service is not showing the website that I expect to see. I can see that the error rate is 100%.

40:38 So something is terribly terribly wrong. Let's go into the logs on Versal and do a little bit of debugging. So I can see that actually none of the pages that I tried to access worked. And if I go into the log I will see that it's complaining about not finding a module Onyx runtime for the note. Well, we use Onyx runtime actually to generate clip embeddings.

41:06 And this is not good because we made an assumption that this module will be available. But the way Versal works and the way it deploys is not fully compatible with the initial setup that we did. Well, we're just going to ask an agent to go and fix this for us. So, let's take a copy of this error, move to the agent, and ask it to fix this. Once again, my prompt was extremely simple.

41:37 I literally copy pasted the errors that I've encountered and I add just a little bit of context. I told my agent that this is a versell deployment and I asked it to create a GitHub PR because I would like to review what the changes uh before actually deploying them and it took it a few minutes but now we have a solution. So let's go to GitHub and check what was actually generated.

42:03 Perfect. So I see we already got a pull request that our agent created. Let's go and check it. And it went through a number of checks. And we can see what Versal is thinking about our change by previewing the build. And look at it. Seems like actually the website loaded this time. So let's return to GitHub. Let's merge this PR. Then let's go back to versal and see how things are looking.

42:39 Opposite side. Well, look at this. Our deployment has succeeded now that we've deployed the fix. And let's go open it up. Let's see what we've got. And congratulations. This is our very small but pretty cute and very cat friendly version of YouTube. Let's just go and click on a video of some kind. I like this one. Amazing. If you're interested in a more in-depth conversation about the engineering history of YouTube, how to build something like this today, have some questions for me, or just want to listen to war

43:32 stories from YouTube days, you can find me at Bite Bitego. When me and other real builders are sharing our industry experience, my name is Mike. It was a pleasure to be with you today and I hope you learned something new and useful. Take care.