AI inference compute hosting by deploying an open model on high-performance GPU infrastructure and selling access through an inference marketplace

Operate a GPU rack, deploy Kimi weights with an inference engine, expose the model through OpenRouter, and earn more revenue from serving the model than the compute costs.

From Dwarkesh Patel — Dylan Patel – Two labs will soon control most of the world's workforce at 14:37

Problem: Provides usable model inference capacity without requiring each user to own and operate high-performance AI hardware.

For: Users of hosted AI models who access them through an inference marketplace.

Products from this video

Antithesis Fable orchestrator Grok Bot OpenAI Codex OpenRouter SGLang vLLM

🔒 Full analysis locked

Unlock more videos and the full analysis

A credit unlocks one video's full analysis for good — the build steps, the tools and how each was used, the methods behind every use case. Pro opens the whole library instead, and raises how many videos you can analyse a day.

Unlock full analysis — free