AI inference compute hosting by deploying an open model on high-performance GPU infrastructure and selling access through an inference marketplace

Operate a GPU rack, deploy Kimi weights with an inference engine, expose the model through OpenRouter, and earn more revenue from serving the model than the compute costs.

From Dwarkesh PatelDylan Patel – Two labs will soon control most of the world's workforce at 14:37

Problem: Provides usable model inference capacity without requiring each user to own and operate high-performance AI hardware.

For: Users of hosted AI models who access them through an inference marketplace.

🔒 Unlock the rest of this idea →