Helps size private model infrastructure by calculating KV-cache, serving-capacity, and concurrency limits, then comparing suitable hardware, sovereign rentals, and enterprise subscriptions with labeled evidence profiles.
Estimates the binding constraint for deploying an on-premises LLM by calculating KV-cache memory, prefill and decode capacity, and runtime session limits, with Markdown or JSON export.