AI product Open source
KTransformers is an open-source research framework for efficient large-language-model inference and fine-tuning through heterogeneous CPU-GPU computing. Its user-facing capabilities are provided through the kt-kernel source tree: high-performance inference and supervised fine-tuning (SFT). For mixture-of-experts models, it uses heterogeneous expert scheduling, placing frequently used or “hot” experts on GPUs and other experts on CPUs with optimized AMX, AVX512, or AVX2 kernels and NUMA-aware memory management. It supports CPU-side INT4/INT8 quantized weights, GPU-side GPTQ, native BF16 and FP8 precision, and multiple CPU, GPU, and NPU backends. The SFT capability integrates with LlamaFactory for hybrid CPU-GPU fine-tuning of large MoE models, including LoRA and full-parameter training, with INT4/INT8 quantization options.
1 use taken from transcripts — each links to the moment in the video.
An open-source framework for inference and fine-tuning of very large language models using CPU-GPU heterogeneous computing. It runs mixture-of-experts models by keeping hot experts on the GPU and the rest on optimized CPU kernels.
1 in the library.