Back to list
High-Potential
Rust

⚡ Pegainfer: Pure Rust LLM Inference Engine

670 stars103 forksRust
cudacuda-kernelsdeepseekgpuinferenceinference-enginekimikimi-k2kv-cachellmllm-inferencellm-serving
A pure Rust and CUDA LLM inference engine that drops the PyTorch dependency entirely. It provides an OpenAI-compatible API and supports serving models ranging from Qwen3 to Kimi-K2. The interesting part is the lightweight, ground-up rewrite. For deployment environments where memory overhead and performance are critical, or where avoiding a massive Python dependency tree is preferred, this compiled-language approach offers a clean alternative.