From-scratch Rust+CUDA inference engine, bit-exact by construction — NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).
Derived from this repo's Cargo dependencies and activity — difficulty is an estimate from codebase size and scope.
Matched by dependency overlap in Cargo manifests — each card notes the most distinctive crates both projects share.
From-scratch Rust+CUDA inference engine, bit-exact by construction — NVFP4, MoE, MTP speculative decoding, tuned against measured limits of one RTX 5090 Laptop (sm_120a).
Camelid: a Rust-native local inference backend with evidence-gated model compatibility.
MLX-based experimental inference engine
A blazing fast inference solution for text embeddings models
Rust library for concurrent data access, using memory-mapped files, zero-copy deserialization, and wait-free synchronization.
Asynchronous Virtio socket support for Rust
Pure Rust Inference Engine
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2