Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Derived from this repo's Cargo dependencies and activity — difficulty is an estimate from codebase size and scope.
Matched by dependency overlap in Cargo manifests — each card notes the most distinctive crates both projects share.
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
Pure Rust + CUDA LLM inference engine
The official Rust library for the Edgee AI Gateway
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Mock REST APIs from JSON with zero coding within seconds.
An adaptive-radix-tree metadata storage engine
Asynchronous Virtio socket support for Rust
Fluxon is a distributed data acceleration plane built in Rust for AI-native systems. It features three standardized interfaces: KV/RPC for inference KV Cache reuse, MQ for state passing in training and data pipelines, and FS for caching and S3-compatible access to model files and checkpoints.
Rust based high-performance Apache Uniffle shuffle-server
Rust client for Qdrant vector search engine