SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Derived from this repo's Cargo dependencies and activity — difficulty is an estimate from codebase size and scope.
Matched by dependency overlap in Cargo manifests — each card notes the most distinctive crates both projects share.
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Open-source Rust/Pingora AI gateway and reverse proxy for self-hosted LLM traffic control: OpenAI-compatible Chat/Responses, provider routing/fallback, virtual API keys, policy and budgets, exact-match cache, MCP/tool execution, observability, Admin APIs/dashboard, cluster ops, and automatic HTTPS.
Private AI Gateway for Attested Confidential Inference
A multiplexed p2p network framework that supports custom protocols
A high-performance Rust implementation of the mihomo (Clash Meta) proxy kernel.
A Rust library for random number generation.
Self-hosted, censorship-resistant VPN with REALITY / anti-DPI obfuscation and post-quantum (X25519 + ML-KEM-768) crypto. Rust core + native Android/Windows/macOS clients. Works on networks with active DPI (Iran, China, Russia).
Async-friendly WebTransport implementation in Rust
Peer-to-peer networking library