SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
Archive
Realistic first-contribution target: active project, 113 open issues, small enough for a newcomer PR to land.
Derived from this repo's Cargo dependencies and activity — difficulty is an estimate from codebase size and scope.
Matched by dependency overlap in Cargo manifests — each card notes the most distinctive crates both projects share.
SSD-streaming inference engine for giant MoE models (Rust + CUDA). GLM 5.2 743B at 2 tok/s and Hy3 295B at 7 tok/s on two consumer 16GB GPUs. Zero-config multi-GPU: measures PCIe bandwidth, places attention and hot experts where they fit.
mDNS, SSDP/DIAL, WSD, WoL reflector
Rust crate package to link to a system libz (zlib)
RDNA-native LLM inference engine in Rust.
RDNA-native LLM inference engine in Rust.
Camelid: a Rust-native local inference backend with evidence-gated model compatibility.
Fast, flexible LLM inference