A lightning-fast search engine API bringing AI-powered hybrid search to your sites and applications.
tiktoken-rs
Counted from the Cargo.toml manifests of the 87 indexed repositories that declare tiktoken-rs as a dependency — not download counts. Dependency data last verified 2026-08-13.
Crates that show up unusually often in tiktoken-rs projects. The most distinctive pairings rank first — crates these projects use far more than the average indexed Rust project does, not just crates that are popular everywhere. Each percentage is the share of tiktoken-rs projects that also use it.
an open source, extensible AI agent that goes beyond code suggestions - install, execute, edit, and test with any LLM
A CLI tool to convert your codebase into a single LLM prompt with source tree, prompt templating, and token counting.
Plano is an AI-native proxy and data plane for agentic apps — with built-in orchestration, safety, observability, and smart LLM routing so you stay focused on your agents core logic.
Desktop GUI Framework
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Next Generation Agentic Proxy for AI Agents and MCP servers
LeanCTX — the Context OS for AI development. One local binary that compresses, remembers, routes, and verifies every token between your code and the model. 63 MCP tools, 10 read modes, up to 99% token savings. Works with Cursor, Claude Code, Copilot, Windsurf, Codex, Gemini.
A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.
Macro is a unified interface for email, messages, tasks, calls, agents, pull requests, docs, crm — linked together with shared AI memory.
Low latency web data collector
A fast Rust based tool to serialize text-based files in a repository or directory for LLM consumption
Bionic is an on-premise replacement for ChatGPT, offering the advantages of Generative AI while maintaining strict data confidentiality
EdegQuake 🌋 High-performance GraphRAG inspired from LightRag written in Rust; Transform documents into intelligent knowledge graphs for superior retrieval and generation
Self-hosted, semantically-connected personal knowledge base
Markdown memory system for you and your AI agent
High Performance LLM Reverse Proxy
Agentic review of Linux Kernel code changes
Flow-Like: Strongly Typed Enterprise Scale Workflows. Built for scalability, speed, seamless AI integration and rich customization.
The self-improving all channels AI agent. Self-healing. Fully autonomous. Single Rust binary.
Fast, streaming indexing, query, and agentic LLM applications in Rust
A multi-agent framework written in Rust that enables you to build, deploy, and coordinate multiple intelligent agents
Own your PaaS
Split text into semantic chunks, up to a desired chunk size. Supports calculating length by characters and tokens, and is callable from Rust and Python.
Auditable workspaces for AI agent teams: sandboxed worktrees, multi-agent peer-review, 95% lower token waste, and persistent memory.
Orca is a DeepSeek-native coding agent.
Orca is a DeepSeek-native coding agent CLI by Blade.
Empower the Shell to think. Evolve Operations.
your terminal, but it talks back
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
Models and examples built with Burn
A comprehensive Rust translation of the code from Sebastian Raschka's Build an LLM from Scratch book.
Local desktop gateway that translates OpenAI Codex CLI's Responses API into Chat Completions for Kimi / DeepSeek / Zhipu GLM / Bailian and other OpenAI-compatible providers.
MoFA - Modular Framework for Agents. Modular, Compositional and Programmable.
Democratizing large model inference and training on any device.
Blockcell is a self‑evolving agent
AURA is an agentic harness that turns an LLM model into a reliable, autonomous service capable of executing real SRE work. AURA provides the guardrails, API servers, state management, authentication, streaming, error handling, and tool integrations necessary to run AI SRE agents safely in production.
Suite of tools containing an in-memory vector datastore and AI proxy
🤖 A Matrix bot for using different capabilities (text-generation, text-to-speech, speech-to-text, image-generation, etc.) of AI / Large Language Models (OpenAI, Anthropic, etc.)
Repo-native governance kernel that enriches context and turns natural-language intent into governed, proof-backed work across looping agent fleets. 🦀
Fast, cross-platform, real-time token usage tracker and cost monitor for Claude Code / Codex CLI / Antigravity CLI / Qwen Code / Cline / Zoo Code / Kilo Code / GitHub Copilot / OpenCode / Pi Agent / Piebald.
gproxy is a Rust-based multi-channel LLM proxy that exposes OpenAI / Claude / Gemini-style APIs through a unified gateway, with a built-in admin console, user/key management, and request/usage auditing.
Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and code before they're sent: -31% input / -74% output, measured live. Any provider, no extra model calls. Also an MCP server and embeddable library (Rust, Python, Ruby, Kotlin, Swift, JS/TS).
Pure Rust PDF library for AI/RAG: structure-aware chunking, no ML, no C deps.
A simple CLI for working with any agent
Graph-native code intelligence that replaces embedding-based RAG with deterministic program understanding.
Graph-native code intelligence that replaces embedding-based RAG with deterministic program understanding.
CI-native agent CLI tool for deterministic pipeline gating.
Tandem is the authority layer for AI-first work: runtime authority for agents, tools, memory, approvals, and audit trails.