Spacedrive is an open source cross-platform file explorer, powered by a virtual distributed filesystem written in Rust.
arrow-array
Counted from the Cargo.toml manifests of the 99 indexed repositories that declare arrow-array as a dependency — not download counts. Dependency data last verified 2026-08-13.
Crates that show up unusually often in arrow-array projects. The most distinctive pairings rank first — crates these projects use far more than the average indexed Rust project does, not just crates that are popular everywhere. Each percentage is the share of arrow-array projects that also use it.
Scalable datastore for metrics, events, and real-time analytics
dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.
Incremental engine for long horizon agents 🌟 Star if you like it!
A data visualization and analytics component, especially well-suited for large and/or streaming datasets.
Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.
Simple, Elastic-quality search for Postgres
⚙️🦀 Build modular and scalable LLM Applications in Rust
Sui, a next-generation smart contract platform with high throughput, low latency, and an asset-oriented programming model powered by the Move programming language
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..
The open-source Observability 2.0 database. One engine for metrics, logs, and traces — replacing Prometheus, Loki & ES.
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Distributed stream processing engine in Rust
Apache Iggy: Hyper-Efficient Message Streaming at Laser Speed
Language model tokenization at GB/s
DORA (Dataflow-Oriented Robotic Architecture) is middleware designed to streamline and simplify the creation of AI-based robotic applications. It offers low latency, composable, and distributed dataflow capabilities. Applications are modeled as directed graphs, also referred to as pipelines.
Official Rust implementation of Apache Arrow
Drop-in Apache Spark replacement written in Rust, unifying batch processing, stream processing, and compute-intensive AI workloads.
A native Rust library for Delta Lake, with bindings into Python
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.
Parseable is an observability datalake built from first principles.
An AI agent for teams, communities, and multi-user environments.
Communication infrastructure for the AI era — one binary, one broker, one storage layer, any protocol
A cloud-native open source distributed time series database with high performance, high compression ratio and high availability.
Tonbo is an embedded database for serverless and edge runtimes.
Apache Iceberg
High performance Rust stream processing engine seamlessly integrates AI capabilities, providing powerful real-time data processing and intelligent analysis.
Lakehouse native graph engine with git-style workflows
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Flow-Like: Strongly Typed Enterprise Scale Workflows. Built for scalability, speed, seamless AI integration and rich customization.
AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache
Postgres Foreign Data Wrapper development framework in Rust.
Fast, streaming indexing, query, and agentic LLM applications in Rust
Intuitive Data Workflows
Scalable graph analytics database powered by a multithreaded, vectorized temporal engine, written in Rust
Rust Agent Development Kit (ADK-Rust): Build AI agents in Rust with modular components for models, tools, memory, realtime voice, and more. ADK-Rust is a flexible framework for developing AI agents with simplicity and power. Model-agnostic, deployment-agnostic, optimized for frontier AI models. Includes support for real-time voice agents.
A single-node analytical database engine with geospatial as a first-class citizen
The cost efficiency of S3 with the speed of local RAM. A multi-tenant vector and full-text search engine featuring a tiered RAM -> NVMe -> S3 architecture for microsecond latency on top of object storage
Command line SQL interface for relational databases and common data file formats
GeoArrow in Rust, Python, and JavaScript (WebAssembly) with vectorized geometry operations
Protocol and libraries for sending and receiving OpenTelemetry data using Apache Arrow
Fully Managed, Streaming Ingestion (CDC) into your Lakehouse
Zorai is a persistent, multi-agent, auditable, learning execution platform where the daemon owns work, memory, approvals, tools, and long-running goals.
SciRS2 - Scientific Computing and AI in Rust., providing SciPy-compatible APIs while leveraging Rust's performance, safety, and concurrency features. Unlike traditional scientific libraries
The native Rust implementation for Apache Hudi, with C++ & Python API bindings.
A command-line tool for querying databases