dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.
arrow-select
Counted from the Cargo.toml manifests of the 32 indexed repositories that declare arrow-select as a dependency — not download counts. Dependency data last verified 2026-08-13.
Crates that show up unusually often in arrow-select projects. The most distinctive pairings rank first — crates these projects use far more than the average indexed Rust project does, not just crates that are popular everywhere. Each percentage is the share of arrow-select projects that also use it.
A data visualization and analytics component, especially well-suited for large and/or streaming datasets.
Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
Simple, Elastic-quality search for Postgres
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Official Rust implementation of Apache Arrow
A native Rust library for Delta Lake, with bindings into Python
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
Parseable is an observability datalake built from first principles.
Tonbo is an embedded database for serverless and edge runtimes.
Apache Iceberg
Lakehouse native graph engine with git-style workflows
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Intuitive Data Workflows
Scalable graph analytics database powered by a multithreaded, vectorized temporal engine, written in Rust
Protocol and libraries for sending and receiving OpenTelemetry data using Apache Arrow
SciRS2 - Scientific Computing and AI in Rust., providing SciPy-compatible APIs while leveraging Rust's performance, safety, and concurrency features. Unlike traditional scientific libraries
The native Rust implementation for Apache Hudi, with C++ & Python API bindings.
A typed, polyglot, functional language
Apache Paimon Rust The rust implementation of Apache Paimon.
Embeddable spreadsheet engine — parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 320+ functions.
We're back! Now firing notebooks out of a t-shirt gun.
DuckLake took Flight. Welcome to SwanLake.
On-device property graph database. Schema-as-code. One CLI → One Folder. No Server. Think: DuckDB for graphs.
Library for bringing distributed capabilities to Apache DataFusion
Graph database native to the cloud. Embedded, multi-tenant, built on object storage.
Open-source streaming SQL engine written in Rust using Apache Arrow and DataFusion. Supports continuous queries, temporal stream joins, tumbling/session windows, and CDC/Kafka connectors. Lightweight, embeddable, and sub-microsecond latency
Lossless storage and search for AI agent sessions, across every agentic client.
A framework to manage data, continuously