Data
Databases, query engines, pipelines, formats and analytics.
6 projects
turbovec fits 10M vectors in 4 GB and outruns FAISS
RyanCodrai/turbovec
A Rust index built on Google Research's TurboQuant: no training step, hand-written NEON and AVX-512 kernels, crash-safe incremental saves and filtered search.
Hivemind mines your team's traces into shared skills
activeloopai/hivemind
Sessions become searchable traces, repeated patterns become SKILL.md files, and every agent on the team inherits them. Benchmarked 25% cheaper on LoCoMo.
Zvec is a vector database that lives inside your process
alibaba/zvec
Alibaba's embedded C++ engine does dense, sparse, full-text and hybrid search with WAL durability, bringing SQLite's deployment model to retrieval.
Hindsight is agent memory that tries to learn, not recall
vectorize-io/hindsight
A memory system claiming state of the art on LongMemEval, with the rare detail that outside researchers reproduced the score instead of the vendor asserting it.
PixelRAG retrieves over screenshots instead of parsed text
StarTrail-org/PixelRAG
A Berkeley project that renders pages as images and searches them visually, keeping the tables, charts and layout that HTML parsing throws away.
GeoLibre runs 1,000+ GIS tools in the browser, offline
opengeos/GeoLibre
A cloud-native GIS built on Tauri, MapLibre and DuckDB-WASM that ships the WhiteboxTools toolbox to WebAssembly, for desktop, mobile, browser and Jupyter.