benchmarks
7 projects
TrueForge is TrueFoundry's open agent harness, priced against Claude
truefoundry/trueforge
An MIT-licensed TypeScript runtime for agents: model calls, MCP tools, skills, sandbox and approvals. Its cost benchmark is 14 tasks, run by the vendor.
FreeToken streams MoE experts to a gaming GPU on demand
FlashML-org/FreeToken
A Python and CUDA engine that keeps hot experts on the GPU, the rest in RAM, and splits misses between PCIe and the CPU. One outside test: 2.25x Ollama.
turbovec fits 10M vectors in 4 GB and outruns FAISS
RyanCodrai/turbovec
A Rust index built on Google Research's TurboQuant: no training step, hand-written NEON and AVX-512 kernels, crash-safe incremental saves and filtered search.
Hivemind mines your team's traces into shared skills
activeloopai/hivemind
Sessions become searchable traces, repeated patterns become SKILL.md files, and every agent on the team inherits them. Benchmarked 25% cheaper on LoCoMo.
T3MP3ST turns your coding agent into a red team
elder-plinius/T3MP3ST
An offensive-security harness driving the agent you already run through recon, exploit and report, with every README number recomputable by one command.
Hindsight is agent memory that tries to learn, not recall
vectorize-io/hindsight
A memory system claiming state of the art on LongMemEval, with the rare detail that outside researchers reproduced the score instead of the vendor asserting it.
Ponytail makes your coding agent stop over-building
DietrichGebert/ponytail
An agent skill that enforces YAGNI and reaches for native platform features first. Its own benchmark walked back its earlier, more flattering numbers.