Open, 🍎-to-🍎 ML benchmarks
Every model scored on the same data, the same metric, in a reproducible sandbox —
across LLM, vision, and audio. Browse the standings here; run your model and submit free on BenchHub.
🎯Same data & metric
Every model on a board runs the identical eval set and scorer.
📌Pinned protocol
LLM boards bake one exact prompt — zero-shot, exact-match, no per-model tuning.
📦Sandboxed, not self-reported
Predictions are scored server-side in a container — reproducible, not paper numbers.