Skip to content
#

open-benchmark

Here are 10 public repositories matching this topic...

Open, independent benchmark on document parsing APIs and frontier VLMs for AI agents and RAG: does a struck-out clause in a redlined contract PDF survive the parse? Open source code + open data. LlamaParse, Reducto, Datalab, Extend, Pulse, Mistral OCR, GPT-6 Astra, Claude Fable 5.1.

  • Updated Sep 26, 2026
  • Python

Open, independent benchmark on web search APIs for AI agents - factual lookup on 300 company news questions, ranked by cost per 1,000 correct answers. TinyFish, Parallel, Perplexity, Linkup, Firecrawl, Brave, You, Exa, Tavily, Google SERP.

  • Updated Sep 16, 2026
  • Python

Open, independent benchmark on LLM inference providers: same GLM 5.3 Flash, 600 paired requests each. Latency, tokens per second and task success. Baseten, DeepInfra, Fireworks AI, Modal, Nebius, Novita AI, Parasail, Telnyx, Together AI, Z.AI. Python runner and public data.

  • Updated Sep 5, 2026
  • Python

Open, independent benchmark of web search APIs for coding agents: Exa, Parallel, Perplexity, Firecrawl, Tavily, Linkup, Brave, You, TinyFish. Scored on grounded task completion against held-out enterprise docs tickets. Search-only and search+fetch boards, model held constant.

  • Updated Sep 16, 2026
  • Python

Open, independent benchmark on company news APIs for AI agents - web search APIs vs dedicated news indexes on 300 company news questions, ranked by cost per 1,000 correct answers. Exa, Parallel, Perplexity, Brave, Firecrawl, Linkup, You, TinyFish, Tavily, Google SERP, Seltz, PredictLeads, Autobound, Datahyena.

  • Updated Sep 16, 2026
  • Python

Open, independent benchmark on web search APIs for deep research agents - unlike BrowseComp, this is a benchmark on real user workflows that are not memorized by models. Exa, Parallel, Perplexity, Linkup, Tavily, Firecrawl, Brave, You, TinyFish, Seltz, Google SERP.

  • Updated Sep 16, 2026
  • Python

High-resolution 2D lid-driven cavity CFD solver & 30 open benchmark datasets (Re = 100 to 100,000, grids up to 1025x1025). Features O(N^2 log N) DST-I Poisson solver, vectorized ADI, pure central differencing, Batchelor core verification, and research preprint submitted to Computers & Fluids.

  • Updated Oct 6, 2026
  • Python

Add this topic to your repo

To associate your repository with the open-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more