CIBench Explorer

The entire experiment-learning surface in one place. Browse the 17×7 capability matrix, run any model on any benchmark, explore paper-grade findings, annotate failure modes, see routing decisions, drift, and the full retraction ledger.

⚠️ ANTHROPIC_API_KEY not set — paid (claude-*) manifests will error.

17-model × NoLiMa 4-bin matrix

Each cell links to a stored result_id. Click any row below to load its F1-F8 annotation in the panel beneath the table — no tab-switch required.

Empty cells are intentional, not missing data: cohort carve-outs (Grok 4.20 + gpt-oss-120b held to MRCR-8n only per paper §5) plus the LB-v2 HF dataset drift documented 2026-04-27 (docs/MODEL-CONFIG.md §4 LB-v2 drift). MRCR-SW (524K-1M) was dropped from the figure on 2026-04-27 and now lives in the paper §5.1 narrative.

MODEL-CONFIG.md §0 — NoLiMa table

MODEL-CONFIG.md §0 — NoLiMa table
model
bin
headline
result_id
qwen3.5-122b-a10b-thinking-on
NoLiMa 128K
0.44
a81eca59

Selected row — F1-F8 annotation

Click a row above to populate this panel. Uses cibench-store-experiments as the default store.