CIBench Explorer
The entire experiment-learning surface in one place. Browse the 17×7 capability matrix, run any model on any benchmark, explore paper-grade findings, annotate failure modes, see routing decisions, drift, and the full retraction ledger.
⚠️ ANTHROPIC_API_KEY not set — paid (claude-*) manifests will error.
17-model × NoLiMa 4-bin matrix
Each cell links to a stored result_id. Click any row below to load its F1-F8 annotation in the panel beneath the table — no tab-switch required.
Empty cells are intentional, not missing data: cohort carve-outs (Grok 4.20 + gpt-oss-120b held to MRCR-8n only per paper §5) plus the LB-v2 HF dataset drift documented 2026-04-27 (docs/MODEL-CONFIG.md §4 LB-v2 drift). MRCR-SW (524K-1M) was dropped from the figure on 2026-04-27 and now lives in the paper §5.1 narrative.
MODEL-CONFIG.md §0 — NoLiMa table
model | bin | headline | result_id |
|---|---|---|---|
qwen3.5-122b-a10b-thinking-on | NoLiMa 128K | 0.44 | a81eca59 |
Selected row — F1-F8 annotation
Click a row above to populate this panel. Uses cibench-store-experiments as the default store.
Run any model on any benchmark
Pick a registered model + a benchmark bin. The result lands in a temp store; the Space hard-caps cost at $1.00/click.
Paper-grade findings backed by stored data
Each finding cites stored result_ids. Click to expand.
Qwen3.5-122B-A10B at NoLiMa 32K scores 0.344 with thinking-OFF (vendor-config artifact) and 0.906 with thinking-ON — a 56-point gap entirely attributable to inference-time reasoning, not the underlying model weights.
| config | headline | n_correct | result_id |
|---|---|---|---|
| thinking-OFF (Phase AA) | 0.344 | 11/32 | 985c3a47 |
| thinking-ON (Phase AB) | 0.906 | 29/32 | c1efdfaa |
| Δ | 0.562 | +18 | (controlled) |
docs/MODEL-CONFIG.md §0 + §2 Qwen3.5-122B-A10B section
The 4.7 retrain trades MRCR-8n recall (0.947 → 0.727) for NoLiMa inferential gains (64K: 0.7 → 0.9; 128K: 0.4 → 0.7) — same vendor, same family, same month, opposite direction.
| axis | Opus 4.6 | Opus 4.7 | Δ |
|---|---|---|---|
| MRCR-8n (131K-262K) | 0.947 | 0.727 | -0.22 |
| NoLiMa 64K | 0.7 | 0.9 | 0.2 |
| NoLiMa 128K | 0.4 | 0.7 | 0.3 |
paper/cibench-neurips-2026.tex §5.1 + §5.2 + Abstract finding (3)
GPT-5.4 trades MRCR-8n (0.654 → 0.564) for NoLiMa 64K (0.6 → 0.7) — the same shape as Anthropic's flip. Three-vendor pattern, not a vendor-specific quirk.
| axis | GPT-5.2 | GPT-5.4 | Δ |
|---|---|---|---|
| MRCR-8n | 0.654 | 0.564 | -0.09 |
| NoLiMa 64K | 0.6 | 0.7 | 0.1 |
docs/MODEL-CONFIG.md §6 OpenAI GPT-5.2 / 5.4
K2.6 trades MRCR-8n (0.905 → 0.798) for NoLiMa 32K/64K gains (0.47/0.50 → 0.56/0.66). Three-vendor convergent direction flip across Anthropic, OpenAI, and Moonshot Q1→Q2 2026 retrains.
| axis | K2.5 | K2.6 | Δ |
|---|---|---|---|
| MRCR-8n | 0.905 | 0.798 | -0.107 |
| NoLiMa 32K | 0.47 | 0.56 | 0.09 |
| NoLiMa 64K | 0.5 | 0.66 | 0.16 |
docs/MODEL-CONFIG.md §1 + §1.5 (K2.6, K2.5)
Opus 4.7 wide MRCR profile shows 35.7% refusal-rate (10/28 items) at 524K-char haystacks — Anthropic's safety classifier declined to engage with the synthetic filler. Any number below the refusal wall is a floor, not a capability ceiling. F6 fires at ≥25%.
| profile | refusal_rate | reliability | F6 fires |
|---|---|---|---|
| Opus 4.7 wide | 35.7% | artifact-contaminated | True |
| Threshold (per F6 wiring MR !111) | ≥25% | — | True |
MODEL-MEMORY-ONTOLOGY.md §4.0.1; F6 implementation in src/cibench/observatory/failure_annotator.py
OpenAI announced MRCR v2 jumped 36.6% → 74.0% (+37pp) from GPT-5.4 to GPT-5.5. The +37pp magnitude REPLICATES on NoLiMa: +27pp at 64K and +47pp at 128K. Cross-vendor benchmark validation.
| bin | GPT-5.4 | GPT-5.5 | result_id |
|---|---|---|---|
| NoLiMa 8K | — | 0.9 | 4334c199 |
| NoLiMa 32K | 1.0 | 0.969 | 11fac703 |
| NoLiMa 64K | 0.7 | 0.969 | dd7714a7 |
| NoLiMa 128K | 0.4 | 0.875 | 577fa00f |
MR !116; docs/MODEL-CONFIG.md §6 GPT-5.5
V4-Flash (284B/13B-active MoE, $0.14/$0.28 per 1M tokens) measures +10pp at 8K, +44pp at 32K, +38pp at 64K, +50pp at 128K vs V3.2. The cheapest open-weight model in the cohort is now competitive with frontier closed-source on long-context inferential retrieval.
| bin | V3.2 | V4-Flash | Δ | result_id |
|---|---|---|---|---|
| NoLiMa 8K | 0.8 | 0.9 | 0.1 | 6cf71098 |
| NoLiMa 32K | 0.44 | 0.875 | 0.44 | 97ddf9bc |
| NoLiMa 64K | 0.53 | 0.906 | 0.38 | 9192685b |
| NoLiMa 128K | 0.25 | 0.75 | 0.5 | 0fc96787 |
docs/MODEL-CONFIG.md §4 DeepSeek V4 Pro/Flash (added 2026-04-26)
Annotate any stored result with F1-F8
Paste a result_id (8-char prefix or full UUID) from any store, click Inspect.
ProcurementRouter
Pick a task type + context length. Returns the routing recommendation plus rationale + confidence + fallback.
F7 cross-store drift detection + retraction ledger
Drift = same manifest_hash, different headlines across runs.
Drift pairs
manifest_hash | result_ids | headline_range | threshold |
|---|---|---|---|
Retraction ledger
Source: configs/retractions/v0/retractions.yaml
Retractions
retraction_id | severity | claim | correction |
|---|---|---|---|
schema-evolution-broke-position-bias-manifest-hash-stability | artifact-contamination | Thinking-on v1 effective_tokens values for 4+ reasoning-default
models (GLM-5, Q | Thinking-on v2 measurements with HF_MAX_OUTPUT_TOKENS=8192 produced
floors 3-4x higher on the same w |
Stateless replay
Pick a pre-registered public manifest. Run it. The result and its replay command are the same artifact.