Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published 14 days ago • 10
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces Paper • 2601.11868 • Published Jan 17 • 37
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources Paper • 2509.25531 • Published Sep 29, 2025 • 11
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL Paper • 2607.04412 • Published Jul 5 • 35
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning Paper • 2605.17077 • Published May 16 • 1
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Paper • 2604.24954 • Published Apr 27 • 26
Privasis: Synthesizing the Largest "Public" Private Dataset from Scratch Paper • 2602.03183 • Published Feb 3 • 13
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text Paper • 2601.22975 • Published Jan 30 • 113
mlfoundations-dev/subsampled_flan_v2_w_system_instructions Viewer • Updated Oct 14, 2024 • 3.56M • 61
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning Paper • 2503.19263 • Published Mar 25, 2025 • 2
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems Paper • 2504.09763 • Published Apr 14, 2025 • 12