Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability Paper • 2610.08448 • Published 5 days ago • 203
WorldAuditBench: Interactive 3D World Auditing with Multimodal Agents Paper • 2609.40325 • Published 11 days ago • 106
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 12 days ago • 553
TokenRouter: Efficient Serving System for Token-Level LLM Routing Paper • 2610.12242 • Published 3 days ago • 130
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight Paper • 2610.08077 • Published 5 days ago • 156
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 14 days ago • 675
Learning from Teacher Continuations at Student States Paper • 2609.36246 • Published 13 days ago • 41
Agent Plasticity: Measuring Self-Improvement Through Experience Paper • 2610.08902 • Published 5 days ago • 5
Does Learning Protein Folding Generalize to Broader Reasoning? Paper • 2609.38879 • Published 11 days ago • 62
MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement Paper • 2610.11959 • Published 3 days ago • 72
CoTrace: Data Recipes for Training Terminal Agents with Harness-Model Co-Evolution Paper • 2610.10426 • Published 4 days ago • 9
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent Paper • 2508.06600 • Published Aug 8, 2025 • 45
nanoMuse: An Open-Source Personal Agent for Every Device You Own Paper • 2610.08699 • Published 5 days ago • 101
Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps Paper • 2610.02740 • Published 9 days ago • 1
Recursive Game Creator: An Agentic Product-Level Experience-Oriented Game Harness Paper • 2610.08621 • Published 5 days ago • 87
QiYao-M: Multimodal Time Series Foundation Model with Role-Aware Modeling of Endogenous and Exogenous Modalities Paper • 2609.34842 • Published 13 days ago • 5
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization Paper • 2609.38169 • Published 12 days ago • 118