VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks Paper • 2610.00972 • Published 10 days ago • 61
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Paper • 2609.39071 • Published 11 days ago • 65
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models Paper • 2610.03665 • Published 9 days ago • 62
Native Action-Prior Learning from Videos for World Action Models Paper • 2610.03391 • Published 9 days ago • 88
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 10 days ago • 279
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 15 days ago • 326 • 5
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 12 days ago • 104
What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 12 days ago • 138
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 15 days ago • 326