SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 2 days ago • 99
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 6 days ago • 273
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published 13 days ago • 106
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 19 days ago • 259
The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI Agents Paper • 2608.06065 • Published 14 days ago • 8
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published 21 days ago • 72
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published 14 days ago • 46
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 15 days ago • 23
MiniWorld: Democratizing the Training of Video World Models from Scratch Paper • 2608.01127 • Published 18 days ago • 19
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published 20 days ago • 24