D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation Paper • 2608.24987 • Published 5 days ago • 25
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning Paper • 2608.23318 • Published 6 days ago • 24
Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher Paper • 2608.26872 • Published 3 days ago • 66
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 6 days ago • 200
Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation Paper • 2608.15062 • Published 4 days ago • 9
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning Paper • 2608.13914 • Published 16 days ago • 3
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published 9 days ago • 44
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving Paper • 2608.19758 • Published 9 days ago • 20
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published 18 days ago • 42
ASI-Bench: At the Dawn of Artificial Superintelligence Paper • 2608.17271 • Published 12 days ago • 62
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 13 days ago • 150
MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling Paper • 2608.14783 • Published 16 days ago • 19