SpatialOPSD: Self-Distilling Spatial Intelligence from Verified Coding Agent Traces
Abstract
Spatial coding agents significantly improve spatial reasoning in Multimodal Large Language Models (MLLMs) by using external tools to generate verified execution traces. However, this paradigm inherently suffers from prohibitive inference-time overhead and external dependencies. In this paper, we explore whether an MLLM can internalize this agentic capability to operate entirely tool-free. We begin with a simple observation: prompting an MLLM with summarized execution traces of a spatial coding agent naturally unlocks the model's internal spatial Chain-of-Thought (CoT). Motivated by this, we introduce SpatialOPSD, an on-policy self-distillation framework that internalizes spatial reasoning into a standalone MLLM by formulating verified agent traces as privileged information. To mitigate privileged-information leakage during distillation, we introduce Repetition-Aware Distillation, which combines repetition masking with unlikelihood regularization. Experiments across multiple benchmarks demonstrate that self-distilling SpatialOPSD achieves higher average accuracy than SFT and GRPO on both spatial and OOD datasets, exhibiting superior performance and generalization.
Community
With the widespread adoption of coding agents, we explore an approach to distilling the spatial reasoning capabilities of coding agents into standalone MLLMs through on-policy self-distillation (OPSD).
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RAISED: Self-Distillation for Robustness to Prompt Injection in LLM Agents (2026)
- PSD: Pseudo Self-Distillation of Memory Representation Capabilities for LLM Agents (2026)
- Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation (2026)
- Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models (2026)
- Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning (2026)
- MemoryWalker: Stop Training Agents on Contexts They Never Saw (2026)
- SciLENS: RL-Driven Autonomous Agents for Scientific Localized Evidence Navigation and Synthesis (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper