Search papers, labs, and topics across Lattice.
The paper introduces Perception Programs (P$^2$), a training-free method to convert dense, pixel-level vision tool outputs into compact, structured, language-native summaries for multimodal language models (MLLMs). This approach addresses the misalignment between raw tool outputs and LLM reasoning capabilities, improving the utilization of visual cues. Experiments across six perception-centric tasks demonstrate that P$^2$ significantly boosts MLLM performance, achieving up to a 45% accuracy increase with GPT-5 Mini and substantial gains even on smaller models, surpassing existing tool-use methods without requiring training.
LLMs struggle with raw visual tool outputs, but a simple trick鈥攖ranslating pixels into language-native summaries鈥攗nlocks massive gains in visual reasoning, outperforming even fine-tuned agents.
Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, despite access to these tool-generated visual cues, MLLMs often fail to benefit from them. Existing approaches typically feed raw tool outputs into the model, but these dense, pixel-level representations are misaligned with the language-native reasoning strengths of LLMs, leading to weak perception and reliance on language priors. We argue that, in problems where vision tools can provide the necessary visual cues, the bottleneck is not more tool calls or larger MLLMs, it is how tool outputs are represented. We introduce Perception Programs (P$^2$), a training-free, model-agnostic method that rewrites tool outputs into compact, structured, language-native summaries that MLLMs can directly parse and reason over. Across six perception-centric tasks in BLINK, P$^2$ consistently yields large improvements over base models and raw tool-augmented baselines. With GPT-5 Mini as the base model, P$^2$ raises its accuracy from 41.35\% to 86.47\% on multi-view reasoning, from 52.42\% to 81.45\% on relative depth, and achieves a 22\% average gain across tasks, setting new state-of-the-art results. Even on smaller MLLMs, e.g., InternVL3.5-4B and Qwen3VL-4B, we observe 15-40\% absolute gains from P$^2$, surpassing prior agentic, supervised, and RL-based tool-use methods-without any training or model modifications.