Search papers, labs, and topics across Lattice.
To resolve the trade-off between local detail and continuous scene context in ultra-high-resolution visual reasoning, the authors introduce GazeEarth, a training-free framework that redistributes a fixed pixel budget using question-guided foveated observation. While frozen MLLMs can identify relevant spatial coordinates, detached cropping strips peripheral context and degrades reasoning; GazeEarth instead applies a deterministic, topology-preserving warp that magnifies target regions while retaining a compressed global view. Across four frozen backbones on three remote sensing benchmarks, the approach boosts accuracy by 4.6 to 9.4 percentage points using at most two inference passes, surpassing task-trained models without fine-tuning.
Disjointed cropping destroys visual context鈥攔edistributing a fixed pixel budget through continuous foveation around an MLLM's self-selected focus points beats task-trained baselines without any fine-tuning.
Multimodal large language models (MLLMs) must balance local detail against scene context when interpreting ultra-high-resolution (UHR) remote sensing (RS) imagery within a limited visual-input budget. Existing selection-based methods either prune tokens and select patches through relevance scoring, or crop actively through repeated inspection. Neither strategy directly redistributes pixels within a continuous full-scene view: the first retains selected tokens or patches, and the second re-encodes a crop detached from its surroundings. Our pilot study finds that a frozen MLLM already produces useful question-guided spatial requests, yet crop-based inspection of the selected regions does not consistently improve its answers. We therefore formulate UHR understanding as a question of where to spend a fixed pixel budget. Based on this, we introduce GazeEarth, a simple-yet-effective training-free framework that couples question-guided region selection with full-scene foveated observation. The MLLM selects evidence cells from an indexed overview; a deterministic, topology-preserving warp resamples the original image onto a fixed-size canvas, enlarging their shared neighborhood while compressing the periphery; the same frozen model answers from this focused view, using at most two MLLM calls and no external selector or iterative search. Across three UHR remote sensing benchmarks and four frozen backbones, GazeEarth improves benchmark-averaged accuracy by 4.6 to 9.4 percentage points over direct answering and 3.4 to 4.3 over overview answering, outperforming task-trained methods. Our analyses show that existing MLLMs can guide where to look in UHR images on their own, and that what they can infer from the selected evidence depends on how that evidence is presented.