Search papers, labs, and topics across Lattice.
University of Oxford
3
0
7
Learning algorithms might excel in memorization but can falter in broader generalization, with RL outperforming SFT in transferring knowledge across contexts.
VLMs can learn to actively reason and plan in 3D environments by distilling view graphs from self-exploration trajectories, enabling them to surpass even larger models like GPT-4 Pro and Gemini 1.5 Pro on interactive view planning.
LLM agents can appear to reason well (high entropy) while completely ignoring the input, and mutual information is a far better metric for catching this failure.