Search papers, labs, and topics across Lattice.
Monash University
3
0
8
6
GRPO fails to improve performance in small language models, revealing a critical scale-dependent limitation in reinforcement learning applications.
Unearthing two distinct planning competencies in LLMs reveals that scaling up models enhances operational reasoning but leaves structural enumeration largely unchanged.
Unlocking the black box of dense retrieval, Xetrieval reveals the interpretable feature-level reasoning that drives relevance scoring.