Search papers, labs, and topics across Lattice.
3
0
5
Task learnability can significantly enhance RL efficiency in LLMs, leading to better performance with less data.
LLMs struggle with compositional reasoning, showing sharp performance drops in executing order-sensitive data refinement tasks.
Directional inconsistency can destabilize LLM training, but geoalign curates rollouts to improve performance and stability, outperforming several established methods.