Search papers, labs, and topics across Lattice.
2
0
5
Results show that scaling consistently improves performance, while multitask learning acts as a beneficial regularizer primarily for smaller-capacity models.
Routing tokens through a contrastive lens boosts expert specialization and yields up to 1.77 points improvement in zero-shot reasoning accuracy.