Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Weak supervisors no longer bottleneck stronger models when their guidance is used solely to accelerate verifier-aligned policy gradients rather than dictate optimization targets.
Sliding Window Attention outperforms Linear Attention by 2 to 10 times on long-context reasoning tasks while being faster and requiring less memory.