Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
2
Despite executing target instructions, LLMs often fail to deliver competitive performance in GPU kernel optimization, especially on complex tasks.
Bridging the gap between inference and adaptation in VLMs could lead to significant performance boosts by ensuring robust pseudo-labels that accurately reflect sample-level relationships.
Parallax achieves a Pareto improvement in LLM pretraining by replacing standard attention with a novel parameterized local linear attention mechanism, demonstrating that architectural innovation in attention can still yield significant gains.