Search papers, labs, and topics across Lattice.
Affiliation:
2
0
4
Get up to 20% faster ViT inference by hot-swapping certain attention heads for depthwise convolutions – without tanking accuracy.