Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
Video diffusion models notoriously mangle hands and facial micro-expressions, but spatially grounded local adapters conditioned on 2D keypoints can recover the fine articulation required for intelligible sign language, outperforming existing baselines by 30% in pose precision.
SignSeek achieves zero-shot generalization to British Sign Language, outperforming models trained specifically on BSL data.
BLEU-4 may mislead SLT progress, masking the true performance differences between gloss-free and gloss-supervised systems.