Search papers, labs, and topics across Lattice.
Affiliation:
2
0
3
Video diffusion models notoriously mangle hands and facial micro-expressions, but spatially grounded local adapters conditioned on 2D keypoints can recover the fine articulation required for intelligible sign language, outperforming existing baselines by 30% in pose precision.
BLEU-4 may mislead SLT progress, masking the true performance differences between gloss-free and gloss-supervised systems.