Search papers, labs, and topics across Lattice.
1
0
3
18
Cross-lingual word-to-speech mappings can be effectively learned from visual grounding without the need for transcriptions or extensive model training.