Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
1
This work introduces URBANCONTRASTIVEQA, a benchmark that asks whether tool-augmented language models can make this baseline-relative comparison between NYC, Chicago, and Seattle, and evaluates six instruction-tuned models under five tool-output formats.
Audio language models can grasp the broad gist of stuttered child speech, but their reasoning completely collapses and leaks multi-speaker context as disfluency rates rise.
Current VLMs struggle with page-level comic interpretation, frequently hallucinating objects and demonstrating that semantic similarity metrics are a poor proxy for true comic understanding.
Synthetic data and tailored preprocessing can significantly boost machine translation for indigenous languages, but generic methods fall short for highly agglutinative languages like Aymara.