Search papers, labs, and topics across Lattice.
Affiliation:
1
0
1
0
This work introduces URBANCONTRASTIVEQA, a benchmark that asks whether tool-augmented language models can make this baseline-relative comparison between NYC, Chicago, and Seattle, and evaluates six instruction-tuned models under five tool-output formats.