Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
Even top-performing models falter in cultural reasoning, scoring below chance on concepts requiring specific regional knowledge.
MLLMs are failing to recognize and effectively utilize physical tools, with top models achieving only 21% task completion in real-world scenarios.