Search papers, labs, and topics across Lattice.
3
0
3
Cultural cues can mislead language models into choosing the wrong answers even when they select the correct normative framework, exposing a significant knowledge gap in AI systems.
EDRAC reveals that existing Arabic LLMs struggle with dialectal fidelity, exposing a critical gap in current benchmarks for Arabic NLP.
A new semantic correctness framework reveals that traditional evaluation metrics often overlook critical distinctions in answer quality, leading to misleading assessments of LLM performance.