Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
The leading commercial LLM outperforms open-weight models by a significant margin, revealing stark disparities in factual knowledge across languages.
Subword composition methods can drastically reduce continued pre-training steps while enhancing accuracy in LLM vocabulary extensions.
A new gold-standard Marathi POS tagging dataset reveals that even under-resourced languages can achieve high accuracy with the right tools and methodologies.
IndicGuard outperforms existing models by significantly enhancing LLM safety and robustness in culturally sensitive contexts across ten Indic languages.