Search papers, labs, and topics across Lattice.
10
3
8
3
Reranking documents by language coherence can dramatically improve retrieval performance, as shown by LAMAR's superior results across multiple languages.
DART boosts reasoning accuracy by up to 22.5 points while slashing thinking token usage by over 50%, all without requiring labeled training data.
ACOER reduces token generation by over 60% while boosting accuracy, solving the reward collapse problem that plagues traditional efficiency training methods.
A simple rescaling of the MLM-head can turn unstable training runs into competitive sparse retrieval models, challenging the notion that bigger encoders alone drive performance.
SHIFT effectively eliminates language bias in multilingual information retrieval, enhancing access to semantically relevant documents across diverse languages.
Unlock high-performance sparse retrieval in any language: SemBridge's smart initialization closes the cross-lingual gap without sacrificing precision.
Synthetically corrupting data with a taxonomy of OCR errors lets you train LLMs to fix real-world OCR mistakes and dramatically improve document understanding.
Multilingual retrievers often prioritize irrelevant English documents over relevant foreign-language documents, even when the query is in that foreign language.
Low-resource languages can get a 15% boost in cross-lingual retrieval accuracy by using English as a Rosetta Stone during training.
Forget just mining hard negatives: the secret to better knowledge distillation for retrieval lies in matching the *entire* score distribution of your teacher model.