Search papers, labs, and topics across Lattice.
Affiliation:
17
0
15
15
This work studies adaptive routing of prompts to large language model experts to maximize response quality in an online setting with limited feedback and proposes algorithms that strategically select and observe rewards to minimize regret.
Diverse Skill Routing is proposed, a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy and improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries.
CRISP recovers up to 28% improvement on retrieval tasks while achieving a remarkable 5.30x speedup in attention computation for long-context LLMs.
A unified layer equation for GNNs reveals how over 200 architectures can be systematically compared and optimized for performance.
Personalization in AI co-scientists could be the key to unlocking novel research insights that generic systems overlook.
Early context pruning can slash token usage by 73% without sacrificing quality, reshaping how we design efficient research agents.
A unified toolbox for time-series dataset similarity could revolutionize how researchers select and evaluate datasets for AI model fine-tuning.
Learning to coordinate retrieval signals can significantly boost multi-hop reasoning accuracy in language models.
MultAttnAttrib achieves superior attribution accuracy while cutting inference latency to one-seventh of traditional prompting methods, revolutionizing multimodal evidence tracing in AI.
Output compression can slash inference costs by up to 3x, but input compression leads to higher costs and accuracy collapse鈥攁n unexpected trade-off for LLMs.
Shifting reasoning to the indexing stage can drastically reduce query latency while enhancing retrieval effectiveness through LLM-generated rationales.
Integrating semantic, acoustic, and engagement signals in music recommendations can boost performance by nearly 95%, challenging the status quo of opaque token-based systems.
Texture, not color, is the secret sauce behind fashion house identity, revealed by probing a multimodal CNN trained on decades of Vogue runway images.
Multimodal agents can now plan more coherently and solve complex tasks thanks to a new anticipatory reasoning framework that forecasts short-horizon trajectories before acting.
Agentic RAG systems can be made significantly more efficient and accurate simply by adding a contextualization module and de-duplicating retrieved documents at test time.
A unified benchmark reveals the fragmented landscape of RAG security, highlighting vulnerabilities to knowledge-extraction attacks and paving the way for robust defense strategies.
LLM judges exhibit a surprising "blindness" to human-written summaries, increasingly preferring machine-generated content as the similarity to human references decreases, challenging their reliability in summarization tasks.