Search papers, labs, and topics across Lattice.
54 papers published across 6 labs.
TP-MPPO achieves up to 87.5% higher goodput for LLM inference in edge networks, revolutionizing how we manage bandwidth and task offloading.
Mixed SFT outperforms next-chunk reasoning RL in reasoning tasks while using over 60 times less compute, reshaping our understanding of effective training strategies.
Rubric-based scoring can transform how vision-language models ground their responses, leading to substantial improvements in reasoning accuracy.
Free-form language reasoning transforms VLMs into powerful robotic reasoners, significantly boosting performance in complex manipulation tasks.
P4-DT outperformed human surrogates in predicting patient preferences, achieving an impressive 81.7% accuracy by leveraging contextual dilemmas.
Rubric-based scoring can transform how vision-language models ground their responses, leading to substantial improvements in reasoning accuracy.
Free-form language reasoning transforms VLMs into powerful robotic reasoners, significantly boosting performance in complex manipulation tasks.
P4-DT outperformed human surrogates in predicting patient preferences, achieving an impressive 81.7% accuracy by leveraging contextual dilemmas.
Filtering out fit sequences during fine-tuning can boost post-training performance by up to 17%, reshaping how we approach model training for RL applications.
Multi-objective RL methods overlook critical interactions between non-linear utility effects across different timescales, leading to suboptimal decision-making.
Pairwise rewards in reinforcement learning can significantly boost the robustness of LLM auditors, enhancing their ability to detect hidden model behaviors with minimal false positives.
GRIP achieves superior accuracy-efficiency trade-offs in reasoning tasks by intelligently interpolating parameters from two distinct model types without retraining.
TP-MPPO achieves up to 87.5% higher goodput for LLM inference in edge networks, revolutionizing how we manage bandwidth and task offloading.
Confidence estimates from LLMs can be misleading when evaluating many candidates, but a new framework ensures high-probability agreement with human judgments.
AutoVerifier learns from its mistakes, transforming verifier errors into reusable strategies that dramatically boost verification accuracy.
Mixed-policy reinforcement learning can enable language models to absorb knowledge more effectively than traditional supervised fine-tuning, especially in complex reasoning scenarios.
Listwise preference-based reward learning from vision-language models can outperform traditional pairwise methods, achieving up to 86% success in complex robotic tasks.
Self-improving search agents thrive when feedback and policy evolution are intertwined, leading to sustained performance gains and reduced hallucinations.
Elevating multi-task vehicle routing performance, this approach reduces solution gaps by over 21% while enhancing generalization across diverse problem variants.
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
On-policy self-distillation can boost diffusion model performance by up to 44% while slashing training time by over 60%.
IAPO redefines credit assignment in multi-turn interactions, showing that leveraging influence-dependency graphs can significantly enhance service agent performance.
CBPO redefines credit assignment in RLVR, enabling precise decision sensitivity that boosts performance across diverse benchmarks.
Jointly training tool creation and use allows LLMs to achieve unprecedented accuracy on procedural reasoning tasks, outperforming larger models and enhancing smaller ones.
Reinforcement learning can significantly boost the efficiency of scheduling heterogeneous satellites, achieving better utility and convergence than traditional optimization methods.
MetaRAG achieves a superior accuracy-efficiency trade-off in agentic RAG by aligning decision-making with the model's internal beliefs, outperforming traditional RL methods.
RePolicy achieves superior safety-policy invocation in language model agents, adapting dynamically to changing contexts and unseen trajectories.
FARCA transforms factual supervision into precise, reliability-weighted training signals, significantly boosting model factuality without sacrificing reasoning performance.
Grounding clinical language models in structured physiological knowledge can boost safety scores by over 21 percentage points, surpassing even state-of-the-art models like GPT-4.
A frozen instruct model can reshape a student's reasoning policy, enabling RL refinement that surpasses traditional methods without costly fine-tuning.
BALIGN filters out high-risk preference samples, preserving foundational model capabilities while optimizing alignment, achieving the best of both worlds.
Continuous skill verification in RL agents leads to a substantial performance boost, outperforming traditional static skill banks.
AdaptRubric's innovative two-stage framework boosts GUI reward modeling performance by over 3.6 F1 points, showcasing the power of task-adaptive criteria.
Self-improvement in LLMs can be achieved without external supervision by leveraging their own evaluative capabilities, leading to substantial performance gains across diverse benchmarks.
TAGR's innovative approach to real-time user intent modeling and ad tokenization leads to significant revenue gains in live-stream advertising.
Persuasion in LLM networks is not just about who speaks, but how the topology and exposure shape stance shifts, revealing a complex interplay of influence that traditional analysis overlooks.
MoPLEx achieves up to 43.7% improvement in clustering accuracy by effectively learning from complex multi-way rankings, revealing the power of leveraging language models for preference optimization.
Fine-tuning may preserve the underlying steering mechanism, but it can drastically undermine the intended behavioral effects, with an average 64% loss in effectiveness.
Ockhamareto achieves a staggering 49.9% mutation score while using 44% fewer tests than the best existing method, revolutionizing unit-test generation efficiency.
Faulty-code-driven test synthesis boosts code generation performance by 3% in LLMs, tackling reward hacking and validation issues head-on.
Emotion preference models can be dramatically improved by addressing both data sparsity and model bias, leading to more accurate emotional assessments in multimodal contexts.
Achieving expressive TTS now hinges on effectively optimizing non-verbal vocalizations, with design choices impacting NV fidelity more than previously understood.
Choosing AI for emotional support not only enhances immediate satisfaction but also reshapes long-term preferences away from human interaction.
CRPO reveals that leveraging English preference data can dramatically enhance multilingual LLM performance, outperforming standard methods across diverse languages.
TailSieve achieves up to 2.59x speedup in LLM rollouts by intelligently routing long-tail requests, transforming how we handle high-concurrency decoding.
SRPO enables LLMs to self-reflect and transform sparse feedback into dense learning signals, achieving state-of-the-art performance with drastically reduced training costs.
Mixed SFT outperforms next-chunk reasoning RL in reasoning tasks while using over 60 times less compute, reshaping our understanding of effective training strategies.
A single neuron can recalibrate LLM investment biases, enabling precise control over decision-making without altering the model's architecture or prompts.
DIAG reshapes practice distribution to maximize informative supervision, leading to significantly improved reasoning performance in LLMs.
Lever-Edit shows that you can effectively optimize image editing policies using T2I rewards, bypassing the need for costly editing-specific rewards altogether.
A carefully designed critic can provide a stable and efficient alternative to traditional group-relative advantage estimation in reinforcement learning for language models.
Shifting regularization to the input side allows for better exploration while maintaining response stability, leading to significant performance gains in LLM policy optimization.
Exploration bias in RL leads models to favor easy instructions, but a novel two-stage framework can unlock their potential for more challenging tasks.
Grounded, multi-dimensional rubrics can boost answer quality in open-domain question answering by over 6%, transforming how we evaluate and train AI systems for complex queries.
Models exhibit a surprising preference for their own prior answers, revealing inconsistencies in LLM value profiling across response formats.
GeoRisk-RAG slashes false confidence rates for location-sensitive queries to just 0.009, a game-changer for decision-making in natural hazard management.
Non-English users pay a significantly higher price for safety alignment in AI models, revealing systemic inequities in current practices.
FIRM-Video reveals that a checklist-driven approach can significantly enhance the reliability of text-to-video reward modeling, achieving state-of-the-art performance in evaluation metrics.
CLEAR slashes harmful completions from 32.3% to just 0.5% while boosting utility performance, redefining the safety-utility balance in LLMs.