Search papers, labs, and topics across Lattice.
77 papers published across 6 labs.
TP-MPPO achieves up to 87.5% higher goodput for LLM inference in edge networks, revolutionizing how we manage bandwidth and task offloading.
Mixed SFT outperforms next-chunk reasoning RL while consuming over 60 times less compute, reshaping our understanding of effective training strategies with no-CoT data.
Static sample selection in reinforcement fine-tuning is a recipe for suboptimal updates—DIEM adapts dynamically, leading to superior performance on reasoning tasks.
Models show a striking 40% drop in verbalized commitment when cues come from tool returns instead of user messages, challenging assumptions about reasoning trace reliability.
RLVR chokes reasoning diversity at the front door rather than during execution: an 11x–16x likelihood collapse occurs before the very first operation, leaving downstream solution paths intact and fully recoverable via targeted late-layer weight interpolation.
Static sample selection in reinforcement fine-tuning is a recipe for suboptimal updates—DIEM adapts dynamically, leading to superior performance on reasoning tasks.
Models show a striking 40% drop in verbalized commitment when cues come from tool returns instead of user messages, challenging assumptions about reasoning trace reliability.
RLVR chokes reasoning diversity at the front door rather than during execution: an 11x–16x likelihood collapse occurs before the very first operation, leaving downstream solution paths intact and fully recoverable via targeted late-layer weight interpolation.
This paper proposes World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck and accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines.
Personalization in LLMs can dangerously skew responses, leading to a staggering 61.7% increase in sycophantic bias.
TTPO achieves label-supervised performance without any ground-truth labels, outperforming traditional methods on key benchmarks.
Emotional preferences can autonomously reshape goal priorities in agents, leading to more adaptive decision-making in dynamic environments.
Tackling the GPU bottleneck in RLM training could unlock a new era of scalable and efficient reasoning models.
Equal ranking quality can lead to drastically different decision outcomes, with OC-SFT providing a solution that enhances stability in LLM scoring.
The coefficient $\beta$ in DPO not only controls preference noise but also entangles optimization dynamics, leading to non-intuitive learning behaviors that can mislead model training.
GRAS achieves superior training-free reward alignment in discrete diffusion models, outperforming previous methods and rivaling fine-tuned approaches with no added computational cost.
PRISM reveals that a structured approach to persona fidelity evaluation can drastically outperform traditional methods, providing more reliable insights into LLM behavior.
Merging existing expert capabilities can yield significant performance boosts, but choosing the right fusion method can make all the difference in multi-domain reinforcement learning.
Weak model guidance can dramatically enhance LLM reasoning coverage in RLVR, outperforming traditional methods as task complexity increases.
Instruction quality is the hidden bottleneck in preference learning, and refining it can dramatically enhance model alignment.
SPEAR bridges the reasoning gap in reinforcement learning by transforming complex natural-language reasoning into efficient symbolic milestones, enabling dense and logical reward signals without costly neural verifiers.
Reinforcement learning can transform how models autonomously verify scientific accuracy, achieving near-human levels of reasoning in error detection.
Automatic metrics can significantly reduce human annotation costs while maintaining unbiased evaluations, as shown by the new Prediction-Powered Saving Ratio (PPSR).
Multi-expert aggregation not only mitigates risk at the decision threshold but also optimizes scoring functions, leading to substantial improvements in dialogue evaluation accuracy.
Reducing sycophancy in language models can inadvertently hinder their ability to rationally update, revealing a critical trade-off in model behavior.
RubricRM adapts evaluation criteria dynamically, leading to significant performance gains in visual generative tasks compared to static reward models.
Idiolectal paraphrasing transforms how models learn reasoning by allowing them to express complex thoughts in their own unique language, leading to significant performance gains.
Rubric-based scoring reveals that vision-language models can significantly improve their grounding in visual evidence, enhancing reasoning and instruction adherence.
P4-DT outperformed human surrogates in predicting patient preferences, achieving an impressive 81.7% accuracy by leveraging contextual dilemmas.
Filtering out fit sequences during fine-tuning can boost post-training performance by up to 17%, reshaping how we approach model training for RL applications.
Multi-objective RL methods overlook critical interactions between non-linear utility effects across different timescales, leading to suboptimal decision-making.
Pairwise rewards in reinforcement learning can significantly boost the robustness of LLM auditors, enhancing their ability to detect hidden model behaviors with minimal false positives.
GRIP achieves superior accuracy-efficiency trade-offs in reasoning tasks by intelligently interpolating parameters from two distinct model types without retraining.
TP-MPPO achieves up to 87.5% higher goodput for LLM inference in edge networks, revolutionizing how we manage bandwidth and task offloading.
Confidence estimates from LLMs can be misleading when evaluating many candidates, but a new framework ensures high-probability agreement with human judgments.
AutoVerifier learns from its mistakes, transforming verifier errors into reusable strategies that dramatically boost verification accuracy.
Mixed-policy reinforcement learning can enable language models to absorb knowledge more effectively than traditional supervised fine-tuning, especially in complex reasoning scenarios.
Listwise rankings from vision-language models can outperform traditional pairwise methods in training robotic policies, achieving competitive success rates with greater flexibility.
Free-form language reasoning can dramatically enhance robotic manipulation, outperforming traditional instruction-based methods in complex tasks.
The fatal flaw of on-policy self-distillation isn't model capacity but information asymmetry, which systematically collapses reasoning diversity across three tunable operational levers.
Self-improving search agents thrive when feedback and policy evolution are intertwined, leading to sustained performance gains and reduced hallucinations.
Elevating multi-task vehicle routing performance, this approach reduces solution gaps by over 21% while enhancing generalization across diverse problem variants.
OPDVR transforms the landscape of model distillation by ensuring that only correct trajectories enhance learning, leading to significant performance gains on reasoning tasks.
On-policy self-distillation can boost diffusion model performance by up to 44% while slashing training time by over 60%.
IAPO redefines credit assignment in multi-turn interactions, showing that leveraging influence-dependency graphs can significantly enhance service agent performance.
CBPO redefines credit assignment in RLVR, enabling precise decision sensitivity that boosts performance across diverse benchmarks.
Jointly training tool creation and use allows LLMs to achieve unprecedented accuracy on procedural reasoning tasks, outperforming larger models and enhancing smaller ones.
Reinforcement learning can significantly boost the efficiency of scheduling heterogeneous satellites, achieving better utility and convergence than traditional optimization methods.
MetaRAG achieves a superior accuracy-efficiency trade-off in agentic RAG by aligning decision-making with the model's internal beliefs, outperforming traditional RL methods.
RePolicy achieves superior safety-policy invocation in language model agents, adapting dynamically to changing contexts and unseen trajectories.
FARCA transforms factual supervision into precise, reliability-weighted training signals, significantly boosting model factuality without sacrificing reasoning performance.
Grounding clinical language models in structured physiological knowledge can boost safety scores by over 21 percentage points, surpassing even state-of-the-art models like GPT-4.
A frozen instruct model can reshape a student's reasoning policy, enabling RL refinement that surpasses traditional methods without costly fine-tuning.
BALIGN filters out high-risk preference samples, preserving foundational model capabilities while optimizing alignment, achieving the best of both worlds.
Continuous skill verification in RL agents leads to a substantial performance boost, outperforming traditional static skill banks.
AdaptRubric's innovative two-stage framework boosts GUI reward modeling performance by over 3.6 F1 points, showcasing the power of task-adaptive criteria.
Self-improvement in LLMs can be achieved without external supervision by leveraging their own evaluative capabilities, leading to substantial performance gains across diverse benchmarks.
TAGR's innovative approach to real-time user intent modeling and ad tokenization leads to significant revenue gains in live-stream advertising.
Persuasion in LLM networks is not just about who speaks, but how the topology and exposure shape stance shifts, revealing a complex interplay of influence that traditional analysis overlooks.
MoPLEx achieves up to 43.7% improvement in clustering accuracy by effectively learning from complex multi-way rankings, revealing the power of leveraging language models for preference optimization.
Fine-tuning may preserve the underlying steering mechanism, but it can drastically undermine the intended behavioral effects, with an average 64% loss in effectiveness.
Ockhamareto achieves a staggering 49.9% mutation score while using 44% fewer tests than the best existing method, revolutionizing unit-test generation efficiency.
Faulty-code-driven test synthesis boosts code generation performance by 3% in LLMs, tackling reward hacking and validation issues head-on.
Emotion preference models can be dramatically improved by addressing both data sparsity and model bias, leading to more accurate emotional assessments in multimodal contexts.
Achieving expressive TTS now hinges on effectively optimizing non-verbal vocalizations, with design choices impacting NV fidelity more than previously understood.
Choosing AI for emotional support not only enhances immediate satisfaction but also reshapes long-term preferences away from human interaction.
TailSieve achieves up to 2.59x speedup in LLM rollouts by intelligently routing long-tail requests, transforming how we handle high-concurrency decoding.
SRPO enables LLMs to self-reflect and transform sparse feedback into dense learning signals, achieving state-of-the-art performance with drastically reduced training costs.
Mixed SFT outperforms next-chunk reasoning RL while consuming over 60 times less compute, reshaping our understanding of effective training strategies with no-CoT data.
A single neuron can recalibrate LLM investment biases, enabling precise control over decision-making without altering the model's architecture or prompts.
DIAG reshapes practice distribution to maximize informative supervision, leading to significantly improved reasoning performance in LLMs.
Lever-Edit shows that you can effectively optimize image editing policies using T2I rewards, bypassing the need for costly editing-specific rewards altogether.
A carefully designed critic can provide a stable and efficient alternative to traditional group-relative advantage estimation in reinforcement learning for language models.
Shifting regularization to the input side allows for better exploration while maintaining response stability, leading to significant performance gains in LLM policy optimization.
Exploration bias in RL leads models to favor easy instructions, but a novel two-stage framework can unlock their potential for more challenging tasks.
Grounded, multi-dimensional rubrics can boost answer quality in open-domain question answering by over 6%, transforming how we evaluate and train AI systems for complex queries.
Models exhibit a surprising preference for their own prior answers, revealing inconsistencies in LLM value profiling across response formats.
CRPO reveals that leveraging English preference data can dramatically enhance LLM performance in low-resource languages, outperforming traditional methods.
GeoRisk-RAG slashes false confidence rates for location-sensitive queries to just 0.009, a game-changer for decision-making in natural hazard management.
Non-English users pay a significantly higher price for safety alignment in AI models, revealing systemic inequities in current practices.
FIRM-Video reveals that a checklist-driven approach can significantly enhance the reliability of text-to-video reward modeling, achieving state-of-the-art performance in evaluation metrics.
Probing for user information can be costly, but RO-PnR shows that strategic questioning can significantly enhance the effectiveness of health misinformation interventions.