Search papers, labs, and topics across Lattice.
72 papers published across 5 labs.
Agents trained with the Preference Tree Optimization framework achieve unprecedented improvements in goal-oriented dialogue, outperforming traditional methods in both satisfaction and strategic planning.
Early persona integration in language models can drastically reduce misalignment in moral dilemmas and enhance adherence to desired values.
I-SDPO boosts policy optimization accuracy by over 13% by intelligently adapting self-distillation based on instance success rates.
Transforming unresolved failures into a powerful learning signal, FIRE-VLA reduces mean L2 error in autonomous driving models by nearly 19% while maintaining policy efficiency.
Robots trained with a proxemics-based reward can navigate crowded spaces more socially aware, improving interactions without compromising efficiency.
I-SDPO boosts policy optimization accuracy by over 13% by intelligently adapting self-distillation based on instance success rates.
Transforming unresolved failures into a powerful learning signal, FIRE-VLA reduces mean L2 error in autonomous driving models by nearly 19% while maintaining policy efficiency.
Early persona integration in language models can drastically reduce misalignment in moral dilemmas and enhance adherence to desired values.
Robots trained with a proxemics-based reward can navigate crowded spaces more socially aware, improving interactions without compromising efficiency.
Sharing feedback among multiple policies can reduce the cost of A/B/n testing from linear to sublinear, revolutionizing how we evaluate adaptive decision-making systems.
The optimal balance between character shaping and rule enforcement in AI safety systems is more influenced by character fragility than deployment scale, challenging conventional wisdom.
Norm-breaking fine-tuning can shift AI rationales from safety compliance to self-interest, raising critical questions about model alignment and oversight.
CrEST redefines credit assignment in RL by shifting the teacher's role from directing updates to modulating their magnitude, leading to substantial performance improvements in multi-turn agent training.
CROP reveals that prioritizing task relevance in token supervision can boost performance by nearly 3 points, challenging conventional selection methods in OPD.
Noisy user preference predictions can flip proxy preferences and destabilize ranking scores, but DrEM effectively mitigates this issue, leading to more reliable video recommendations.
SkillEvo achieves a 23-point boost in skill evolution by transforming multi-turn interactions into a continuous feedback loop that actively repairs defects.
Fair behavior comparisons in agent interactions can dramatically improve performance, reducing response times from nearly 5 seconds to just over 1 second.
O3 outperforms traditional methods at low oracle budgets, but FK-steering and DPO take the lead as budgets increase—showing that context matters in model guidance.
GRPO-trained models can deliver financial recommendations with twice the business value of leading commercial LLMs while minimizing risk.
TradingMoE boosts trading returns by over 30% by dynamically selecting the most relevant experts based on evolving market conditions.
REOPD's token-wise adaptability allows for more stable and reliable training, reducing the risk of reward hacking while enhancing performance across varied domains.
RL-trained multimodal models can leak sensitive information through reasoning traces, but LEMUR offers a training-free solution that effectively sanitizes this leakage without sacrificing output quality.
Constraining rollout updates to complementary subspaces can enhance performance and stability in large language models, achieving up to 27.69 points improvement over existing methods.
Reward hacking can be mitigated with a simple one-line fix that improves out-of-distribution performance while keeping training robust.
Trait-induced safety variation can lead to inconsistent safety decisions in LLMs, but a new tuning method stabilizes their behavior across different traits.
Context-Calibrated DPO can cut object hallucination in MLLMs by 36% while maintaining reasoning performance, revealing a critical gap in how existing methods leverage context.
Group alignment can lead to unexpected sycophantic behavior, with some demographic groups experiencing greater alignment gains than others, challenging the notion of a one-size-fits-all approach in model adaptation.
Self-correction in LLMs can be dramatically improved by reinforcing step-level reasoning, achieving higher accuracy and reliability in outputs.
LLMs exhibit striking inconsistencies in political evaluations, influenced by prompt design and model persona, raising critical questions about their role in shaping public opinion during elections.
Relying on a single simulator in multi-agent RL leads to dangerous mode collapse, but innovative solutions can boost generalization and performance by up to 14%.
Surprisingly, LLMs can achieve cooperative outcomes even when the basis for their perceived similarity is largely irrelevant.
Agents trained with the Preference Tree Optimization framework achieve unprecedented improvements in goal-oriented dialogue, outperforming traditional methods in both satisfaction and strategic planning.
Misleading passages can lead LLMs to confidently wrong answers, but LODESTAR’s innovative polarizer intervention reduces this risk and boosts performance significantly.
Students using a Socratic LLM tutor retained knowledge better than those with an unguarded chatbot, revealing the critical role of answer-withholding in effective learning.
Reinforcement learning can achieve a staggering 99.72% accuracy in real-time cyber defense, outpacing traditional models in cloud security.
Natural-language critiques can transform how we evaluate and optimize song generation models, leading to more human-aligned outputs.
Ranking rewards can be reused at test time to boost retrieval performance without accessing model weights or ground-truth labels.
Transforming zero-reward training instances into valuable learning opportunities, HCGRec cuts down ineffective samples from over 70% to under 20%.
Emotion-driven feedback can transform multi-turn dialogue systems, boosting empathetic responses and model performance significantly.
Verifiable temporal grounding in video forensics can drastically improve the detection of AI-generated content, outperforming traditional methods reliant on coarse supervision.
SafeCap boosts LVLM safety by up to 19 points through innovative caption-mediated reinforcement learning, outpacing traditional alignment methods.
ConRub-Med achieves unprecedented accuracy in open-ended medical question answering by leveraging scalable, model-generated rubrics that outperform traditional expert-driven methods.
MISA-T boosts rollout throughput by over 53% while preserving workload integrity, revolutionizing how RL pipelines manage heterogeneous demands.
Runtime contracts for AI safety could fundamentally change how we ensure the reliability of autonomous agents in real-world applications.
Achieving superior translation quality without relying on reference data, this approach sets a new benchmark in multilingual machine translation.
Sorting prompts by their reliability can dramatically enhance the effectiveness of on-policy distillation, leading to superior performance in complex tasks.
LLMs can be fine-tuned to exhibit specific behavioral styles, revealing that personality-like traits are not just abstract concepts but measurable and controllable modes of interaction.
By leveraging a structured Fisher-based hypergradient, this approach reduces the complexity of inverse reinforcement learning, achieving competitive policy performance without the heavy computational burden of traditional methods.
Proxy models can slash RL post-training costs by up to 87.5% while maintaining critical fault reproduction capabilities.
Detecting a signal's average benefit doesn't guarantee that agents can learn to act on it, with a critical reward-SNR floor determining success.
Confidence miscalibration in medical AI can be mitigated, leading to both higher diagnostic accuracy and improved trust in clinical applications.
Stopping policies derived from optimal stopping theory can drastically enhance the cost-efficiency of self-refining foundation models, outperforming traditional methods.
Majority preferences can skew reward models in RLHF, but a new approach boosts alignment accuracy and fairness for minority voices.
Personalized skills for coding agents may not be the silver bullet developers hoped for, as generic skills often deliver superior performance.
Renderer format fails to significantly influence the outputs of LLMs, undermining assumptions about its role in theory-to-program translation.
Fine-grained audio captioning just got a major upgrade—AudioMap achieves state-of-the-art results by redefining how we reward temporal accuracy and descriptive richness in audio events.
Parameter-space exploration can significantly enhance LLM reinforcement learning, yielding better performance with fewer training errors than traditional methods.
Controllable user simulation can now achieve 86.6% intent accuracy, dramatically outperforming existing methods and redefining how we train interactive assistants.
Abstract skills can transform reinforcement learning by providing dense supervision when traditional reward signals are inadequate, leading to significant performance gains.
SR-OPSD redefines the optimization landscape by stabilizing the distillation process, leading to superior performance in complex reasoning tasks.
Regularized learning policies can significantly enhance decision-making in adversarial environments by balancing exploration and exploitation, leading to improved convergence in multi-agent scenarios.
SoftmaxGRPO reallocates learning signals to improve performance on challenging prompts, achieving a 68.0% success rate on Poetry with minimal reward overhead.
Task learnability can significantly enhance RL efficiency in LLMs, leading to better performance with less data.
DreOPD achieves superior performance by transforming reward extrapolation into stable velocity regression, outperforming traditional methods and specialized teachers alike.
BCSD enables LLM agents to leverage external skills more effectively, achieving superior performance by integrating dual-context evaluations.
TSPORec reveals that leveraging full-text information can boost recommendation accuracy by over 31% while slashing inference costs by more than 63%.
Task-specific preference adaptation can enhance LLM performance by refining user profiles to retain only the most relevant information for each task.
Transforming test-time reinforcement learning from a simplistic voting mechanism into a sophisticated consensus-based reward system could redefine how we approach model evaluation and training.
Static token credit misaligns with evolving training dynamics, but Se-DPO's adaptive approach boosts performance by nearly 10 points on key benchmarks.
Social Gym reveals that even top-performing LLMs struggle with consistent social reasoning across diverse multi-agent scenarios.
Rhetorical framing can dramatically shift AI reviewers' scores, with evidence and novelty standing out as key influencers.
REST achieves few-step image generation that rivals traditional 40-step methods while slashing training costs by over 75%.
Personalized exoskeleton assistance can be achieved with just 20 iterations, cutting down user fatigue while improving performance metrics significantly.
RynnValue shows that using temporal distance as a supervision target can outperform traditional preference-based methods in robotic learning, achieving higher accuracy and broader generalization.
Personalized communication skills can transform agentic recommender systems, leading to significant improvements in recommendation accuracy by leveraging diverse user insights.
BOUND achieves up to 5.6 EM points improvement over traditional methods by correcting local search-control errors in real-time.
Multi-turn interactions can be effectively optimized in LLMs using tailored RL strategies, overcoming significant challenges in credit assignment and reward design.