Search papers, labs, and topics across Lattice.
71 papers published across 6 labs.
Visual insensitivity in multimodal LLMs can be effectively mitigated with a new framework that enhances alignment and reduces hallucinations.
By making environment design a learnable process, SPADE unlocks a new frontier in self-improvement for language agents, leading to substantial performance gains across diverse tasks.
CLEAR slashes harmful completions from 32.3% to just 0.5% while boosting utility performance, redefining the safety-utility balance in LLMs.
RGA-Designer cuts token consumption by over 20% without sacrificing task accuracy, revolutionizing how we design communication topologies in multi-agent systems.
Truncating low-ranked completions can outperform traditional rank-based policies by reducing reliance on brittle top rankings while still achieving strong alignment results.
CLEAR slashes harmful completions from 32.3% to just 0.5% while boosting utility performance, redefining the safety-utility balance in LLMs.
RGA-Designer cuts token consumption by over 20% without sacrificing task accuracy, revolutionizing how we design communication topologies in multi-agent systems.
Truncating low-ranked completions can outperform traditional rank-based policies by reducing reliance on brittle top rankings while still achieving strong alignment results.
A novel reinforcement learning reward derived from positive pairs alone boosts keypoint learning stability and performance, even in low-texture environments.
A simple Gaussian counterexample reveals that existing concentration inequalities for discounted least-squares estimators are fundamentally flawed, necessitating significant corrections.
Adaptive reasoning in language models can reduce token usage by 41% while maintaining high accuracy, reshaping efficiency in AI computations.
SAPO achieves a 15.1 percentage point improvement over PPO while slashing memory costs and runtime by a third, revolutionizing how we optimize agentic RL.
Routing LLM judges based on their roles can significantly enhance evaluation efficiency and effectiveness, revealing when to stop calling judges to optimize performance.
Visual insensitivity in multimodal LLMs can be effectively mitigated with a new framework that enhances alignment and reduces hallucinations.
Continual adaptation to evolving prompt injection attacks can enhance LLM defenses by up to 6.3 times, addressing a critical vulnerability in AI systems.
SSR-GRPO significantly reduces noise in retrieval systems, enhancing relevance assessments and enabling more accurate e-commerce search results.
Nested sequential Monte Carlo methods can dramatically improve inference-time control in text generation, outperforming traditional techniques in steering towards desired outcomes.
FAR-DPO boosts the success rate of cyclic peptide design by over 10% while ensuring robust performance across challenging targets.
Optimizing latent visual representations can boost multimodal reasoning performance by over 9% on complex tasks.
Manifold drift can lead to substantial misalignment in generative models, but ThermoDPO offers a powerful solution that anchors preference optimization to the pretrained data manifold.
Human-mediated AI guidance can transform how families engage with emergency preparedness, making it more interactive and tailored to children's needs.
Subjectivity-Adaptive soft-Label Training (SALT) transforms LLM training by embracing the inherent variability in human responses, leading to more robust social simulations.
Annotations can be transformed into powerful oracle rollouts, dramatically enhancing the efficiency of reinforcement learning for video MLLMs.
A curriculum that intelligently adapts to diverse user preferences can double population satisfaction while slashing training time.
RP1 achieves near-perfect planning success with 1,000 times fewer world-model rollouts and is up to 67 times faster than traditional methods.
LLMs are miscalibrated in their reasoning, failing to distinguish between scenarios where valid solutions exist and where they do not, which could undermine their effectiveness in real-world applications.
VAKE reveals that over 80% of the knowledge activated through explicit priming is crucial for answering questions, showcasing a new pathway to enhance LLM factual accuracy.
Pairwise ranking not only outperforms single-action RL in explanation selection but also slashes serving costs and latency, making it a game-changer for industrial applications.
Cost-bounded self-verification allows LLMs to self-correct efficiently, cutting down on response generations while maintaining accuracy.
A judge's benchmark accuracy can significantly overstate its true evaluative competence, revealing critical insights for optimizing skill evaluation in AI systems.
MR-IQA-2 reveals that decoupling reasoning from rating can significantly enhance the reliability of image quality assessments, achieving human-level alignment without sacrificing interpretability.
RL-driven perception models can now achieve unprecedented accuracy in extremely dense visual scenes, eliminating the need for complex hyperparameter tuning.
Clinically structured rewards can enhance medical image captioning accuracy, yielding up to 5.8% better factuality in generated descriptions.
Reinforcement learning can significantly enhance the robustness of point cloud quality assessment, achieving state-of-the-art performance across diverse datasets.
PALATE enables personalized portrait retouching at a fraction of the cost, requiring only 512 bytes of user-specific data while achieving superior preference prediction accuracy.
Teacher rewards can mislead learning, but R2-OPD filters out conflicting feedback to boost reasoning performance in language models.
HARP outperforms traditional CVE prioritization methods by dynamically adapting to implicit operational preferences, revealing the power of leveraging historical data without explicit prompts.
Reward function design just got a major upgrade—MLREF achieves 25.2% better performance by reusing and evolving reward components across iterations.
Transition-level comparisons in Dream2Reward reveal that even subtle missteps in robotic motion can be effectively penalized, leading to significantly improved learning outcomes.
By making environment design a learnable process, SPADE unlocks a new frontier in self-improvement for language agents, leading to substantial performance gains across diverse tasks.
VA-Judger transforms reward modeling by integrating human preference feedback, leading to coherent and high-quality video-audio generation that outperforms traditional metrics.
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.
GUPO reveals that accounting for gradient uncertainty can dramatically improve policy optimization in post-training LLMs, leading to more effective reasoning capabilities.
Bridging fragmented research, this survey offers a unified taxonomy that redefines human-centric intelligence in the age of foundation models.
Next-turn user reactions can boost multi-turn agent performance by over 10 percentage points, revealing the critical role of local feedback in reinforcement learning.
Achieving over three times the accuracy of an untrained model, this research uncovers the transformative potential of domain-specific training in signal mathematical reasoning.
LLM-derived rewards can maintain optimal policy invariance even when the feedback is inaccurate, challenging the limitations of conventional reward shaping methods.
Harnessed agentic RL can boost agent performance significantly, as shown by a 14.6-point improvement in coding tasks with minimal training data.
LLM-derived preference judgments reveal significant inconsistencies, undermining the reliability of using a single utility function for decision-making.
Sharing rollout feedback across related samples can significantly boost exploration efficiency in RLVR, leading to better reasoning capabilities in large language models.
Optimization success in LLM unlearning can be misleading, as different evaluation metrics may yield conflicting insights about model behavior.
Reader-specific preferences can enhance performance, but they don't ensure consistent intervention success across contexts—highlighting a critical gap in our understanding of evidence utility in ML systems.
Immediate feedback-driven corrections can enhance robotic manipulation performance without the overhead of retraining entire policies.
Speeding up reinforcement learning algorithms by up to 5.7 times, rl-triton revolutionizes how we handle credit assignment on GPUs.
Cooperative multi-agent training can unlock unsupervised reasoning capabilities in RL, yielding performance gains that rival supervised methods without the need for costly annotations.
TSFT redefines how we allocate fine-tuning resources in CRL, achieving near-oracle performance while maximizing task coverage.
Achieving a 15.4% shift in LLM sycophancy control with a method that ensures predictable and gradual adjustments could redefine user interactions with AI.
Reallocating optimization effort based on reward saturation can boost performance by up to 9.2% in complex reasoning tasks.
QVIRL achieves robust apprenticeship learning from raw pixel data while quantifying uncertainty, a breakthrough for safety-critical AI applications.
Privileged Value Functions can inject crucial token-level signals into LLM reinforcement learning, leading to substantial performance gains over traditional methods.
Human preferences in ethical decision-making for autonomous vehicles reveal a troubling preference for self-sacrifice over minimizing casualties, challenging traditional ethical frameworks.
PertMind reveals that leveraging cellular perturbation data can significantly enhance LLMs' biological reasoning capabilities without extensive task-specific retraining.
Semantically informative labels can skew LLM decision-making, leading to either enhanced performance or catastrophic failures depending on alignment with reward structures.
CAPO enables LLMs to meet stringent operational constraints without sacrificing task performance, achieving feasible prompts in every evaluated domain.
Annotator identity can drastically alter benchmark outcomes, yet current leaderboards mask this variability, misleading researchers about model performance.
Reinforcement learning can transform whole-slide image analysis, achieving rapid tumor segmentation without sacrificing accuracy.
Expert-guided policy revisions in language models can drastically improve diagnostic accuracy, achieving up to a 32.7 percentage point increase in Recall@1 for rare diseases.
Models trained with ACA-RL not only outperform on missing-premise tasks but also redefine how we evaluate reasoning under uncertainty in NLP.
BabelSteering boosts harmful request refusals across multiple languages by an average of 11 percentage points, all while maintaining task performance.
Timing the entry of preference dimensions can lead to substantial performance gains in multi-preference alignment for LLMs.
OPD transfers reasoning skills rather than answers, revealing a complex interplay between teacher-student origins that can either enhance or hinder model capabilities.
Reducing uncertainty in LLM interactions can dramatically enhance recommendation quality without needing ground-truth data.
Evolving AI avatars can achieve near-human performance in live-stream interactions while adapting to real-time changes without retraining.
Sentiment drift in RLHF can strip emotional nuance from summaries, but a new regularization technique can mitigate this effect without sacrificing quality.
GAINS reveals that effectively modeling human imperfection can boost task success rates by over 20% in robot manipulation tasks.
Achieving a mean visual order consistency of 0.9872, Robo-Dopamine 2.0 significantly outperforms traditional reward models, demonstrating its superior ability to handle OOD scenarios in robotic manipulation.