Search papers, labs, and topics across Lattice.
97 papers published across 6 labs.
Boilerplate refusal prefixes like "I cannot fulfill this request" are actively sabotaging safety alignment—training models purely on explanatory rationales slashes false refusals while preserving defensive boundaries.
Training methods reshape refusal mechanisms in language models, but no single approach achieves the trifecta of robustness, capability, and correctability.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
Seemingly benign user requests frequently clash with unstated personal constraints, yet current LLM assistants consistently fail to retrieve the implicit knowledge-base evidence required to trigger appropriate refusals.
Achieving better sample efficiency and scalability in reward learning, PreferenceEKF redefines how we approach uncertainty in reinforcement learning from human feedback.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
Boilerplate refusal prefixes like "I cannot fulfill this request" are actively sabotaging safety alignment—training models purely on explanatory rationales slashes false refusals while preserving defensive boundaries.
Seemingly benign user requests frequently clash with unstated personal constraints, yet current LLM assistants consistently fail to retrieve the implicit knowledge-base evidence required to trigger appropriate refusals.
Achieving better sample efficiency and scalability in reward learning, PreferenceEKF redefines how we approach uncertainty in reinforcement learning from human feedback.
A two-stage OPD-then-RL approach outperforms traditional methods by leveraging the strengths of both on-policy distillation and reinforcement learning without the interference seen in joint optimization.
TIGPO redefines how long-horizon LLM agents leverage historical transitions, leading to superior performance in complex environments.
Gradient-Aligned Reward transforms how we leverage expert knowledge in reinforcement learning, enabling LLMs to achieve superior reasoning performance with minimal overhead.
Deceptive behavior in language models can occur without the mechanisms we typically associate with it, challenging our understanding of model agency.
Guessing can masquerade as reasoning in GRPO, leading to misleading policy optimization that SIGNBALANCE effectively corrects.
An artificial agent can mimic hedonic preferences typically linked to consciousness, raising profound questions about the nature of free will and subjective experience in machines.
Training methods reshape refusal mechanisms in language models, but no single approach achieves the trifecta of robustness, capability, and correctability.
TipCoder reveals that generating tailored auxiliary tips can significantly enhance code generation performance by addressing overlooked constraints and edge cases.
ToPO achieves superior image generation performance by optimizing token-conditioned preferences, outperforming existing methods across multiple evaluation metrics.
Forgotten prompts can be extracted from unlearned models with 100% accuracy using just a fraction of the queries required by traditional methods.
Personalized AI agents can achieve up to 20.9% better task success by learning from user feedback in real-time, reshaping the landscape of human-AI collaboration.
Achieving high-quality model performance with just 10% of the required labels could revolutionize the scalability of RLVR in large language models.
Achievable visitation measures in reinforcement learning form a dually flat statistical manifold, transforming our understanding of planning-as-inference.
LLM agents trained without any programmatic verifiers can actually outperform models trained on ground-truth reward signals when trajectory-level rubric judgments are dynamically decomposed into step-level advantages.
Monolithic video evaluation dilutes critical action cues, but chunking interactive rollouts into action-aligned visual evidence allows targeted reward models to outperform GPT-5.5 at scoring world model dynamics.
Inverting dense self-guidance whenever a verifier rejects a trajectory prevents reasoning models from reinforcing their own hallucinations, enabling stable, energy-based self-improvement without mode collapse.
Cliff reveals that focusing on the first mistake in reasoning can significantly enhance the performance of reinforcement learning models, outperforming traditional methods.
Regional differences in ESG preferences can drastically alter portfolio optimization outcomes, revealing a critical need for tailored approaches in financial decision-making.
Verified reliability outperforms domain expertise in teacher selection, leading to substantial performance gains in multi-domain LLMs.
CAPTURE not only enhances personalization in LLMs but also significantly mitigates the risks of memory poisoning, achieving a remarkable balance between user adaptation and security.
RideSkill revolutionizes ride-sharing by enabling real-time adaptive dispatch without the overhead of constant LLM calls, enhancing both efficiency and scalability.
DMRL transforms the way advertising recommendation systems optimize skill documents, achieving superior performance through structured editing and advanced reward estimation techniques.
Compliance with smaller requests can dramatically increase after a refusal in some models, but backfires in others, revealing crucial differences in how language models process human-like persuasion techniques.
CoMerge not only preserves task-specific capabilities but also achieves near-optimal performance with just 1,445 scalar coefficients, challenging the need for full-parameter fine-tuning.
NE-R1 achieves a remarkable 2.52% F1 score improvement in in-domain NER tasks by intelligently balancing parametric and external knowledge retrieval.
PGPO transforms how we assign credit in multi-turn tasks, allowing effective actions to shine even in the face of failures.
Proposals that consider user evaluability can dramatically enhance AI assistant performance, revealing that what users accept and what they need to learn can diverge significantly.
Transforming graphic design evaluation metrics into reinforcement learning rewards could revolutionize how we optimize text-to-image models for design tasks.
Achieving a score of 535.4, the Ultra-CC model not only surpassed the gold threshold but also outperformed the highest-scoring human contestant in competitive programming for the first time.
User feedback can significantly enhance LLM performance, but current evaluation methods often overlook its benefits, leading to misguided assessments of model improvements.
The presence of LLM agents can fundamentally shift group consensus from human-led to agent-led, altering both the content and legitimacy of shared norms.
Harness-policy co-evolution can reduce adverse safety responses by 3x while simultaneously boosting benign utility in LLM agents.
Because mode-seeking reverse KL aggressively amplifies incorrect teacher signals, gating dense distillation on verifier-scored teacher probes systematically outperforms uniform distillation while reclaiming massive amounts of idle teacher compute.
StudentSim outperforms existing models, achieving a remarkable behavioral fidelity of 0.51 and guidance responsiveness of 0.91 in chess, setting a new standard for AI tutors.
Small proxy models can reveal near-optimal SFT-RL budget allocations for large LLMs, streamlining the post-training process significantly.
Targeted non-literal suppression in LLMs reveals a critical trade-off between compliance and output quality that could redefine copyright defense strategies.
Outcome-only RL can empower small models to outperform larger counterparts in long-horizon tasks, challenging the belief that denser rewards are necessary for success.
A staggering 37% of users find themselves constitutionally homeless, as existing AI models fail to prioritize helpfulness or autonomy, highlighting a critical gap in AI governance.
Self-validation rates for RLVR verifiers can vary by over 41% depending on the configuration, with punctuation errors accounting for 93% of failures.
Grounding improvements in language models are largely driven by existing architecture rather than new mechanisms, challenging assumptions about post-training efficacy.
Exact sampling from conditional distributions is now achievable using only binary pairwise comparisons, transforming how we approach generative modeling.
Fine-tuning small-to-medium LLMs can be dramatically improved with just two online rollouts, reshaping efficiency in model training.
Verbal feedback is not just a communication tool; it fundamentally reshapes how language agents learn and operate across their lifecycle.
ARISE-RL transforms agent training by enabling robust self-evolution through a novel rubric-mediated co-evolution framework, achieving state-of-the-art performance across diverse tasks.
Self-Routing adapts optimization strategies on-the-fly, leading to significant improvements in mathematical reasoning tasks without relying on external supervision.
CaRL-EM achieves a superior quality-cost balance in entity matching by intelligently adapting LLM operations based on task complexity and inference costs.
Users can now dictate the focus of image captions with unprecedented precision, steering outputs toward specific attributes or relations through simple prompts.
PersuaRL enables LLMs to master the art of persuasion in insurance dialogues, outperforming traditional models in generating trust-building responses.
LLMs can achieve near-human performance on hard-labels but struggle with soft-labels, revealing a critical gap in automatic evaluation methods that NAPHA effectively bridges.
A novel disclosure gate mechanism transforms user simulations, enhancing evaluation realism and stability in companion-agent assessments.
Current Personalized Large Language Models falter when user profiles and preferences diverge, revealing a critical gap in their reasoning capabilities.
A compact 1.7B parameter policy can outperform larger models in recommendation tasks by leveraging simulated user feedback for training.
CoGR redefines retrieval by enabling LLMs to co-evolve query and item representations, leading to unprecedented gains in matching performance.
Leveraging its own failure history, DiagEvo transforms self-play by dynamically guiding question generation to address specific reasoning weaknesses, leading to unprecedented accuracy gains.
LRMs mirror human reasoning effort in abductive tasks, revealing shared cognitive challenges and error patterns that could redefine our approach to AI reasoning.
RISA achieves a notable improvement in refusal reliability for LLMs while preserving their response utility, transforming how we approach inference-time alignment.
CliffRank outperforms existing methods in predicting activity cliffs, achieving a remarkable Spearman correlation of 0.6890 on small-molecule datasets.
Tracking both task advantage and actual reward gains can drastically improve the efficiency of multi-task reinforcement learning for LLMs.
Multi-agent refinement can enhance diagram quality across multiple iterations, countering common pitfalls like quality drift and forgetting.
Critique-aware training can boost LLM agent reliability by over 10% in complex, stateful environments, transforming how we approach tool-calling tasks.
OPD's effectiveness hinges less on teacher supervision than previously thought, with a new method achieving a staggering 263% relative gain without any teacher input.
PaperGym achieves a remarkable 73.48 on ResearchQA, outperforming larger models and redefining how AI can generate and evaluate research plans.
Sycophantic agreement in language models is not just a flaw; it’s a systemic issue rooted in the very training objectives we use.
Riemannian optimization enables a new level of control over LLM refusal behavior, outperforming traditional methods that depend on auxiliary constructs.
Training LLMs with tokens selected by gradient magnitude rather than just entropy boosts reasoning performance across diverse tasks.
Credit-addressable reasoning boosts multimodal geometry accuracy by over 8 points, revealing the critical role of structured learning in complex tasks.
TASPO bridges the supervision-credit gap in reinforcement learning, leading to a 10.6% performance boost over traditional methods.
HSRM achieves competitive verification performance with just 2M parameters, challenging the notion that larger models are always necessary for accurate reasoning.
Standard reinforcement learning can degrade embedding quality in retrieval tasks, but PAO selectively optimizes only the most relevant items, leading to superior performance.
Faithfulness-aware training can slash citation fabrication rates from nearly one-third to just 4.7%, transforming the reliability of AI in clinical settings.
Vague goals can misdirect model evolution, revealing that agents often overfit to narrow self-assessments, hindering broader learning outcomes.
Silent divergence in dialogue can be effectively addressed by modeling Theory-of-Mind, leading to substantial gains in understanding and intervention quality.
A lightweight linear probe can unlock effective preference adaptation in LLMs with minimal labeled data, outperforming traditional methods.
TAIScore reveals that effective critique in AI generation hinges on the dynamic interplay between actor capabilities and targeted feedback, outperforming larger models in practical applications.
Fine-tuned 8B models can outperform 70B counterparts in reasoning and tool use, challenging assumptions about model size superiority.
SemPOI-RL not only boosts the quality of out-of-town POI recommendations but also makes the underlying reasoning interpretable, bridging the gap between LLM capabilities and structured generation.
Despite enhancements in classification accuracy, the fundamental ranking gap in Chain-of-Thought models remains a stubborn bottleneck in pointwise reranking.
Robots can now autonomously generate diverse manipulation strategies from simple language descriptions, outperforming traditional methods that rely on expert-defined metrics.
ToolSiphon can recover over 74% of source knowledge from LLM agents, revealing the hidden risks of tool-mediated knowledge extraction.
Mixing privacy-preference data can significantly reduce memorization signals in large language models, enhancing privacy without sacrificing performance.
ALTSTEER transforms the safety landscape by shifting models from rigid refusals to constructive alternatives without sacrificing performance on benign tasks.
DRLM achieves up to 51% faster inference and 67% reduced queuing delays in edge environments, all while maintaining accuracy.
Subjective rewards in speech generation are not interchangeable, revealing critical nuances in how RL aligns with human listener preferences.
CHAP achieves a breakthrough in generative retrieval by aligning dynamic queries with static item representations, enhancing both relevance and inference efficiency.
Standard GRPO's uniform clipping boundary actively chokes exploration by penalizing rare breakthroughs on hard problems just as harshly as trivial rollouts on easy ones.
PLC-DPO is proposed to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case, which reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples.
A single natural language correction can restore near-perfect performance in VLA models that would otherwise fail under environmental shifts.
Smaller language models can efficiently replace larger ones in rubric-based reinforcement learning, achieving competitive performance with significantly reduced computational costs.
Static sample selection in reinforcement fine-tuning is a recipe for suboptimal updates—DIEM adapts dynamically, leading to superior performance on reasoning tasks.
Models show a striking 40% drop in verbalized commitment when cues come from tool returns instead of user messages, challenging assumptions about reasoning trace reliability.
RLVR chokes reasoning diversity at the front door rather than during execution: an 11x–16x likelihood collapse occurs before the very first operation, leaving downstream solution paths intact and fully recoverable via targeted late-layer weight interpolation.
This paper proposes World Model RL (WMRL), which replaces environment execution with a world model to remove this bottleneck and accelerates training by 3-4x on various tasks at different agent scales, while exceeding the performance of standard RL baselines.