Search papers, labs, and topics across Lattice.
80 papers published across 9 labs.
VLMs struggle to provide accurate annotations in video games, revealing significant gaps in their understanding of dynamic environments.
Annotating video game datasets with VLMs can transform the way RL agents learn by simplifying reward extraction and conditioning.
Reinforcement learning can transform persistence diagrams into dynamic entities, balancing complexity reduction with the preservation of essential topological features.
A single lightweight adapter can enhance language model personalization without the need for per-user fine-tuning or costly forward passes, achieving robust performance across diverse user contexts.
LLMs can achieve 7.32% better goal completion in social negotiations by strategically optimizing reward signals based on dialogue context.
Reinforcement learning can transform persistence diagrams into dynamic entities, balancing complexity reduction with the preservation of essential topological features.
VLMs struggle to provide accurate annotations in video games, revealing significant gaps in their understanding of dynamic environments.
Annotating video game datasets with VLMs can transform the way RL agents learn by simplifying reward extraction and conditioning.
A single lightweight adapter can enhance language model personalization without the need for per-user fine-tuning or costly forward passes, achieving robust performance across diverse user contexts.
LLMs can achieve 7.32% better goal completion in social negotiations by strategically optimizing reward signals based on dialogue context.
Uncertainty-guided feedback can dramatically improve the reliability of reward models in visual diffusion, leading to superior optimization and quality outcomes.
Parser-derived supervision can replace costly human annotations, achieving up to 93.8% parse success in Danish and 80% preference in native speaker comparisons.
Latent context in inverse reinforcement learning may obscure more than it reveals, as it can detract from performance by failing to capture genuinely hidden preferences in Arctic shipping.
Adapting supervision weights based on the evolution of divergence histories boosts reasoning performance in language models without extra computational overhead.
PSRS affects up to 56% of responses in LLMs, revealing a critical vulnerability in AI alignment that can lead to harmful outcomes.
Targeted citation strengthening can eliminate 100% of attack success rates in retrieval systems, revealing critical vulnerabilities in existing auditing methods.
CodeGrep slashes token usage and rounds by over 15% while maintaining high resolve rates, revolutionizing how LLM coding agents handle file retrieval.
Test-time self-correction can boost LLM accuracy by over 30% on challenging reasoning tasks without the need for external reward models.
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
Generative reward models can finally unlock their full potential in RL, leading to substantial performance improvements through innovative ranking strategies.
OPD$^2$ not only boosts multilingual math reasoning but also narrows the performance gap between English and Korean models, revealing the hidden potential of language-specific training signals.
Self-distillation conditioned on privileged information may lead to a model that is less capable of reasoning, as it optimizes for a misleading signal rather than task success.
OCSD reveals how isolating observation effects can lead to more effective token-level updates in reinforcement learning, outperforming traditional methods.
End-to-end training outperforms traditional decision-blind methods, achieving higher policy value while respecting capacity constraints in resource allocation.
Frontier LLMs struggle to match human expert performance, solving less than 57% of complex European executive tasks.
Calibrating guilt from human neural data enables AI agents to mimic human prosocial behavior more accurately than conventional reward shaping methods.
Reasoning-oriented LLMs may excel in Theory of Mind tasks not due to specialized abilities, but because they are more robust to variations in prompts and tasks.
Open-ended revisions in LLMs suffer from poor peer input, leading to significant declines in answer quality across various models and benchmarks.
Reversible action cycles enable self-verification in long-horizon planning, cutting state returning drift by 44% and boosting accuracy nearly 4x.
Rollout generation can be transformed from a static process into a dynamic, learning-driven strategy that adapts to policy changes in real-time.
PrivDPO achieves robust LLM alignment while maintaining privacy, outperforming traditional methods in balancing privacy and utility.
ABSeeker's innovative credit assignment method allows it to achieve performance levels comparable to much larger models, redefining expectations for long-horizon search agents.
SPOT redefines on-policy distillation by ensuring that probing decisions directly enhance downstream reasoning performance, not just teacher alignment.
State-Matched Routing and Contextualized Self-Distillation boosts task success rates by over 15% in complex interactive environments by aligning guidance with the agent's actual execution state.
TurnSight reveals that leveraging turn-level hindsight can dramatically enhance LLM performance in complex tool interactions, outperforming conventional reinforcement learning techniques.
The intrinsic geometry of GFlowNets reveals surprising insights into when temporal interactions can be ignored, fundamentally changing how we approach forward-policy training.
Latent Reward Registers enable real-time preference alignment in diffusion models, achieving up to 33x reduction in computational costs while enhancing accuracy and perceptual quality.
By refining VLM-derived reward signals with structural priors, SAFT transforms noisy feedback into a reliable guide for faster and more aligned policy learning.
SFT leads to task conflicts that can cripple multi-task learning, while RL's variance-limited updates enable seamless task coexistence.
Hybrid agents that combine LLM-driven planning with RL optimization achieve superior performance in complex decision-making tasks, outperforming traditional methods.
Utility misspecification can lead to significant performance drops in RL, but this new framework ensures robustness against such deviations, enhancing real-world applicability.
Rarity-aware credit redistribution can drastically improve reinforcement learning performance by ensuring that rare solutions receive the recognition they deserve.
Steering a language model's intertemporal preferences can induce significant shifts in decision-making, impacting how AI systems advise on delayed costs and benefits.
Adaptive sampling can cut human evaluation costs while boosting the accuracy of model rankings in NLP tasks.
CARE-X achieves a remarkable 94.0% accuracy in visual question answering, outperforming existing models by 6 percentage points while also enhancing report quality and spatial localization.
Forgetting can be drastically reduced by nearly 80% in continual learning scenarios without sacrificing performance on new tasks.
Consensus strength in reinforcement learning can make or break model performance—Hi-TTRL offers a solution that fine-tunes this critical factor for better outcomes.
StructPO achieves a coherent academic introduction in a single pass, outperforming traditional multi-stage workflows while maintaining competitive quality against advanced LLMs.
Trajectory-guided test-time sampling can enhance LVLM accuracy while sidestepping the resource burdens of traditional alignment methods.
CVPO redefines LLM training by leveraging value-variance to enhance reasoning accuracy and exploration, outperforming traditional methods.
EvoHIL achieves unprecedented improvements in robotic manipulation by dynamically adapting reward models and ensuring action coherence, setting a new standard for human-in-the-loop learning.
Sparse rewards can be effectively optimized without losing the advantages of fine-grained feedback through a novel two-stage training approach.
Token-level credit assignment in multi-turn language agents can be effectively enhanced by integrating teacher preferences through a novel self-distillation approach.
Reflecting on failed expert trajectories can boost reasoning performance more than tackling problems directly from scratch.
Despite advances in LLMs, they fail to effectively integrate user preferences over time, with accuracy rates stagnating around 39% even in ideal conditions.
ProCAVE achieves a remarkable improvement in video streaming efficiency by leveraging predictive modeling and DRL, setting a new standard for edge caching frameworks.
A new flow matching prior can drastically enhance humanoid motion tracking by guiding policy exploration with geometric insights from unordered pose data.
Position bias in LLM-based rerankers can lead to significant inconsistencies in preference rankings, undermining their reliability in recommendation systems.
GROW achieves a 22.7% reduction in word error rate while accelerating training by 2.9x, redefining efficiency in TTS reinforcement learning.
Skill-α outperforms traditional skill generation methods by leveraging a novel rollback reward mechanism, leading to significant improvements in agent performance on downstream tasks.
RoMeRL achieves an 80% reduction in the Cold-Q ratio while enhancing feedback density sixfold, revolutionizing how LLM agents manage memory and rewards.
Categorical value learning can significantly enhance the performance of PPO critics in reinforcement learning, leading to better calibration and lower variance in advantage estimation.
Bridging the gap between transient experience and persistent capabilities, SPEE enables LLMs to self-improve by evolving their knowledge through a unique experience distillation process.
Instruction-Conditioned Exploration boosts LLM performance by 5% on complex reasoning tasks by strategically diversifying training instructions.
Translating LTL to LTLf+ unlocks efficient finite automata techniques for a wide array of AI applications without sacrificing complexity.
Effective curling strategies can be learned entirely through self-supervision, matching expert heuristics without human-annotated data.
Optimizing multiple moments of failure probabilities can dramatically enhance LLM reasoning performance, outperforming traditional single-moment approaches.
IACM-RL reduces infinite loops and stale context errors by proactively managing dynamic user intents, setting a new standard for robust tool invocation.
AdaThinkV achieves 40.79% accuracy in video reasoning while using 22.7% fewer tokens than its strongest adaptive baseline, showcasing a breakthrough in token-efficient reasoning.
Misalignment between visual evidence and predicted timestamps in video grounding can lead to substantial performance drops, but CAVE effectively bridges this gap with boundary-specific rewards.
CoKL enables LLMs to learn new tasks without sacrificing previously acquired capabilities, striking a balance that traditional methods fail to achieve.
Competitive adversarial self-play fails to improve legal reasoning performance, revealing that the environment's verifiability is more crucial than the competition.
Hindsight critiques can transform failed trajectories into powerful learning signals, boosting search-augmented RL performance by nearly 40%.
Personalization in LLM agents is more complex than previously thought, with existing benchmarks failing to capture the dynamic nature of user preferences and their impact on task execution.
PCSD boosts reinforcement learning performance by 15.6 points over existing methods, demonstrating that persistent teacher signals can effectively guide agents through sparse reward landscapes.
Achieving a 91.51% acceptable-answer rate in enterprise question answering reveals the potential of staged adaptation to balance proprietary knowledge acquisition with general capabilities.
Incorrect answers reveal critical insights about LLM behavior, enabling a more refined understanding of model capabilities that binary scoring overlooks.
Self-evolving rubric rewards can dramatically enhance audio reasoning in models, outperforming traditional methods by adapting to the model's evolving capabilities.
Grounding preference learning in psychological and cultural constructs yields a significant boost in population alignment for LLMs, outperforming traditional demographic methods.
EviSD achieves state-of-the-art performance in question-answering tasks by leveraging privileged evidence, outperforming existing methods while maintaining efficiency in response generation.
Leading LLM investment advisors can significantly enhance long-term investor outcomes, revealing a critical gap in traditional evaluation methods.
Selective distillation can unlock critical learning signals in reinforcement learning, leading to significant performance gains in complex tasks.
Enhancing on-policy reinforcement learning can paradoxically reduce the diversity of successful behaviors, leading to a trade-off that challenges future trainability.
Multi-dimensional Evaluation-Verification Reward transforms multi-reference image editing by providing a structured approach to evaluate and enhance visual consistency, yielding superior results over existing models.
Combining RLVR and OPD through SAF not only prevents entropy collapse but also boosts performance across multiple benchmarks, revealing the potential for more effective reinforcement learning strategies.