Search papers, labs, and topics across Lattice.
We track OpenAI, DeepMind, Anthropic, and 17 other labs daily - with AI-powered summaries, trend charts, and a weekly digest.
We read everything so you don't have to. One email, zero noise.
This hands-on problem-solving tutorial provides both a rigorous algorithmic and practical introduction to multi-turn RL finetuning for LLMs, and covers state-of-the-art multi-turn RL finetuning algorithms, turn-level vs. trajectory-level reward design, and production grade monitoring for reward hacking detection.
Real-time character animation is now feasible with Wan-Animate-2, which achieves high-fidelity results without the pitfalls of traditional motion representation methods.
Test-time self-correction can boost LLM accuracy by over 30% on challenging reasoning tasks without the need for external reward models.
MicroEvo achieves a staggering 10.6x increase in search efficiency while improving Pareto-front quality by 36.2%, revolutionizing microarchitecture design exploration.
ShimGen not only matches but surpasses manually-designed protocols in consistency, revealing critical performance gains in heterogeneous memory systems.
World rehearsal enables LLM agents to internalize environment dynamics, leading to superior performance without the need for costly external interactions.
AMS reveals that safety training modifications can significantly alter the activation landscape of language models, impacting their compliance with safety protocols.
Smart-home agents struggle to differentiate between real commands and misleading ambient noise, with traditional detectors and MLLMs both failing in complementary ways.
SJRL not only overcomes collision challenges in multi-agent navigation but also adapts dynamically to real-world constraints, outperforming traditional methods in complex environments.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
GeniWorld achieves robust zero-shot generalization in robotic manipulation, outperforming traditional models even with minimal training data.
PaDoc achieves a remarkable 67.4-118% increase in valid-page throughput while maintaining top-tier parsing accuracy, revolutionizing document parsing efficiency.
We read everything so you don't have to. One email, zero noise.
MameLoshnLM not only excels in performance but also reveals the shortcomings of existing multilingual models in capturing the nuances of low-resource languages like Yiddish.
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
Operator-residual feedback slashes the rate of misleading score-only decisions from nearly 40% to under 2%, ensuring that autonomous agents make choices grounded in physical reality.
Reducing token overhead by pruning redundant communication edges allows multi-agent systems to achieve better performance without the computational burden.
F$^2$Agent achieves over 20% better annualized returns than existing models by dynamically capturing inter-modality dependencies and resisting market noise.
Rollout generation can be transformed from a static process into a dynamic, learning-driven strategy that adapts to policy changes in real-time.
Constraint-First Reasoning reveals that explicitly managing answer-space constraints can dramatically enhance the accuracy of mathematical problem-solving in language models.
Extending context in conversations can significantly amplify the risk of LLMs promoting delusional behaviors, challenging assumptions about model size and reasoning capabilities.
RTCF boosts the success of frozen VLA policies by leveraging past experiences without the need for retraining or extra GPU power.
Current instruction-based video editing models are far from satisfactory, revealing critical gaps in evaluation that could reshape the field.
Training LLM agents without expert supervision can lead to better performance and generalization across diverse environments.
We read everything so you don't have to. One email, zero noise.
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
The first publicly available dataset for early pregnancy fetal ultrasound screening could revolutionize automated diagnostics and standardization in prenatal care.
Formalizing the soundness of symbolic execution tools reveals that path-merging can significantly enhance software reliability without sacrificing correctness.
Fine-grained robot manipulation can be significantly enhanced by anticipating future wrist interactions through a novel task-conditioned modeling approach.
Output extrapolation in STEP-OPD enables a unified student model to surpass the performance of its specialized teachers across all evaluated tasks.
SFC redefines semantic understanding in spoken language tasks, achieving superior accuracy and adaptability in open-domain contexts.
Event-adaptive compression allows EvtGraph to outperform traditional models while maintaining efficiency, proving that less can be more in temporal data representation.
Pre-computed embeddings can outperform traditional supervised models in estimating global biomass, reshaping our approach to carbon stock monitoring.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.
SmartMage's dynamic modality orchestration reveals that tailored modality selection can significantly enhance 3D scene understanding performance.
By integrating fresh and delayed signals, this approach boosts viewer engagement and revenue while reducing model complexity by nearly 42%.
Modality Balance can be harnessed as a powerful form of privileged information, leading to substantial gains in reasoning performance for multimodal models.
We read everything so you don't have to. One email, zero noise.
Formal verification of a compiler for asynchronous dataflow could redefine the reliability and efficiency of parallel computing architectures.
PrivDPO achieves robust LLM alignment while maintaining privacy, outperforming traditional methods in balancing privacy and utility.
Path-level pretraining in MultiPathFormer leads to a dramatic 59% improvement in wireless propagation estimations, reshaping the landscape of wireless foundation models.
Self-distillation conditioned on privileged information may lead to a model that is less capable of reasoning, as it optimizes for a misleading signal rather than task success.
Argus achieves a 78% success rate on long-horizon reasoning tasks while using 21% fewer tokens in mature workflows, showcasing a revolutionary approach to agentic autonomy.
SpineSegDiff not only matches the best in vertebral segmentation but also reveals critical insights into degenerative disc conditions through uncertainty mapping.
Spectral fingerprints can differentiate between unique molecular structures with identical 2D connectivity, revolutionizing how we assess chemical similarity.
Diff-Symbo achieves unprecedented quality and diversity in text-controlled music generation, outperforming leading models by leveraging a novel latent diffusion framework.
SVI-DAG outperforms existing Bayesian methods by effectively quantifying uncertainty in causal inference while leveraging prior knowledge and edge dependencies.
ASTELD uncovers a critical gap in the autonomous AI landscape: no evaluated systems achieve both local-first deployment and enterprise-grade security.
AFD-Ledger reveals that optimizing deployment for AFD can drastically cut evaluation costs while exposing the nuanced performance dynamics between homogeneous and heterogeneous setups.
The most effective playback similarity metric, CLEWS, is also the least expensive to implement, revolutionizing evaluation strategies in music transcription.
We read everything so you don't have to. One email, zero noise.
Dynamic adaptation in vision-language models can significantly boost performance while cutting down computational costs.
H2S achieves a remarkable 48.54 mAP in audio-visual instance segmentation, setting a new benchmark in the field.
SG-TULA achieves competitive performance in sampling from complex non-convex distributions while offering explicit convergence guarantees that traditional methods fail to provide.
Generative reward models can finally unlock their full potential in RL, leading to substantial performance improvements through innovative ranking strategies.
TS-RAG redefines time series forecasting by effectively merging input data with retrieved sequences, leading to unprecedented accuracy improvements.
Co-optimizing network and ML parameters can accelerate training by up to 42%, unlocking new efficiencies in AI workloads.
Super-resolution techniques can significantly erase small white matter lesions in MRI scans, challenging assumptions about their reliability in clinical settings.
TFMs, despite being the leading approach for tabular predictions, fail to consistently represent joint distributions, raising concerns about their reliability in practical applications.
SEAM reveals that even accurate local predictions can mask significant global inconsistencies in scientific explanations, challenging traditional validation approaches.
Self-calibrating quantum fault tolerance can achieve provable efficiency, allowing for continuous error correction without the need for disruptive recalibrations.
Despite aggregate accuracy gains, multimodal LLMs often misuse visual tool-use, revealing a critical disconnect between visual input and causal influence on model outputs.
Action segmentation accuracy skyrockets with D-CLOT, achieving up to +12.7 F1 by resolving representation-prototype inconsistencies in unsupervised settings.
Visual context can dramatically enhance knowledge graph completion, as shown by ViSR-KGC's superior accuracy over traditional methods.
We read everything so you don't have to. One email, zero noise.
G-STEER refines user research queries with unprecedented efficiency, asking one-third as many questions while maximizing personalization and target coverage.
Game-hopping proofs can now be mechanized in Lean with unprecedented clarity and integration, enabling more robust cryptographic security verification.
RustGo prunes 78.49% of irrelevant paths, accelerating bug discovery in unsafe Rust code while uncovering 13 previously unknown vulnerabilities.
IcFuzz uncovers critical bugs in NVIDIA Isaac Sim that existing fuzzing methods completely miss, achieving over 200% code coverage.
Noise-aware residual correction boosts the realism of autoregressive audio-visual generation, tackling issues of identity drift and desynchronization head-on.
Achieving high-quality long video generation without any model retraining, Diff-VF redefines the capabilities of existing short-video diffusion frameworks.
Achieving seamless identity replacement in videos, Vorch-IR can handle multiple subjects and backgrounds without requiring precise pose matching.
Decoupling coordinate frame selection from box regression leads to a remarkable 11% accuracy boost in 3D visual grounding tasks.
G$^2$ARD-GS achieves up to 30x compression while enhancing image quality and preserving geometric fidelity, setting a new standard for 3D scene representation.
Achieving real-time audio-video generation at 27.12 FPS, Vorch-Streamer tackles the dual challenges of exposure bias and causal speech generation in long-form content.
Aligning robot scene geometry with ARGUS enables manipulation policies to learn 4-6 times faster from diverse viewpoints, transforming their generalization capabilities.
Overlapping electronic states in YbO lead to unexpected vibrational irregularities that challenge conventional understanding of molecular spectra.
We read everything so you don't have to. One email, zero noise.
Nonlinear modal synthesis can now be computed with reduced costs while retaining high fidelity in frequency control, revolutionizing simulations of musical instrument dynamics.
Standard Young tableaux can reveal nuclear-spin symmetry without complex projections, streamlining molecular spectroscopy analysis.
Novice users benefit significantly from explanations in product recommendations, while expert users remain unaffected by additional information complexity.
Cleo transforms conversational commerce by combining transparent ranking with controllable language generation, enabling users to make informed decisions without the pitfalls of LLM unpredictability.
Unlearnable perturbations can safeguard copyright by ensuring models learn irrelevant features, thwarting both unauthorized training and data leakage.
OPD$^2$ not only boosts multilingual reasoning in LLMs but also reveals that English-centric training can inadvertently skew responses toward English, raising questions about data balance in multilingual contexts.
Graph propagation over a temporal knowledge graph can predict clinical advancement with unprecedented accuracy, especially for cases lacking direct evidence.
Staging social interactions in trajectory prediction leads to remarkable gains in accuracy and consistency, reshaping how we approach agent behavior modeling.
Multi-layer circuit steering can achieve robust behavioral control in LLMs without sacrificing text quality, outperforming traditional single-point interventions.
Structured video-grounded semantics can boost synthetic IMU generation, leading to a 19.86% improvement in tail-class recognition over traditional methods.
MLLMs may signal hazards with over 95% accuracy, yet they struggle to identify the underlying causes, revealing a critical gap in proactive safety capabilities.
Current agentic AI systems fall short, with none demonstrating more than two out of nine critical maturity criteria, exposing a significant gap in their capabilities.
We read everything so you don't have to. One email, zero noise.
PSRS affects up to 56% of responses in LLMs, revealing a critical vulnerability in AI alignment that can lead to harmful outcomes.
Unconstrained decoding in dLLMs can lead to a staggering 90% collapse into answer-only outputs, highlighting a critical flaw in reasoning capabilities.
EchoPrompt reveals that by restoring latent prompts, we can significantly enhance the detection of LLM-generated text, achieving state-of-the-art results without any training.
SkillZip achieves a remarkable 3.46x compression ratio while preserving 99.2% of dependencies and 98.7% verifier reachability, revolutionizing how we manage agent skill libraries.
AgentExecutor outperforms existing methods by achieving up to 94% code coverage while slashing execution time by over 80%.
WasmMend achieves a remarkable 70% fix rate for discrepancies between WebAssembly and native binaries, showcasing the power of divergence-guided reasoning in automated repair.
JDomInO keeps Java code and domain models in sync, preventing the costly drift that undermines effective Domain-Driven Design.
Fine-tuned foundation models can significantly outperform human-inspired methods in compositional analysis, but at the expense of interpretability and generalization.
StreamMind achieves superior performance in streaming video understanding by effectively managing the trade-off between real-time interaction and long-term memory retention.
Pretraining on relevant scientific images can significantly boost quality assessment performance, outperforming larger datasets.
MultiMoQ achieves smoother viewport playback by increasing goodput and reducing latency, even under challenging network conditions.
Ambient temperature fluctuations can be harnessed to create robust adversarial attacks that consistently deceive multimodal perception systems.
We read everything so you don't have to. One email, zero noise.
Hierarchical post-training can significantly enhance robotic manipulation by enabling agents to better navigate complex tasks through effective subgoal decomposition.
U-Nets unexpectedly show greater robustness to resolution changes than anticipated, challenging assumptions about neural operator architectures in inverse imaging.