Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.
Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.
It is demonstrated that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and Agent as Policy (AGP) is introduced, which places task planning and execution under the agent's control.
MaP-WAM is introduced, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history.
Latent Interface Training (LIT), a framework-agnostic two-stage strategy that first establishes a spatial-goal-conditioned action prior without images, then constrains visual conditioning through a pose-supervised latent interface, improves overall LIBERO-Plus success while preserving or improving average LIBERO success.
Dynin-Robotics achieves competitive performance on LIBERO and zero-shot LIBERO-Plus, together with a 78.4% average success rate across four manipulation conditions on a Franka Research 3 robot.
It is demonstrated that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training and Agent as Policy (AGP) is introduced, which places task planning and execution under the agent's control.
MaP-WAM is introduced, a Memory-as-Plans framework that decomposes memory-dependent world-action modeling into memory-grounded planning and plan-conditioned execution, and uses long-term multimodal episodic context as planning-time evidence rather than repeatedly conditioning the executor on the full history.
Modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask, suggests modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.
Experimental results demonstrate that DRG-MAPPO achieves a state-of-the-art win rate of 87%, suggesting that the framework effectively balances relational modeling, interpretability, and optimization stability for cooperative air combat.
This work studies a route-conditioned order of unavoidable stages that every successful executor must traverse, recoverable from offline trajectories and belonging to none of them: a route-conditioned order of unavoidable stages that every successful executor must traverse.
ActSafeGuard is introduced, a differentiable and training-aligned safeguard layer for flow-matching based policies that integrates hard action feasibility into policy learning, not merely treating safety as an inference-time external component.
DeFiFusion is presented, a dual-modal PMA detection framework that closes this gap by jointly modeling transaction events and smart contract semantics within a unified pipeline, and proposes a Dual-Modal Projection-Fusion Transformer with T5-style relative positional encoding.
A reproducible pipeline that converts multi-dataset SMPLX motion into executable loco-manipulation behavior for the Galaxea R1 Pro wheeled humanoid and provides a complete bridge from human motion data to physically trackable wheeled-humanoid loco-manipulation rather than a visualization-only retargeter is presented.
OststaDiff is presented, a decomposed diffusion-policy framework with a lightweight obstacle-aware visual encoder that extracts a structured target-obstacle-background representation, enabling the downstream alignment policy to generate end-effector trajectories toward a target-centered bottleneck pose while reasoning about surrounding obstacles.
A novel Dual-Latent Space Reinforcement Learning (DLSRL) framework, which complements initial-noise steering with representation-level control inside the frozen generator and effectively accelerates online robot policy adaptation and achieves competitive performance.
Dist-GPRL is presented, a distance-aware and safety-guided reinforcement learning framework for structured robot skill adaptation that sequentially adapts overlapping local windows of sparse trajectory via-points rather than modifying the complete skill at every policy step.
This paper investigates freehand sketching as an end-user programming interface for specifying robot swarm geometries and supports freehand sketching as an intuitive interaction abstraction for human-swarm collaboration without requiring robotics or programming expertise.
2AM shows that task memory can remain Agent-side and shows that Action Model capability depends not only on what the policy has learned, but on how precisely the Agent can steer it, and isolates this question through a deliberately constrained design.
CAP is proposed, a single-stage humanoid locomotion policy that recovers this signal with a perceptive world-model encoder trained as a learned denoiser to reconstruct clean depth from a corrupted input, together with a co-active proprioceptive variational encoder that supplies depth-free body-state information.
GeoTrussRover combines an electrically actuated VGT, a wheeled base, and contact-semantic morphology planning and control, which stores task coordination in a hyper-redundant, load-bearing morphology and reuses it during locomotion.
Experiments on multi-agent LTLDiff manipulation tasks demonstrate improved task success rates compared to the baseline and demonstrate the effectiveness of LTLDiff for coordinated multi-agent manipulation.
This research presents an embodied control approach based on real-time task Jacobian estimation of the combined hand and object system on the physical robot and showcases an alternative to compute- and data-heavy approaches such as RL and IL for achieving dexterous manipulation through computationally simple and data-efficient algorithms.
UniMPA, a Unified Memory-Prediction-Action model that addresses transition ambiguity by modeling the intended future state evolution through a shared action-grounded transition interface, is proposed.
Attention-DP3 is proposed, a spatially object-aware 3D diffusion policy that injects object-level geometric cues via attention while keeping the DP3 diffusion backbone unchanged and shows consistent improvements over DP3, achieving state-of-the-art performance across benchmarks.
Cross-channel knowledge distillation allows a tiny 123K-parameter model to decode five-finger motor intent from stroke survivors using only four surface EMG channels without sacrificing multi-label precision.
Identical physical AI behaviors often mask fundamentally incompatible developmental origins, yet every major paradigm from generative world models to morphological co-design reduces to just seven core formation sources.
Monolithic federated updates are a fundamental bottleneck for embodied AI: decoupling vision, language, and action streams cuts uplink payloads by 96% and beats standard FedAvg by 22 percentage points under real-world wireless interference.
This work evaluates knee-angle estimation under a missing-ankle-keypoint condition and test a first-order temporal interpolation scheme as a recovery mechanism to recover a critical missing joint, without resorting to learned reconstruction models.
TANGO enables humanoid robots to navigate cluttered spaces with unprecedented language-guided precision, achieving state-of-the-art performance without real-world training.
Real-time object detection and tracking in autonomous racing can be achieved with a multi-modal fusion approach that significantly improves performance under challenging conditions.
Ostrich is presented, a GPU-accelerated rigid-body simulator that resolves hard contacts and friction with non-smooth Newton iteration at large timesteps, and differentiates the converged residual via the implicit function theorem, reusing the forward Schur complement to compute the adjoint at O(1) memory per timestep.
This paper proposes SUccessor-to-Novelty (SUN), an indicator derived from successor value functions to identify goals that are both novel and reachable and presents an adaptive goal-selection strategy that leverages these properties, and an accurate yet lightweight pseudocount to avoid the overhead of classic methods.
DeCAL achieves a remarkable 71% success rate in dexterous manipulation tasks by seamlessly integrating tactile sensing with advanced vision-language-action modeling.
BIFTA draws on the brain's rapid sensory adaptation mechanism to adapt a frozen encoder to an unknown tactile sensor from a small labeled support set, and it substantially improves adaptation to unknown sensors.
Egocentric motion forecasting fails without explicit spatial goals; anchoring residual flow matching to predicted continuous 3D interaction targets eliminates structural drift and aligns full-body trajectories with human intent.
Closed-loop world models degrade when unexecuted rollouts pollute memory: strictly separating planning timescales and barring latent states from factual feedback drops 8-second multi-agent prediction error to 1.20 m.
Structuring heterogeneous multi-robot coordination into a modular search-handoff-verify protocol skyrockets zero-shot VLM search success from 8.6% to 55.7% on complex urban benchmarks without requiring any model fine-tuning.
This work introduces Chunk-Aligned Semantic Distillation (CASD), which derives semantic targets for entire action chunks from the current observation, robot state, and task instruction, and freezes the generator to train a policy conditioned on its predictions.
Hallucinated objects and unexecutable actions in embodied planning don't require fine-tuning to fix: pairing a frozen VLM with inference-time symbolic masking and an HMM state lookahead guarantees feasible, grounded plans out of the box.
This work introduces a reinforcement learning approach that forgoes predefined plans entirely, instead generating construction sequences adaptively as the structure is built, and evaluates the algorithm, HSAC, against the prior method hybrid-PPO (HPPO), demonstrating significantly higher asymptotic performance and good sample efficiency.
Offline JSTC slashes computation time and joint motion while ensuring non-revisiting coverage paths for redundant manipulators.
Transforming noisy visual observations into reliable spatial guidance, AeroBelief achieves unprecedented navigation success rates for UAVs in complex environments.
Hard safety guarantees no longer require hours of offline pre-computation: an online, interval-based verification pipeline brings formal reachability directly into sampling-based MPC while eliminating over 99% of trajectory violations.
More accurate 3D scene completion doesn't reliably improve autonomous robot exploration—unsupported neural hallucinations quietly sabotage navigation planners unless continuously gated against real-time sensor observations.
ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
Monocular camera localization within dense LiDAR maps no longer requires complex dual-encoder architectures: converting both inputs to unified depth representations lets a single vision foundation model outperform specialized cross-modal baselines under extreme environmental shifts.
Multimodal LLMs can cleanly disentangle camera ego-motion from real-world scene dynamics without any drone telemetry or pose sensors, using just four optical-flow-derived residual tokens per slice to unlock robust aerial video reasoning.
Whether a benchmark can rank models is an empirical property is an empirical property, and three checks are given that establish it.
Copying human head-neck kinematics leaves over 60% of reachable manipulation space unobservable; independently actuating onboard RGB-D cameras surges visible-reachable coverage to 97% while cutting manipulation energy by 19%.
HiBRIDGE combines strong predictive performance with a structured decision process that supports more interpretable explanations of robot behaviour, and is demonstrated to be feasible for autonomous real-time group interaction.
This work presents the first synergistic framework that tightly couples hierarchical memory with proactive tool invocation in a closed reasoning loop, and reveals the complementary roles of hierarchical memory and tool strategies.
Locomotion recognition models collapse from 93% to 68% F1 precisely when assistive exoskeletons need them most: during movement transitions in clinical populations.
DYAD (DYadic Assistance Dataset), a synchronized multimodal record of human-human assistance during gearbox assembly, is introduced, a linked interaction structure spanning help seeking, intervention choice, execution, and outcome under egocentric and workspace sensing.
Physical robot kinematics can double as an out-of-band data link, letting agents transmit arbitrary bitstreams purely through camera-observable action noise without degrading policy performance or touching an RF transceiver.
Manual and teleoperated intraocular instrument motion was compared with the trocar constraint, the instrument, the eye model and the tracking source common to both conditions, so that the control interface was the only factor varied.
A single IMU is all that is needed to diagnose structural punctures in soft actuators and dynamically reconfigure inflation pathways to prevent catastrophic loss of actuation force.