Search papers, labs, and topics across Lattice.
20
0
19
0
Current MLLMs excel at visual reproduction but falter in generating the necessary data semantics and interaction logic for coordinated multi-view interfaces.
Reciprocal cross-stream addressing in Dual Attention Residuals leads to significant improvements in Transformer performance without sacrificing depth-wise diversity.
Shared descriptive patterns in skill libraries can obscure task-relevant signals, but SkillSight reveals and calibrates these biases to boost retrieval accuracy by over 20%.
Achieving over 89% accuracy in attributing generated images to their source models, DNA redefines the landscape of digital forensics in image generation.
Instruction-level feedback from audio-aware LLMs can drastically enhance the accuracy of multi-event audio generation, bridging a critical gap in current models.
Achieving over 2.2x inference speedup in Vision Transformers while maintaining accuracy reveals the untapped potential of hardware-software co-design in optimizing model performance.
Stealthy attacks in multi-agent systems can be detected more effectively without relying on explicit interaction graphs, leading to a significant boost in detection accuracy.
AgenticPD redefines physical design optimization by enabling stage-aware decision-making, drastically cutting costs and improving timing results.
KAT-Coder-V2.5 outperforms existing models in agentic tool-use, showcasing a new paradigm for autonomous coding agents within executable environments.
Achieving a 90% reduction in data volume without sacrificing accuracy, this integrated optoelectronic architecture revolutionizes robotic visual inspection.
Achieving over 95% success in real-world robotic tasks after just 1.5 hours of training, this model-agnostic framework could redefine the deployment of VLA models in industry.
OpenPCC achieves robust privacy for LLMs without the constraints of proprietary hardware, paving the way for more secure AI applications.
MiniMax-M2 proves that massive parameter counts don't always translate to better agentic performance; strategic activation of a smaller subset can unlock frontier-level intelligence.
Achieve real-time, proactive video understanding with StreamOV, which uses bounded memory and a novel response trigger to overcome the limitations of offline methods.
Transfer learning from a large, pre-trained speech synthesis model unlocks high-quality Tibetan TTS, even with limited Tibetan-specific data.
Current evaluation metrics for trajectory inference can mislead researchers, but functional KL divergence offers a clearer, more reliable comparison of methods in sparse data conditions.
DQPOPE achieves the same sample efficiency as traditional OPE methods while providing a comprehensive return distribution, leading to significantly more accurate policy evaluations.
LIME leverages LLM-generated narratives to create a scalable surgical dataset, but SurgLIME's innovative approach ensures that noisy text doesn't compromise model performance.
Speculative decoding can be sped up by >2x without sacrificing accuracy by rescuing previously rejected tokens that are semantically valid but lexically different.
Achieve flicker-free, scalable reconstruction of long dynamic videos by blending the best of stream and clip-based Gaussian Splatting.