Search papers, labs, and topics across Lattice.
9
0
15
3
StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.
This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
GeoTrussRover combines an electrically actuated VGT, a wheeled base, and contact-semantic morphology planning and control, which stores task coordination in a hyper-redundant, load-bearing morphology and reuses it during locomotion.
Shrinking a 102M-parameter nnU-Net by 81脳 costs barely 2.5% in segmentation Dice while actually boosting lesion-level detection F1 by over 5 points.
Analog over-the-air model aggregation can match the theoretical $\mathcal{O}(1/\sqrt{T})$ convergence of ideal FedAvg without requiring instantaneous channel state information or strict phase alignment.
Shifting the focus from topical relevance to answerability, CLEAR reveals that many conversational retrievers miss the mark by overlooking critical answer-supporting passages.
H2Table achieves a remarkable 22.88% improvement in complex table reasoning, revealing the power of hierarchical hypergraph representations in LLMs.
Traditional trajectory-based training collapses policy diversity, but DART-SD's topology-aware approach ensures agents can explore optimal paths without penalty.
Even without retraining, a simple dual-system approach can significantly boost the performance of self-supervised talking head forgery detectors by refining the ordering of uncertain samples.