Search papers, labs, and topics across Lattice.
7
0
11
2
StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop, achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in real time, and resolves the tension between deep deliberation and latency via Think-While-Speaking.
This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
UI-Venus-2 achieves unprecedented environment coverage and task reliability, paving the way for dependable multimodal GUI agents in real-world applications.
Achieving optimal watermarking trade-offs is now a guided optimization problem rather than a heuristic guessing game.
Forget specialized architectures: StepAudio 2.5 proves a single audio-language foundation, shaped by RLHF, can dominate ASR, TTS, and real-time dialogue simultaneously.
Planning in latent space, rather than directly in action space, overcomes the exponential decay in feasible trajectories, unlocking efficient long-horizon decision making for embodied agents.