Search papers, labs, and topics across Lattice.
5
0
9
5
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
Independently scaling long-term memory in language models can yield better performance with fewer parameters than simply increasing model size.
MemSFT enables LLMs to gain specialized domain knowledge without sacrificing their general performance, effectively sidestepping the alignment tax.
Key contribution not extracted.
Stop wasting compute: PonderLM-3 learns to spend extra inference FLOPs only on the tokens that actually need them, outperforming fixed-step pondering methods.