Search papers, labs, and topics across Lattice.
Tencent
2
0
2
ARGUS achieves unprecedented fine-grained observability in 10,000+ GPU clusters with less than 2% overhead, revolutionizing performance diagnosis at scale.
Forget agonizing over checkpointing and restarts: LiveR slashes LLM training downtime from minutes to seconds by hot-swapping model state between parallel training worlds.