Search papers, labs, and topics across Lattice.
Affiliation:
6
0
11
0
Scalar metrics fail to capture the true diversity of AI-generated content, but diversity profiles offer a robust, multi-dimensional evaluation framework that reveals hidden biases.
Dynamic scaling in D2-ScaleAgent allows for adaptive retrieval and reasoning, leading to superior long document comprehension compared to traditional fixed workflows.
LOTUS accelerates RL convergence by over 20% and boosts success rates by up to 45% in unseen tasks, redefining how agents can generalize across diverse scenarios.
Unified multimodal models may excel in generation and understanding, but they often falter when reasoning about their own outputs, revealing hidden weaknesses in their capabilities.
A new benchmark reveals that many AI-generated videos fail to adhere to fundamental physical principles, challenging the reliability of existing quality assessment metrics.
PhyAI achieves up to 4.65x speedup in Physical AI tasks by unifying disparate inference processes into a single, efficient runtime.