Search papers, labs, and topics across Lattice.
10
0
8
6
RayOrch is presented, a programming model and distributed execution engine that preserves parent child relations throughout execution and reduces end to end time by 13.1 percent versus Ray Data and 29.0 percent versus Daft on MinerU, and by 16.0 percent versus Ray Data on Docling.
DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
Agents excel at selecting knowledge sources but struggle with actual task completion, achieving only 56.1-75.3% accuracy in answers despite near-perfect routing.
Even state-of-the-art models only achieve pass rates below 60% on a new benchmark that spans 1,431 diverse tasks, exposing critical weaknesses in general AI capabilities.
GraspLLM achieves unprecedented zero-shot generalization on Text-Attributed Graphs, outperforming existing methods by effectively merging graph structure with LLM semantics.
LLMs can now be benchmarked for their ability to prepare training data, revealing that a new evaluation metric outperforms traditional methods in predicting downstream utility.
Curriculum-aligned training can dramatically improve educational LLM performance, revealing that existing models are underprepared for structured knowledge tasks.
Stop reinventing the wheel: OpenWorldLib offers a unified framework and codebase for advanced world models, finally bringing standardization to a fragmented field.
DataFlex makes data-centric LLM training dramatically easier, unifying disparate methods for data selection, mixing, and reweighting into a single, efficient, and reproducible framework.
Stop wrestling with finicky evaluation codebases: One-Eval lets you specify LLM evaluation tasks in natural language and automatically executes them end-to-end.