Search papers, labs, and topics across Lattice.
Affiliation:
7
0
11
12
Urban spatial optimization can be solved far more effectively by framing it as an interactive code-and-eval sandbox with decomposed build-and-refine RL than by relying on rigid, task-specific RL solvers.
Even top-tier vision-language models consistently invert a subject's left and right, exposing a blind spot in camera- versus subject-centric spatial reasoning that rubric-guided GRPO over anatomical priors can finally resolve.
DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.
TCS not only generates sound test cases but also adapts to the model's weaknesses, leading to a marked improvement in code generation performance.
By separating control from data flow, this approach ensures that prompt optimization enhances performance without risking protocol integrity.
Explicit context compilation boosts LLM performance on in-context learning tasks, lifting accuracy from 15.4% to 21.4% on complex benchmarks.
Stop wrestling with finicky evaluation codebases: One-Eval lets you specify LLM evaluation tasks in natural language and automatically executes them end-to-end.