Search papers, labs, and topics across Lattice.
3
0
4
11
Regrading model outputs can shift correctness labels by 9% and double the performance spread, revealing critical flaws in conventional evaluation methods for LLM code generation.
TSFMs can forecast heart rate variability from consumer wearables with unprecedented accuracy, outperforming traditional methods without any fine-tuning.
Epic-organized LLM-generated Gherkin scenarios are rated significantly higher in quality than those generated by a naive baseline, despite similar semantic coverage.