Search papers, labs, and topics across Lattice.
Affiliation:
3
0
4
WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
Idiolectal paraphrasing transforms how models learn reasoning by allowing them to express complex thoughts in their own unique language, leading to significant performance gains.