Search papers, labs, and topics across Lattice.
2
0
4
2
General-purpose helpfulness metrics fail to reliably signal effective pedagogy in LLM tutoring, revealing a critical gap in evaluation methods.
Even the best LLMs struggle to maintain correct intermediate states when solving university-level STEM problems, often taking more steps than necessary and accumulating errors along the way.