Search papers, labs, and topics across Lattice.
Affiliation:
2
0
5
Full rationale supervision burns up to 254× more tokens without guaranteeing reliability, whereas isolating representations with high rationale-boundary sensitivity selectively immunizes medical LLMs against choice-order brittleness.
Forcing clinical agents to run exhaustive diagnostic workups can actually degrade diagnostic accuracy from 28.3% to 34.3%, exposing the critical need for statistically bounded stopping rules over agent self-termination.