Search papers, labs, and topics across Lattice.
1
0
3
0
Evaluator bias can dramatically alter the perceived safety of medical AI, with LLM judges showing a leniency that could misrepresent model performance. WHY_IT MATTERS: This insight challenges the reliability of current evaluation methods for medical AI, emphasizing the need for standardized assessment frameworks to ensure safety in clinical applications.