Search papers, labs, and topics across Lattice.
2
0
4
0
Evaluation reference choices can drastically alter model performance and rankings, challenging the reliability of current CXR machine learning assessments.
Multilingual LLM evaluators can misjudge content, favoring lower-resource languages and potentially allowing harmful material to slip through safety filters despite high accuracy scores.