Search papers, labs, and topics across Lattice.
Affiliation:
3
0
6
4
Penalizing shifts in safety representations during reasoning fine-tuning can restore LLM safety without sacrificing performance, revealing a critical interplay between reasoning and safety in model training.
System prompts in commercial AI products are often a mixed bag, with 40% harboring instructions that can undermine user interests, revealing a critical gap in accountability.
A Qwen3-8B model, trained with a new SFT+RLAIF recipe on a challenging new benchmark, SWE-QA-Pro, beats GPT-4o in repository-level code understanding.