Search papers, labs, and topics across Lattice.
Affiliation:
4
1
8
13
FairTPT not only improves fairness in vision-language models but also prevents catastrophic forgetting, setting a new standard for test-time adaptation in AI.
Even with corrupted human feedback, surprisingly tight guarantees for multi-agent reinforcement learning are possible.
Forget perplexity: DMAP offers a mathematically grounded, model-agnostic representation of text that unlocks new insights into generation quality, machine-generated text detection, and forensic analysis of synthetic data influence.
RLHF and DPO are surprisingly vulnerable to data poisoning, with even a small number of carefully crafted preferences capable of steering the learned policy towards a desired (potentially harmful) target.