Search papers, labs, and topics across Lattice.
This study investigates the impact of inference-time sampling strategies on multimodal protein language models (pLMs) through a three-stage framework, focusing on vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search. The findings reveal that default inference protocols are often suboptimal, and task-oriented sampling can significantly enhance performance across various tasks without altering model parameters. Notably, the research challenges previous assumptions about base model capabilities, providing new insights into the exploration-exploitation trade-off in pLMs.
Task-oriented sampling strategies can substantially elevate multimodal protein language model performance, revealing that default protocols are often inadequate.
Multimodal protein language models (pLMs) learn joint protein sequence-structure distributions, and their generation performance should also depend critically on inference-time sampling strategies. Yet prior work has focused more on model training than on how inference-time strategies behave. In this paper, we establish a three-stage investigation framework to empirically study the inference design space of multimodal pLMs across three representative pLMs and four fundamental tasks. We evaluate vanilla sampling, task-specific classifier-free guidance, and reward-guided beam search on multimodal pLMs, corresponding to controls over sampling distributions, per-step logits, and parallel trajectories. Throughout the complementary advancements centered on exploration-exploitation trade-off, we (1) reveal the suboptimality of default inference protocols and identify task-oriented sampling preferences; (2) observe substantial quantitative gains across tasks, consistently boosting the upper bound performance of multimodal pLMs without updating model parameters; (3) derive conclusions about base models that differ from prior consensus.