Search papers, labs, and topics across Lattice.
This paper introduces AMPBench-MT, a comprehensive benchmark designed to evaluate antimicrobial peptides (AMPs) by integrating multiple assay-derived metrics such as potency, safety, and spectrum within a homology-controlled framework. The study reveals that high performance in binary recognition does not correlate with actual assay behavior across 161 evaluations, highlighting the limitations of existing benchmarks. Key findings indicate that while frozen protein-language-model embeddings cluster around pMIC errors, traditional regression methods provide more reliable predictions, underscoring the need for a shift towards endpoint-aware evaluations in AMP research.
High binary recognition performance in AMP models fails to predict real-world assay outcomes, revealing critical gaps in current evaluation methods.
Computational AMP discovery is often evaluated through AMP/non-AMP recognition, yet follow-up decisions depend on assay-derived evidence such as target-species potency, hemolysis, toxicity, and selectivity. Existing AMP and peptide benchmarks cover binary recognition, multilabel annotation, assay regression, or broader peptide-model comparison, but they do not jointly place AMP recognition, species-conditioned potency, spectrum, safety-facing proxy endpoints, and cross-endpoint behavior within one sequence-homology-controlled protocol. To address this problem, we introduce AMPBench-MT, a provenance-preserving benchmark that standardizes canonical peptide records and organizes them into binary recognition, species-conditioned pMIC regression, and endpoint-specific potency and safety-facing readouts. Across 161 endpoint-specific model evaluations, high binary performance does not reliably indicate assay-endpoint behavior. Frozen protein-language-model embeddings form the leading pMIC error cluster, while graph and classical regressors remain close. Spectrum labels further reveal that PR-oriented metrics can be misleading under scarce observed negatives, whereas low-toxicity, HC50 hemolysis, and selectivity expose smaller but more assay-facing signals. AMPBench-MT shows that AMP evaluation should move beyond recognition leaderboards toward endpoint-aware evidence auditing. Our proposed benchmark is available at https://huggingface.co/datasets/ZihengZhou06/AMPBench-MT.