Search papers, labs, and topics across Lattice.
This study integrates Mamba, a state space model, into the YOLO26 framework by introducing MambaPSA, a lightweight alternative to the conventional C2PSA block, and enhances the architecture with a bidirectional Vision Mamba (BiViM) module. The experiments conducted on the PASCAL VOC 2007+2012 dataset reveal that MambaPSA achieves a 2.9% reduction in parameters and a 12.1% decrease in FLOPs while improving CPU inference throughput by 17.6%, all with a negligible accuracy drop. Notably, the addition of the BiViM module at the P4 level results in a significant accuracy improvement of +0.9 mAP50:95, indicating a promising efficiency-accuracy trade-off in lightweight object detection models.
MambaPSA achieves a remarkable 17.6% boost in CPU inference speed while maintaining competitive accuracy in YOLO26, showcasing the potential of state space models in object detection.
State space models (SSMs), notably Mamba, have recently emerged as efficient alternatives to self-attention with linear computational complexity. We investigate the integration of Mamba into YOLO26, the latest non-maximum suppression (NMS)-free object detection framework, by proposing MambaPSA, a lightweight Mamba-based replacement for the C2PSA block at the end of the backbone. To complement this study, we additionally insert a bidirectional Vision Mamba (BiViM) module at the P3, P4, and P5 levels of the neck. Experiments on PASCAL VOC 2007+2012 show that MambaPSA reduces parameters by 2.9%, FLOPs by 12.1%, and improves CPU inference throughput by 17.6% (from 17 to 20 FPS) with negligible accuracy change (-0.1 mAP50:95), while the P4 BiViM placement yields the best accuracy gain (+0.9 mAP50:95). These results suggest that SSMs offer a favorable efficiency-accuracy trade-off when replacing attention-based blocks in NMS-free lightweight detectors.