Search papers, labs, and topics across Lattice.
Affiliation:
8
0
7
3
EXPL-FR reveals that you can achieve interpretable face recognition without any text training, using a simple adapter to bridge vision and language spaces.
Transferable presentation attack representations can be learned without relying on facial content, achieving high performance across standard benchmarks.
Adapting foundation models with synthetic data can dramatically enhance face recognition accuracy, with some methods even surpassing traditional baselines.
Introducing register tokens transforms ViT attention maps from opaque artifacts into clear, interpretable structures, boosting face recognition performance.
Exiting from intermediate layers of Vision Transformers can yield significant speedups in face recognition with minimal accuracy loss, revolutionizing deployment on resource-constrained devices.
Face recognition systems, beware: DCMorph's dual-stream diffusion morphs achieve unprecedented attack success rates while remaining stealthy to existing detection methods.
Forget relying solely on final-layer features: intermediate layers in Vision Transformers hold untapped potential for boosting face image quality assessment.
Turns out, your pre-trained face recognition ViT already knows which faces are high quality, just by looking at the attention maps.