Search papers, labs, and topics across Lattice.
This study introduces a comprehensive benchmark for fine-grained isolated handshape recognition in sign language, utilizing the Hamburg Notation System (HamNoSys) to create a dataset of 144,000 RGB images across 160 handshape classes. Evaluating various model architectures, including ResNet-18 and ViT-B/16, the research highlights the challenges of generalizing recognition to unseen participants, revealing a significant drop in performance during leave-one-subject-out evaluations. The findings establish reproducible reference performance metrics, paving the way for advancements in computational sign-language transcription and recognition technologies.
Recognition accuracy for isolated handshapes drops significantly when models encounter unseen signers, highlighting a critical gap in current methodologies.
Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited. This work introduces a benchmark grounded in the language-independent Hamburg Notation System (HamNoSys). Methods: A balanced dataset of 144,000 RGB images was collected from 15 participants for 160 handshape classes defined by the official HamNoSys 4 Handshapes Chart. ResNet-18 and ViT-B/16 were evaluated as appearance-based models, while a graph convolutional network and XGBoost were evaluated from hand landmarks. Both a class-stratified subject-dependent split and a 15-fold leave-one-subject-out (LOSO) protocol were used. The same model families were additionally assessed on LSWH100 and ASL Fingerspelling Dataset A for external context. Results: The subject-dependent benchmarks established reproducible reference performance across all four model families, whereas LOSO evaluation exposed a substantial reduction when recognition was required to generalise to unseen participants. On ASL Fingerspelling Dataset A, mean LOSO top-1 accuracy ranged from 82.20% to 87.40%. Conclusion: The documented acquisition, curation, and complementary evaluation protocols pro-vide a reproducible resource for fine-grained isolated-handshape research and for developing more accessible sign-language technologies.