Search papers, labs, and topics across Lattice.
This paper introduces EventKitchen, a novel stereo event camera dataset specifically designed to capture human cooking activities in natural settings, addressing the gap in human-centric event-based perception datasets. Collected egocentrically from 10 participants across 13 kitchens, the dataset includes 5.5 hours of stereo event recordings along with synchronized RGB, depth, and IMU data, annotated for over 10,000 action segments and bounding boxes. The authors demonstrate the dataset's utility by training baseline models for action recognition, object detection, and stereo depth estimation, highlighting its potential to advance neuromorphic vision research beyond traditional automotive applications.
EventKitchen reveals the complexities of real-world cooking activities, setting a new benchmark for event-based perception that challenges existing datasets focused on scripted actions.
Event cameras, also known as neuromorphic cameras, have gained significant attention in recent years due to their high temporal resolution, high dynamic range, and low power consumption. While many studies and datasets in neuromorphic vision have focused on automotive and drone applications, human-centric daily-life scenarios remain largely underrepresented, despite their importance for developing and benchmarking event-based perception systems. Moreover, the few existing event-based human activity datasets are typically recorded with scripted human actions, limiting their ability to capture natural human behaviors. In this paper, we introduce EventKitchen, a large-scale stereo event camera benchmark dataset of human cooking activities in the kitchen. EventKitchen is egocentrically collected from 10 participants in 13 diverse kitchens, where the participants wear a helmet with multiple sensors and naturally perform cooking activities, without any scripted actions. EventKitchen comprises 5.5 hours of stereo event recordings with synchronized RGB, depth, and IMU data. We provide human annotations for 10,762 action segments and 13,482 bounding boxes. We train baseline models on EventKitchen to perform multiple event-based tasks, including action recognition, object detection, and stereo depth estimation. By capturing natural, real-world human activities, EventKitchen establishes a challenging benchmark for neuromorphic vision beyond autonomous driving.