Search papers, labs, and topics across Lattice.
This paper introduces Always Alert, a smartphone-based distress detection system that utilizes the device's microphone to monitor vocal expressions of fear, such as screaming and crying, in real-time. Employing a two-stage supervised learning framework with support vector machines, the system effectively balances high distress detection rates with a manageable false alarm rate, allowing users to maintain their daily activities without significant interruption. The framework demonstrates robust performance in diverse audio environments, achieving a detection rate that translates to minimal user overhead, approximately equivalent to one social media post every 3 to 4 hours.
Always Alert can detect distress signals with minimal disruption, achieving high accuracy while users go about their daily lives.
We investigate an unobtrusive and $24\times7$ human distress detection and signaling system, Always Alert, that requires the smartphone, and not its human owner, to be on alert. The system leverages the microphone sensor, at least one of which is available on every phone, and assumes the availability of a data network. We propose a novel two-stage supervised learning framework, using support vector machines (SVMs), that executes on a user's smartphone and monitors natural vocal expressions of fear---screaming and crying in our study---when a human being is in harm's way. The challenge is to achieve a high distress detection rate while ensuring that the false alarm rate is a manageable overhead, while a typical smartphone user goes about living life as usual. We train the learning framework with carefully selected audio fingerprints of distress and of varied environmental contexts. The audio is used to tune the learning framework to obtain a desirable distress detection rate and false alarm rate (FAR). The ability of the proposed framework to detect distress in rather challenging audio environments is demonstrated. Exploiting the time contiguous nature of false alarms further allows us to reduce the FAR. We show the feasibility of using our framework anytime and anywhere by testing it over many hours of audio fingerprints recorded by volunteers on their smartphones, as they went about their daily routines. We are able to achieve high distress detection rates at an average overhead that is equivalent to about 1 facebook post every 3 to 4 hours.