SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
- URL: http://arxiv.org/abs/2609.08038v2
- Date: Wed, 09 Sep 2026 06:49:22 GMT
- Title: SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
- Abstract summary: We introduce SAFER-Activities, a dataset for fall detection and physical activity monitoring.<n>It comprises over 66 hours of video data captured by multiple cameras, with 85,310 action instances and frame-level annotations for 30 action classes.<n>We benchmark action recognition on SAFER-Activities with 2D and 3D skeleton models, RGB models with frozen backbones, and multimodal fusion strategies.
- Score: 0.4893345190925178
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely intervention in critical situations such as falls, particularly for mobility-challenged individuals. Existing datasets are often clip-based, lacking the frame-level detail needed to recognize actions online, as they unfold. To address this, we introduce SAFER-Activities, a dataset for fall detection and physical activity monitoring, with a dedicated subset for wheelchair use scenarios. It comprises over 66 hours of video data captured by multiple cameras, with 85,310 action instances and frame-level annotations for 30 action classes. We benchmark action recognition on SAFER-Activities with 2D and 3D skeleton models, RGB models with frozen backbones, and multimodal fusion strategies, and evaluate on in-lab, out-of-distribution, and cross-dataset test sets. Skeleton-based models generalize best under domain shift; fusing frozen RGB features with the skeleton stream improves in-domain recognition over the baseline CNN1D, most clearly on the wheelchair subset, but degrades out of distribution. Cross-dataset and qualitative evaluations confirm that models trained on SAFER-Activities transfer well to unseen environments and external fall data. To support research on robust fall detection and activity monitoring, we release the dataset and code at https://safer-activities.github.io/.
Related papers
- FEMOT: Multi-Object Tracking using Frame and Event Cameras [57.75021868843091]
Bio-inspired event cameras offer high temporal resolution and high dynamic range, providing complementary cues under extreme scenarios.<n>FEMOT is a large-scale RGB-event multi-object tracking dataset that covers diverse real-world scenarios and 14 challenging attributes.<n>FEMOTR is a multimodal tracking framework that decouples RGB and event features and fuses them in the frequency domain.
arXiv Detail & Related papers (2026-06-12T04:17:41Z) - Towards Video Anomaly Detection from Event Streams: A Baseline and Benchmark Datasets [71.53287557600177]
We take the first major step toward establishing event-based anomaly detection as a unified research direction.<n>We first construct multiple eventstream based benchmarks for video anomaly detection, featuring synchronized event and RGB recordings.<n>We then propose an EVent-centric Video Anomaly Detection framework, namely EWAD, with three key innovations.
arXiv Detail & Related papers (2026-03-26T03:33:33Z) - Concept-based Explainable Data Mining with VLM for 3D Detection [0.0]
This paper proposes a novel cross-modal framework that leverages 2D Vision-Language Models to identify and mine rare objects from driving scenes.<n>Our approach synthesizes complementary techniques such as object detection, semantic feature extraction, dimensionality reduction, and multi-faceted outlier detection.<n> Experiments on the nuScenes dataset demonstrate that this concept-guided data mining strategy enhances the performance of 3D object detection models.
arXiv Detail & Related papers (2025-12-05T07:18:45Z) - Real-Time Detection and Tracking of Foreign Object Intrusions in Power Systems via Feature-Based Edge Intelligence [4.60587070358843]
This paper presents a novel framework for real-time foreign object intrusion (FOI) detection and tracking in power transmission systems.<n>The framework integrates: (1) a YOLOv7 segmentation model for fast and robust object localization, (2) a ConvNeXt-based feature extractor trained with triplet loss to generate discriminative embeddings, and (3) a feature-assisted IoU tracker.<n>To enable scalable field deployment, the pipeline is optimized for deployment on low-cost edge hardware using mixed-precision inference.
arXiv Detail & Related papers (2025-09-16T17:17:03Z) - TimberVision: A Multi-Task Dataset and Framework for Log-Component Segmentation and Tracking in Autonomous Forestry Operations [2.0499240875881997]
We introduce the TimberVision dataset, consisting of more than 2k annotated RGB images containing a total of 51k trunk components.<n>We introduce a generic framework to fuse the components detected by our models for both tasks into unified trunk representations.<n>Our solution is suitable for a wide range of application scenarios and can be readily combined with other sensor modalities.
arXiv Detail & Related papers (2025-01-13T14:30:01Z) - Visual Context-Aware Person Fall Detection [52.49277799455569]
We present a segmentation pipeline to semi-automatically separate individuals and objects in images.
Background objects such as beds, chairs, or wheelchairs can challenge fall detection systems, leading to false positive alarms.
We demonstrate that object-specific contextual transformations during training effectively mitigate this challenge.
arXiv Detail & Related papers (2024-04-11T19:06:36Z) - Learning from Temporal Spatial Cubism for Cross-Dataset Skeleton-based
Action Recognition [88.34182299496074]
Action labels are only available on a source dataset, but unavailable on a target dataset in the training stage.
We utilize a self-supervision scheme to reduce the domain shift between two skeleton-based action datasets.
By segmenting and permuting temporal segments or human body parts, we design two self-supervised learning classification tasks.
arXiv Detail & Related papers (2022-07-17T07:05:39Z) - Self-supervised Pretraining with Classification Labels for Temporal
Activity Detection [54.366236719520565]
Temporal Activity Detection aims to predict activity classes per frame.
Due to the expensive frame-level annotations required for detection, the scale of detection datasets is limited.
This work proposes a novel self-supervised pretraining method for detection leveraging classification labels.
arXiv Detail & Related papers (2021-11-26T18:59:28Z) - ETRI-Activity3D: A Large-Scale RGB-D Dataset for Robots to Recognize
Daily Activities of the Elderly [6.597705088139007]
We introduce a new dataset called ETRI-Activity3D, focusing on the daily activities of the elderly in robot-view.
The proposed dataset contains 112,620 samples including RGB videos, depth maps, and skeleton sequences.
We also propose a novel network called four-stream adaptive CNN (FSA-CNN)
arXiv Detail & Related papers (2020-03-04T07:30:16Z) - Stance Detection Benchmark: How Robust Is Your Stance Detection? [65.91772010586605]
Stance Detection (StD) aims to detect an author's stance towards a certain topic or claim.
We introduce a StD benchmark that learns from ten StD datasets of various domains in a multi-dataset learning setting.
Within this benchmark setup, we are able to present new state-of-the-art results on five of the datasets.
arXiv Detail & Related papers (2020-01-06T13:37:51Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.