Physical activities enable scalable foundation modelling for broad-spectrum health prediction

· Source: cs.LG updates on arXiv.org · Field: Technology & Digital — Artificial Intelligence & Machine Learning, Internet of Things (IoT) & Connected Devices, Data Science & Analytics · Depth: Expert, extended

Summary

StepFM is a novel foundation model designed for broad-spectrum health prediction, utilizing only low-dimensional step counter data. Pre-trained on 141.12 million minute-level step-count observations, this model, with 3.4 million trainable parameters, offers a practical, privacy-preserving, and computationally efficient alternative to traditional high-frequency sensor-based approaches. StepFM employs a scalable pre-training framework that captures temporal dynamics and behavioral patterns from large-scale step sequences. It achieves a mean AUROC of 0.7318, outperforming existing foundation models like NormWear (0.7079) and TimeSiam (0.6704) across 20 of 21 health risk prediction tasks. The model demonstrates strong generalizability across diverse sensing devices, geographical regions, and novel disease types, including multiple-sclerosis-related endpoints unseen during pre-training.

Key takeaway

For AI Scientists and Machine Learning Engineers developing health monitoring systems, StepFM demonstrates that focusing on low-dimensional, privacy-preserving data like step counts can yield highly scalable and generalizable foundation models. You should consider integrating multi-scale temporal modeling and hierarchical activity phenotype alignment in your pre-training frameworks to capture nuanced behavioral patterns. This approach can significantly reduce computational overhead and privacy concerns associated with high-frequency raw sensor data, enabling broader deployment and more accessible health inference across diverse populations and devices.

Key insights

StepFM uses low-dimensional step data to create a scalable, privacy-preserving foundation model for broad-spectrum health prediction.

Principles

Method

StepFM uses a log-scaled tokenizer and temporal rhythm encoding for hourly step data. A dual-stream Step-Mamba encoder, modulated by minute-level CNN features via FiLM, is pre-trained with next-token prediction and hierarchical activity phenotype alignment.

In practice

Topics

Best for: AI Scientist, Machine Learning Engineer, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by cs.LG updates on arXiv.org.