Ethics assertion
Participation was voluntary, and contributors acquired compensation in accordance with UK laws and the accepted institutional protocol. Previous to knowledge assortment, all contributors have been knowledgeable concerning the research goals, experimental procedures, potential dangers, and the deliberate public launch of the dataset. The launched knowledge can be anonymised and would include solely motion-related radar level clouds, skeleton coordinates, motion labels, and topic ID, with none instantly identifiable private data. Written knowledgeable consent was obtained from all contributors for participation within the research and for sharing the anonymised motion-related knowledge in a public repository. Individuals have been additionally knowledgeable earlier than the experiment that they may contact the analysis group in the event that they wished to request elimination of their knowledge from future public releases. All knowledge have been dealt with in accordance with institutional moral tips and relevant knowledge safety necessities. The research was accepted by the Queen Mary Ethics of Analysis Committee (Reference No: QME24.0392).
Experimental setup
The experimental setup consists of a vertical platform supporting each the mm-wave radar sign processing machine and the Azure Kinect digicam, a big display screen positioned above the platform to show a steering video for helping contributors with the required actions, and a laptop computer used to manage and coordinate all parts of the system (Fig. 1). The radar machine was positioned on a desk with a peak of 73 cm and securely fastened to make sure stability all through the experiment. The Azure Kinect was mounted instantly above the radar and used to supply reference skeleton annotations for subsequent knowledge validation and evaluation.

Information assortment atmosphere and setup.
The experiments have been performed in a typical managed indoor room with plain painted partitions, a carpeted ground, and several other atypical tables with laminate board tops and gray powder-coated metallic legs across the exercise space. No massive metallic objects, mirrors, glass partitions, or different deliberately extremely reflective surfaces have been positioned on this room. The principle surrounding surfaces consisted of frequent indoor supplies somewhat than robust specular reflectors.
The bodily room was roughly 6.5 m × 7 m, whereas the efficient seize space proven in Fig. 1(b) was roughly 4.5 m × 3 m and corresponded to the retained sensing area used for participant movement recording and radar level filtering. Radar factors exterior this predefined exercise space have been discarded, as they have been extra more likely to originate from environmental reflections, multipath results, or irrelevant background objects. Information assortment was performed beneath two lighting situations: a well-lit atmosphere and a low-light atmosphere. These two situations have been included to file participant actions beneath totally different visible situations.
Participant recruitment
26 wholesome adults with no recognized mobility limitations have been recruited for this research. Participant ages vary from 25 to 57 years ((mu ,=,34.15,sigma ,=,mathrm{7.45)}), heights from 159 cm to 186 cm ((mu ,=,172.1,cm,sigma ,=,7.00,cm)), weights from 52 kg to 97 kg ((mu ,=,71.4,kg,sigma ,=,11.5,kg)), and physique mass index (BMI) values from 17.99 to 30.11 ((mu ,=,24.11,sigma ,=,mathrm{3.49)}).
Sensor modalities
Level cloud from mm-wave radar
A Texas Devices IWR6843AOPEVM was employed, which is a Frequency-Modulated Steady Wave (FMCW), multiple-input multiple-output (MIMO) radar system working within the unlicensed 60 GHz frequency band. It was configured particularly for the detection and monitoring of human motion15,16 utilizing the TI 3D Folks Counting firmware17. Detection of factors is carried out utilizing the fixed false alarm fee (CFAR) algorithm utilized to the radar knowledge dice. For every detected level (p), the output contains native three-dimensional coordinates (({X}_{p},{Y}_{p},{Z}_{p})), and Doppler velocity (({D}_{p})), all measured relative to the radar sensor. The radar’s native spherical coordinate system (({rho }_{i},{varphi }_{i},{theta }_{i})) is remodeled into Cartesian coordinates for additional evaluation and integration. Detailed configuration parameters of the radar system are summarised in Desk 2.
RGB-D frames from Azure Kinect
To supply a reference for radar-based movement evaluation, we synchronously recorded skeleton data utilizing Azure Kinect throughout knowledge assortment. Azure Kinect provides markerless RGB-D sensing via a time-of-flight depth digicam and an RGB digicam, and its Physique Monitoring SDK estimates 3D skeletons with 32 joints at 30 Hz for every tracked physique18. It was chosen as a sensible reference modality as a result of it’s simpler to deploy than multi-camera marker-based motion-capture techniques, whereas earlier research have reported helpful depth accuracy and body-tracking efficiency in managed settings19,20. The Kinect-derived skeletons are used for spatial calibration, temporal alignment, knowledge high quality inspection, and supervised analysis of radar-based fashions. They aren’t used as mannequin enter throughout radar-only inference.
Information assortment
Determine 2 illustrates the general system pipeline. Every participant spent roughly 20 minutes finishing the experimental process, together with the research introduction, rationalization of experimental necessities, consent affirmation, motion demonstration, and knowledge recording.

Overview of the information gathering system pipeline with three benchmark duties.
Throughout recording, contributors stood at a set place roughly 3.5 m from the sensor platform to stay inside the shared sensing area of the mm-wave radar and Azure Kinect. Individuals have been instructed to comply with the reference video (https://youtu.be/WBGPFIWC_50) displayed on the monitor and reproduce the demonstrated actions as constantly as attainable. The motion set contained 21 predefined classes, proven in Desk 3. Every motion was repeated thrice. After every motion, contributors returned to the preliminary standing posture, with each arms naturally relaxed on the sides of the physique, ft collectively, and the physique stored nonetheless till the following motion started. The length of every motion and the brief pauses between consecutive actions are listed in Desk 3.
To make sure that the total motion sequence was captured, roughly 5 s of static standing was recorded earlier than and after every group of actions. These extra static intervals have been used to facilitate segmentation and high quality checking, and have been eliminated as a lot as attainable throughout knowledge preprocessing. The complete motion set was recorded beneath each normal-light and low-light situations.
Information processing
Calibration
The radar operates inside a coordinate system centred at its transmitter, whereas the Azure Kinect defines its coordinate system with the optical centre of the depth digicam because the origin. To reinforce the accuracy of downstream duties, it’s important to align the coordinate frames by remodeling the keypoint positions from the digicam coordinate system to that of the radar. For this objective, a nook reflector is positioned at a number of recognized places to function an anchor level (Fig. 3).

Calibration of Kinect and radar coordinates.
For radar, alerts are predominantly mirrored when the reflector’s centre is aligned with the radar’s line of sight. Within the case of the Azure Kinect, which integrates each depth and RGB cameras, object recognition methods are employed to detect the triangular reflector. Subsequently, the centre level’s place is calculated inside the Kinect coordinate system.
Desk 4 enumerates the chosen calibration factors. An actual level is concurrently noticed by each the radar machine and the Azure Kinect. These gadgets independently compute the purpose’s place inside their respective coordinate techniques. The transformation is carried out utilizing Singular Worth Decomposition (SVD) as a part of the usual Procrustes alignment methodology, which includes the next steps:
-
1.
Centroid Alignment: We compute the centroids of each level units and subtract them from the unique factors to acquire centred units ({{bf{P}}}_{c}) and ({{bf{Q}}}_{c}).
-
2.
Singular Worth Decomposition (SVD): We calculate the covariance matrix ({bf{H}}={{bf{P}}}_{c}^{high }{{bf{Q}}}_{c}) and carry out SVD: ({bf{H}}={bf{U}}{boldsymbol{Sigma }}{{bf{V}}}^{high }). The optimum rotation is then given by ({bf{R}}={bf{V}}{{bf{U}}}^{high }).
-
3.
Reflection Correction: If ({rm{d}}{rm{e}}{rm{t}}({bf{R}}mathrm{) < 0}), we appropriate for improper rotation by flipping the signal of the final row of ({{bf{V}}}^{high }) and recomputing ({bf{R}}).
-
4.
Translation Vector: The interpretation vector is computed as ({bf{t}}=bar{{bf{Q}}}-{bf{R}}bar{{bf{P}}}).
-
5.
Transformation and Error Analysis: The unique Kinect factors are remodeled utilizing ({{bf{P}}}_{{rm{aligned}}}={bf{R}}{bf{P}}+{bf{t}}).
Determine 3 reveals the calibration factors within the digicam and radar coordinate techniques earlier than and after alignment. The ensuing transformation parameters (({bf{R}},{bf{t}})) are saved and subsequently used to align Kinect-derived skeleton labels with radar knowledge. After making use of the computed transformation, the Kinect factors align intently with the corresponding radar factors, with a imply Euclidean residual of 39.0 mm, an RMSE of 44.0 mm, and a most residual of 75.5 mm. The imply absolute residuals alongside the x/y/z axes have been 10.9/33.1/12.3 mm.
Synchronisation
The radar machine operates at 18.2 Hz, whereas the Azure Kinect captures at 30 Hz, resulting in a pure temporal misalignment between their knowledge streams. In our setup, each sensors have been managed by a single host laptop, which triggered knowledge seize and recorded a system timestamp instantly after every body. This shared system clock ensured constant timestamp formatting and comparability throughout gadgets. To realize frame-level alignment, a nearest-neighbour timestamp matching methodology was utilized. The residual temporal mismatch after this software-based synchronisation was quantified utilizing absolutely the timestamp distinction between matched radar and Kinect frames, and the empirical synchronisation-error statistics are reported in Desk 5.
As proven in Desk 5, the deployed timestamp-based pipeline achieved a imply error of 8.52 ms, a median error of 8.42 ms, and an RMSE of 12.19 ms, with p95, p99, and p99.9 errors of 16.99 ms, 22.14 ms, and 30.86 ms. Remoted timestamp interruptions affected 61 out of 137,717 frames and produced a most error of 491.25 ms; after excluding the affected recordings, the utmost error decreased to 77.32 ms whereas the imply and median remained at 8.41 ms. General, the process offers steady synchronisation for the present managed acquisition protocol, though residual timing errors should still have an effect on quick or abrupt actions. {Hardware} triggering was not used as a result of the Azure Kinect wired synchronisation mode is designed for multi-Kinect setups, and the radar pipeline didn’t present a instantly appropriate shared set off or frequent clock interface.
Movement segmentation
Every recording group comprises roughly 4 consecutive actions separated by brief static pauses. Movement segmentation was carried out to extract the person motion intervals from every group. We used frame-wise Movement Power (ME) computed from the Kinect-derived skeleton sequence as the first segmentation cue, as a result of the skeleton trajectories present a compact illustration of whole-body motion.
Earlier than ME calculation, every joint coordinate trajectory was smoothed utilizing a one-dimensional Gaussian filter to cut back frame-level jitter. For a skeleton sequence with T frames and J joints, let ({{bf{p}}}_{t}^{(j)}in {{mathbb{R}}}^{3}) denote the 3D coordinate of joint j at body t. The ME at body t is outlined as
$${{rm{ME}}}_{t}=mathop{sum }limits_{jmathrm{=1}}^{J}{Vert {{bf{p}}}_{t}^{(j)}-{{bf{p}}}_{t-1}^{(j)}Vert }_{2},,t=mathrm{2,}ldots ,T,$$
(1)
with ME1 = 0 to protect the unique sequence size.
The segmentation threshold was decided dynamically for every recorded group somewhat than fastened throughout topics or actions. Let ({rm{ME}}={{{rm{ME}}}_{1},{{rm{ME}}}_{2},ldots ,{{rm{ME}}}_{T}}) denote the motion-energy sequence of a recorded group. We utilized a sliding window of size L = 60 to this sequence, the place the i-th window is outlined as
$${W}_{i}={{{rm{ME}}}_{i},{{rm{ME}}}_{i+1},ldots ,{{rm{ME}}}_{i+L-1}mathrm{}}.$$
For every window Wi, we computed the native movement variation as
$${D}_{i}=mathop{{rm{m}}{rm{a}}{rm{x}}}limits_{tin {i,ldots ,i+L-mathrm{1}}}{{rm{ME}}}_{t}-mathop{{rm{m}}{rm{i}}{rm{n}}}limits_{tin {i,ldots ,i+L-mathrm{1}}}{{rm{ME}}}_{t}.$$
(2)
If ({D}_{i} < delta ), the window was used as a candidate low-motion window for estimating the adaptive threshold. In our implementation, (delta ,=,100), which was chosen empirically to differentiate steady pauses from intervals with substantial motion variation. Desk 6 current the variety of legitimate frames and factors retained for every motion sort in any case processing steps.
Let ({{mathscr{W}}}_{s}) denote the set of candidate low-motion home windows. The adaptive threshold TME was then computed as
$${T}_{{rm{ME}}}=(start{array}{cc}{rm{m}}{rm{a}}{rm{x}}{{{rm{ME}}}_{t}|{{rm{ME}}}_{t}in {W}_{i},,{W}_{i}in {{mathscr{W}}}_{s}}, & {rm{if}},{{mathscr{W}}}_{s}ne varnothing , {P}_{90}({rm{ME}}), & {rm{in any other case}},finish{array}$$
(3)
the place ({P}_{90}({rm{ME}})) denotes the ninetieth percentile of the total ME sequence. When candidate low-motion home windows can be found, ({T}_{{rm{ME}}}) is the most important ME worth noticed inside these home windows. If no candidate low-motion window is detected, the ninetieth percentile of the total ME sequence is used as a fallback threshold.
Frames with ({{rm{ME}}}_{t} < {T}_{{rm{ME}}}) have been marked as stationary, and consecutive stationary frames have been grouped into stationary intervals. In our implementation, a stationary interval was used as an motion boundary provided that it lasted for at the very least (Ok,=,60) frames, similar to roughly 2 s for Kinect recordings at 30 Hz. This minimum-length criterion was used to keep away from splitting actions at brief pauses or transient fluctuations inside a motion. Lengthy stationary intervals have been then used to separate consecutive motion segments inside every recorded group. This adaptive technique accounts for variations in participant motion amplitude, execution velocity, and skeleton monitoring noise. All robotically detected motion boundaries have been visually inspected and adjusted the place obligatory earlier than developing the ultimate aligned motion segments.
After the above acquisition, alignment, segmentation, and quality-checking procedures, the ultimate launched dataset was constructed from the retained legitimate motion segments.