Simulation Datasets for Home-Based Elderly Care: A New Framework Redefines How Embodied Intelligence Is Trained

A research team drawn from Ping An Technology (Shenzhen) Co., Ltd., The Chinese University of Hong Kong (Shenzhen), the Hubei Meteorological Engineering and Technology Center and Hubei University of Technology has published a systematic methodology for constructing high-quality embodied intelligence simulation datasets tailored specifically to residential elderly care scenarios. The work, released in the journal Big Data Research (ISSN 2096-0271, CN 10-1321/G2), proposes a dual-engine architecture that couples the Marble world model with the NVIDIA Isaac Sim physics engine and establishes a four-tier pipeline spanning scene generation, physical modeling, control execution and data acquisition. The authors argue that embodied intelligence is the pivotal technology capable of relieving the shortage of caregiving labor and the latency of emergency response, yet the training of embodied intelligence models is bottlenecked by the absence of domain-specific, physically faithful and safety-constrained interaction data.

The contribution arrives at a moment when the global discourse around embodied intelligence has moved decisively from laboratory demonstrations toward deployment questions. Researchers in the field increasingly describe embodied intelligence as the fusion of artificial intelligence with robotics, in which an agent perceives, reasons, decides and acts within a closed loop inside a physical environment. For service robots operating in the homes of older adults, that closed loop must function under conditions that are far less forgiving than those found in factories or warehouses: irregular floor plans, dense obstacles, adaptive fixtures such as grab bars and nursing beds, unstable lighting, and, above all, a fragile human body positioned at the center of every interaction.

1. An Aging Population Places New Demands on Embodied Intelligence

The seventh national population census recorded that people aged 60 and above account for 18.7 percent of the total population, an increase of 5.44 percentage points compared with the sixth census. That shift has moved the country into a stage of deep aging, and it has consolidated home-based elderly care as the dominant care modality because it aligns with established living habits and the physical settings in which older adults already reside. The consequence is a widening gap between demand and supply: caregiving labor is scarce, and emergency response is frequently delayed.

Against this backdrop, embodied intelligence has been identified as a key technical means of breaking through the caregiving bottleneck. Robots endowed with embodied intelligence can autonomously perceive their surroundings, make dynamic decisions and physically interact with objects and people, which allows them to assume repetitive, strenuous or time-critical tasks. However, the very properties that make elderly care such a compelling application domain for embodied intelligence also make it one of the hardest domains in which to collect training data.

Home-based elderly care environments are unstructured, multimodal, age-adapted and weakly standardized. There is no uniform spatial layout and no consistent rule governing where objects are placed. Obstacles are distributed irregularly. The space is typically compact and often dimly lit, and it must accommodate adaptive equipment such as handrails, walkers, anti-slip mats and nursing beds. From an interaction standpoint, robots must perform high-risk physical contact tasks, including assisting a person to stand, delivering objects and responding to a fall, which obliges them to respect strict physical safety rules covering force feedback and collision thresholds while simultaneously accommodating the physiological and psychological characteristics of older adults. From an information standpoint, the scene fuses vision, depth, force and language instruction into heterogeneous streams, and the behavior of older adults together with the state of the environment is strongly dynamic, with sudden risk events occurring without warning. These characteristics determine that data construction for embodied intelligence in elderly care cannot simply reuse technical solutions designed for general-purpose scenes.

2. Why Real-World Data Cannot Keep Pace With Embodied Intelligence

High-quality interaction datasets are the foundation on which embodied intelligence models are trained, control policies are optimized and deployment safety is guaranteed. Real-world datasets possess native physical authenticity and natural behavioral quality, but they are constrained by prohibitive collection costs, long collection cycles and the impossibility of reproducing dangerous scenarios at scale. In elderly care specifically, the collection of interaction data involving robots, older adults and household environments faces elevated safety risk, strict ethical constraints, strongly individualized scene configurations and insufficient sample diversity. As a result, real-world collection cannot provide stable data support for core tasks such as fall detection and response, object delivery and stand-up assistance.

Low-cost, scalable and highly controllable simulation data has therefore become the mainstream data solution for embodied intelligence research. Simulation allows scene parameters to be controlled precisely, interactions to be reproduced safely, and task requirements to be customized. For embodied intelligence systems destined for the homes of older adults, simulation is not merely a convenience but a prerequisite, because the alternative would require exposing frail individuals to immature robot policies during the earliest and most failure-prone stages of development.

3. Five Stages in the Evolution of Embodied Intelligence Simulation Datasets

The research team organizes the development of embodied intelligence simulation datasets into five successive stages, tracing a path from static scene reconstruction to dynamic task generation and finally to on-demand simulation driven by world models and generative data engines.

  • Basic simulation environments: research concentrated on the digital modeling of real-world scenes, with representative datasets including SUN3D, Matterport3D and Gibson. The core task involved acquiring high-fidelity indoor three-dimensional structures through RGB-D scanning and multi-view reconstruction to support navigation and perception research.
  • Navigation and instruction datasets: the paradigm shifted from scene modeling toward language-guided behavioral decision-making, with representative datasets including R2R, CVDN and REVERIE. The key innovation was the tight binding of natural language instructions to visual observations and navigation trajectories, producing a three-way coupling of instruction, perception and action.
  • Manipulation and interaction datasets: physical interaction and long-horizon tasks were systematically introduced, with representative datasets including ALFRED, iGibson and ManiSkill. The central breakthrough was the expansion from navigation goals to long-horizon task completion, although the stage still suffered from limited data scale, weak task generalization and the absence of a unified representation across different robots.
  • Large-scale and foundation-model stage: simulation datasets moved toward large scale, multiple tasks and cross-embodiment generalization, with representative datasets including ManiSkill2, BEHAVIOR-1K and Open X. Data scale and task complexity rose exponentially, and the fusion of simulated and real data became an important research direction, yet data acquisition cost, insufficient long-tail task coverage and static data distributions emerged as core bottlenecks.
  • World models and generative simulation: a new generation of representative work has driven rapid evolution of this paradigm. Genie 2 and Genie 3 function as interactive foundation world models that autoregressively generate operable, spatiotemporally consistent three-dimensional interactive environments from a single image or a text prompt. Marble targets persistent three-dimensional world construction and can generate structurally complete, freely explorable scenes from reference images or text, exporting them in standard formats for downstream physics simulation. V-JEPA 2 is a video-based self-supervised world model that learns physical dynamics in latent space through joint embedding prediction, offering predictive capability over future environment states and zero-shot robot planning potential.

These world models push embodied simulation data from collection and reconstruction toward on-demand generation, but the content they produce still faces pronounced challenges in physical consistency, long-horizon controllability and the reliability of fusing simulated with real data. Those challenges motivated the dual-engine design adopted in the new work, in which the Marble world model is responsible for scene generation and the NVIDIA Isaac Sim physics engine performs physical constraint validation.

4. Comparing Today’s Simulation Datasets for Embodied Intelligence

From the perspective of application space, embodied intelligence simulation datasets divide into outdoor and indoor branches, with indoor spaces serving as the core research vehicle for embodied intelligence in domains closely coupled to daily human life. The team systematically classifies indoor embodied intelligence simulation datasets along an axis that runs from industrial space through commercial and service space to residential space and medical care space. Residential space and medical care space are identified as the core carriers for embodied intelligence in home-based elderly care, since residential space maps directly onto the daily care needs of older adults while medical care space provides simulation support for health monitoring, rehabilitation assistance and home-visit nursing.

Within residential space, datasets divide into general-purpose residential datasets and age-adapted residential datasets. General-purpose residential datasets are organized around robot hardware configuration and form four typical technical pathways: fixed-base single-arm robots, represented by RLBench and ManiSkill; mobile-base single-arm robots, represented by ALFRED and ProcTHOR; pure mobile platforms, represented by Gibson; and dual-arm or mobile dual-arm configurations, represented by BEHAVIOR-1K and AgiBot World Alpha.

Age-adapted residential datasets are customized for elderly home scenarios but remain at an early exploratory stage. KIST SynADL centers on mobile-base single-arm robots and covers activities of daily living, while ElderSim supports mobile and humanoid robots and focuses on elderly behavior in age-adapted environments. Both achieve preliminary adaptation to elderly care, yet they still suffer from insufficient scene coverage, coarse modeling of elderly behavior, limited dimensionality of caregiving tasks and an absence of multi-agent interaction, which makes it difficult to satisfy the training requirements of the full emergency response and daily care workflow for embodied intelligence in elderly care.

Simulation space Sub-scenario Core dataset Scene scale Task complexity Robot configuration
Residential General residential RLBench Configurable manipulation scenes 100 graded-difficulty tasks Fixed-base single arm
Residential General residential ManiSkill Configurable manipulation scenes 100+ task types Fixed-base single arm
Residential General residential ProcTHOR 10,000 scenes Generalized tasks Mobile-base single arm
Residential General residential ALFRED 120+ scenes 25,000+ instruction trajectories Mobile-base single arm
Residential General residential Gibson 572 scanned spaces Generalized tasks Pure mobile platform
Residential General residential BEHAVIOR-1K 50 complete residences 1,000+ complex activities Dual-arm configuration
Residential General residential AgiBot World Alpha 100+ scenes 92,000+ real robot trajectories Mobile dual-arm
Residential Age-adapted residential ElderSim 4 background scenes 55 tasks Mobile and humanoid
Residential Age-adapted residential KIST SynADL 5 core scenes 55 activity classes Mobile-base single arm
Medical care Nursing area Isaac for Healthcare Generalizable scenes Generalized tasks Fixed-base single arm
Medical care Nursing area Syn Mediverse 13 scene types, 31 semantic categories Generalized tasks Mobile-base single arm

The comparison makes the gap explicit. General residential datasets have matured into a systematized architecture layered by robot hardware configuration, yet their modeling targets are ordinary households rather than age-adapted environments, and they do not cover age-friendly facilities, elderly behavioral patterns or elderly-care service tasks. Age-adapted datasets, meanwhile, remain sparse. ElderSim constructs four background scenes covering 55 elderly care tasks and supports both mobile and humanoid interaction, making it the first simulation platform focused on elderly behavior and age-adapted environments, while KIST SynADL builds on five core scenes covering 55 classes of elderly daily activity to support mobile-base single-arm assistance tasks. Neither offers the scene coverage, behavioral fidelity or task dimensionality required for a full embodied intelligence training regime in home-based elderly care.

5. Four Generation Paradigms Reshape Embodied Intelligence Data

The generation paradigm behind an embodied intelligence dataset determines its physical fidelity, scalability, annotation completeness and cross-scenario generalization, and therefore shapes training convergence efficiency, interaction policy robustness and real-world deployment performance. The team identifies four mainstream paradigms, each representing a distinct trade-off.

Generation method Core mechanism Representative work Scalability Physical fidelity Key limitation
Physics-simulation-based procedural generation Simulation randomization and execution RLBench, ManiSkill High High Sim-to-real gap, limited generalization
Teleoperation-based data collection Human teleoperates the robot RoboCasa, MimicGen Low High Labor intensive, low throughput
Three-dimensional reconstruction and digital twins NeRF, 3D Gaussian Splatting RoboTwin, D3Fields Medium Medium Time-consuming reconstruction, missing physical attributes
Generative model driven synthesis Large language and vision model driven generation GenSim, RoboGen, UniSim, Genie High Medium to low Requires validation pipelines, prone to hallucinated tasks, can violate physical laws

Physics-simulation-based procedural generation relies on high-fidelity physics engines and high-performance computing to randomize scene layout, object pose and physical attributes while driving agents to execute interaction tasks automatically. It delivers extremely high data throughput and complete ground-truth annotation covering six-dimensional object poses, joint torques and contact forces, but the reality gap persists: insufficient contact modeling accuracy and deviations in high-frequency physical effects degrade generalization when policies are transferred to real robots, which is especially problematic for caregiving and nursing tasks that demand high contact precision.

Teleoperation-based collection places a human operator in direct control of a physical or virtual robot through virtual reality or haptic interfaces. The resulting data naturally satisfies real physical constraints and captures natural human operational logic together with fine-grained interaction semantics, making it well suited to scenarios where safety and behavioral naturalness matter. Its weaknesses, however, are structural: the process is labor intensive, throughput is extremely low, and differences in operator style, device transmission latency and interaction instruction noise introduce distribution bias that undermines training stability for embodied intelligence.

Three-dimensional reconstruction and digital twins employ advanced reconstruction techniques to produce photorealistic virtual replicas of real interiors, providing high-realism visual input and enabling safe pre-deployment testing of policies. Reconstruction is nonetheless time-consuming and costly at scale, and the resulting scenes typically lack accurate physical attributes such as mass, friction coefficient and joint damping, which limits the reproducibility and diversity of interactive behavior.

Generative model driven synthesis leverages the world knowledge of large language models, multimodal models or diffusion models, and splits into two paths. The first combines large models with simulators, using language models to generate task descriptions, plan scene layouts and design reward functions before a physics simulator produces structured interaction data; this path broadens task diversity and reduces manual design cost, but physical plausibility requires additional verification and the end-to-end pipeline remains unstable. The second generates visual frames and action trajectories directly from noise or text, achieving high scalability and photorealistic visuals without an explicit simulator, but such data frequently violates physical laws through object penetration, non-inertial motion and distorted contact interaction, and it lacks precise state annotation and temporal action consistency.

The team concludes that no single paradigm is universally optimal, and that multi-paradigm fusion is becoming the development trend for scenarios such as elderly care that demand high safety and physical fidelity. This conclusion directly informs the proposed construction method.

6. Three Core Technical Challenges for Embodied Intelligence in Elderly Care

Combining existing research with practical requirements, the construction of simulation datasets for embodied intelligence in home-based elderly care confronts three core technical challenges.

  • High-fidelity automated generation of home-based elderly care scenes, which must reconcile realism, plausibility and completeness of physical interaction.
  • Closed-loop control and data acquisition for safety-critical tasks such as emergency response, which requires full automation of event triggering, robot control and data annotation.
  • Safety-oriented modeling of caregiving actions, which requires embedding physical rules such as force feedback and collision constraints into action representation and data generation.

To address these challenges, the proposed method adopts scenario-specific customization, dual-engine collaboration and process automation as its core construction principles, and organizes its technical framework into four layers: scene generation, physical modeling, control execution and data acquisition.

7. A Four-Layer Architecture for Home-Care Embodied Intelligence Datasets

The architecture is driven top-down by two core home-based elderly care service tasks, with every layer designed around the actual interaction logic of elderly care so that the dataset remains tightly matched to real service requirements. It is built bottom-up as an incremental progression, where lower layers provide foundational support and upper layers define task orientation, with standardized interfaces carrying data transmission, instruction dispatch and state feedback so that the data flow and control flow are fully connected end to end.

  • Scene generation layer: powered by the Marble world model pipeline, this layer automates the transformation from two-dimensional reference images or text descriptions into complete three-dimensional home scenes. Its outputs include floor plans consistent with the spatial characteristics of Chinese home-based elderly care, indoor object assets with physically realistic materials and textures, and scene variants adapted to different task requirements. All generated scenes are exported in OpenUSD format to ensure seamless connection with the downstream physics simulation engine and to support parametric control over scene elements for later large-scale domain randomization. Age-friendly spatial priors are injected into the Marble generation capability through an asset library containing handrails, nursing beds, walkers and anti-slip mats, and parameters such as compact floor plans, barrier-free circulation and low illumination are incorporated as generation constraints.
  • Physical modeling layer: this layer receives three-dimensional scene assets and completes physical-level accurate modeling of both scene and robot system inside NVIDIA Isaac Sim. It assigns precise physical attributes including mass, friction coefficient, restitution coefficient and collision volume to all objects, and constructs rigid-body dynamics and contact mechanics simulation consistent with real physical laws. High-frequency interaction objects in elderly care receive refined collision mesh modeling and physical parameter calibration. The layer also imports Unitree series robot URDF models, configures joint degrees of freedom, sets driver parameters and mounts sensors. In addition, it constructs parameterized elderly human body models with body dimensions, restricted joint ranges of motion and typical fall postures characteristic of older populations, and introduces human-robot safety constraints that set contact force limits, collision thresholds and compliant contact models for a fragile interaction partner.
  • Control execution layer: this layer binds general motion and manipulation policies to elderly care tasks. Emergency response introduces a fall event trigger mechanism and safety navigation constraints, while daily care introduces force control and obstacle avoidance constraints for operations in proximity to the human body. It adopts a hierarchical control architecture separating high-level decision planning from low-level motion control, configures differentiated algorithm combinations for different task scenarios, and applies domain randomization to object positions, lighting conditions, robot initial states and terrain parameters to strengthen the diversity and generalization of the generated data.
  • Data acquisition layer: this layer systematically extracts, aligns, annotates and stores multimodal data, synchronously recording RGB images, depth maps and proprioceptive data from the robot’s onboard sensors and covering both the first-person robot view and a global monitoring view of the environment. Beyond general multimodal recording, it extends elderly care specific annotation fields including fall posture category, rescue target point, human-robot relative distance and bearing, care task semantics and contact force.

Because the entire generation chain communicates through OpenUSD, the resulting data formats are compatible with mainstream humanoid and quadruped robot simulation development frameworks as well as physical robot driver interfaces, enabling subsequent sim-to-real transfer testing once physical prototypes are available.

8. Quadruped Robots and Emergency Response: A Closed-Loop Pipeline

To support emergency response capability in home-based elderly care, the team developed a dedicated dataset pipeline built on high-fidelity three-dimensional home scenes generated by the Marble model. The pipeline targets high-frequency emergency events, above all falls among older adults, and binds the robot’s cooperative response logic to the data collection process. Quadruped robots are assigned responsibility for whole-home mobility and environmental monitoring, ensuring that the dataset covers the full dimensional range of perception, motion and interaction data required for emergency embodied intelligence.

Data is acquired through a simulation generation fusion model built on the Marble-USD standardized scene set and the IsaacLab simulation engine, achieving large-scale, high-fidelity generation. Through the OpenUSD data exchange channel, the pipeline outputs high-resolution multimodal visual data including RGB images, depth maps, semantic segmentation maps and three-dimensional point cloud data across varied lighting and environmental conditions as well as flat and undulating terrain.

The control execution layer for this scenario adopts a three-tier architecture combining high-level perception and planning, mid-level instruction conversion and low-level motion control. The high-level module builds on the NaVILA vision-language navigation model, taking multimodal visual data and rescue task instructions as input, extracting visual features from RGB and depth images through a convolutional neural network, and mapping natural language rescue instructions to a global sequence of waypoints through a vision-language alignment mechanism. The path planning process is formalized as a maximum a posteriori estimation problem in which the optimal waypoint sequence is obtained by maximizing the posterior probability conditioned on the RGB image, the depth image, the task instruction and the goal state.

The mid-level instruction conversion module receives the waypoint sequence and task constraints and decouples the abstract path planning result into a standardized instruction set executable by low-level motion control, comprising a desired linear velocity command and a desired angular velocity command defined over the navigation duration, with smooth interpolation and obstacle avoidance optimization ensuring instruction continuity and safety. The low-level controller adopts the RSL_RL locomotion policy for the Unitree Go1, using an actor-critic neural network policy trained with proximal policy optimization to convert mid-level velocity instructions into torque control signals for the robot’s twelve joints, enabling stable and agile locomotion across complex domestic terrain. The locomotion policy uses independently parameterized actor and critic networks: the actor maps observations to a probability distribution over joint actions through a three-layer fully connected network that outputs action means with independently learnable standard deviation parameters, while the critic adopts the same network topology with independent parameters to output a scalar state value estimate used in advantage function computation. Training proceeds in synchronous batch mode, with thousands of parallel simulation instances executing policy sampling simultaneously to accelerate data collection.

The data acquisition layer records heterogeneous data streams synchronously at each simulation time step and aligns them through the global clock provided by the simulation engine. The visual perception stream records first-person images from the head-mounted RGB and depth cameras at a rate of 20 Hz, while a global overhead monitoring camera records a bird’s-eye view of the scene at the same rate for annotation verification and behavioral visualization. The proprioceptive stream synchronously records the robot’s complete body state vector at a control frequency of 50 Hz, including position, velocity and torque measurements for all twelve joints together with three-axis angular velocity and acceleration from the inertial measurement unit. The scene state stream records the time-varying state of all dynamic objects in the simulation, covering the three-dimensional pose of the fallen older adult, the fall state label distinguishing lateral, supine and prone postures, and the Euclidean distance and relative bearing between the robot base and the target person.

9. Bipedal Humanoids and Daily Care: A Closed-Loop Pipeline

For the daily care dataset, the guiding requirement is the set of high-frequency assistance needs that arise in home-based elderly care, giving rise to typical caregiving tasks such as dietary assistance, medication management, object organization and delivery. The three-dimensional environment generation layer again relies on the Marble pipeline, emphasizing home environments with dense functional manipulation spaces including kitchen areas, bedrooms and living rooms. During scene generation, particular attention is paid to fine-grained modeling of manipulable objects: every graspable object is configured with accurate geometric dimensions, mass, center of mass, surface friction coefficient and grasp keypoint annotation to ensure the physical realism of subsequent robotic manipulation.

Generated care scenarios are imported into Isaac Sim in OpenUSD format for full physical attribute configuration. For the specific requirements of daily care, the physical modeling layer concentrates on rigid-body dynamics modeling of interactive objects, including precise collision mesh generation, friction cone parameter calibration and inter-object contact stiffness setting, ensuring realistic physical simulation of grasping, placing, pushing and pulling. The layer also completes system modeling of the Unitree G1 humanoid robot, importing its URDF model, configuring joint drive parameters for twenty-three degrees of freedom across the whole body, and mounting a head-mounted RGB-D camera for scene perception and object detection together with wrist-mounted RGB cameras for close-range manipulation guidance, torque sensors for grasp force sensing and whole-body joint encoders.

The control execution layer combines the Psi-Zero end-to-end learning framework with large language model driven task planning to achieve end-to-end execution from natural language care instructions to coordinated whole-body operation. Psi-Zero functions as a general control policy purpose-built for humanoid robots and can handle the coupled control of locomotion and manipulation in a unified manner, completing fine-grained dexterous operations such as grasping, carrying and placing in daily care scenarios. Its architecture follows a layered design comprising a vision-language model backbone, an action expert and a lower-limb controller, with the three systems dividing responsibilities while operating in coordination: the vision-language backbone handles high-level perception and semantic understanding, the action expert generates whole-body motion, and the lower-limb controller safeguards low-level locomotion stability. This architecture suits the complex mobile manipulation tasks involved in home care, maintaining whole-body balance and stable locomotion while the upper body performs fine manipulation.

The highest-level decision module is built on the Qwen3-VL-2B-Instruct pretrained vision-language model and undertakes scene understanding, object recognition and task semantic parsing. It receives multi-view visual observations from the head and both wrists together with the natural language care instruction and outputs a unified multimodal feature representation that provides semantically rich input for downstream action generation. RGB images from the head-mounted RGB-D camera and the two wrist cameras are divided into fixed-size patches, linearly embedded and positionally encoded before being fed into a multi-layer Transformer encoder. The visual encoder inherits pretrained weights from Qwen3-VL on large-scale image-text pairs and possesses open-world object recognition capability, enabling accurate identification of everyday items in the home environment such as medicine bottles, water cups and tableware. The language encoding process tokenizes the care instruction and generates text features through the language encoder of Qwen3-VL, so that the instruction is embedded into the same representational space as the visual observations.

The mid-level execution module adopts a multimodal diffusion Transformer architecture and is responsible for mapping the multimodal features produced by the high-level module into a whole-body action sequence for the robot. The lowest-level execution module builds on an adaptive model-based optimization reinforcement learning tracking policy and maps the eight-degree-of-freedom locomotion instruction output by the mid-level module into fifteen-degree-of-freedom lower-limb joint angles, safeguarding whole-body balance and locomotion stability while the humanoid performs upper-limb manipulation tasks.

The data acquisition layer synchronously records heterogeneous modality streams during simulation runs and enforces strict timestamp alignment through the Isaac Sim global simulation clock to form a standardized care task dataset. Given the three-tier architecture, the acquisition layer records the input and output states of each level to support model training, inference analysis and system-level debugging. The visual perception stream records images from the robot’s multi-view camera system at a rate of 10 Hz, while the proprioceptive stream records the complete body state at a control frequency of 100 Hz, including joint states, inertial measurement data, end-effector states and foot states.

Robot platform Task domain Visual data rate Proprioceptive rate Recorded streams
Unitree Go1 quadruped Emergency response and fall detection 20 Hz 50 Hz RGB, depth, semantic segmentation, point cloud, joint states, inertial data, scene state
Unitree G1 humanoid Daily care and manipulation 10 Hz 100 Hz Multi-view RGB, joint states, inertial data, end-effector states, foot states, object interaction states

10. Evaluation, Limitations and the Research Outlook for Embodied Intelligence

The team emphasizes that the current stage focuses on the design of the overall methodology and architectural construction for home-based elderly care simulation datasets. The complete technical scheme, the four-layer framework, the dual-engine collaboration logic and the two robot task pipelines have been systematically designed; large-scale dataset production has not yet been completed because of substantial real-world safety risks. The published assessment is therefore a theoretical qualitative comparison and an argument for architectural soundness rather than an empirical benchmark report.

Comparing the design logic against existing age-adapted datasets, the team identifies three advantages. In scene generation, existing age-adapted datasets depend on manual modeling and face high scene expansion costs, whereas the proposed approach uses the Marble world model to achieve automated parametric scene generation driven by text and images, addressing the root cause of sample scarcity in elderly care scenarios. In physical constraints, most existing approaches perform only human motion animation rendering and lack safety physics such as human-robot contact mechanics and protection thresholds for older adults, whereas the proposed approach completes full-object physical parameter calibration in Isaac Sim and embeds human-robot safety limits directly into the modeling stage, matching the safety imperatives of close-range interaction in elderly care. In task design, current age-adapted datasets lack closed-loop robot interaction workflows, whereas the proposed approach separates a quadruped emergency rescue pipeline from a bipedal daily care pipeline, designing control chains around real elderly care business tasks and achieving a complete closed loop from scene through modeling and control to data acquisition.

Looking ahead, the researchers outline several directions for deepening the work on embodied intelligence datasets for elderly care.

  • Deepening sim-to-real transfer calibration. Building on visual, force, trajectory and interaction data collected from real quadruped and bipedal robots in actual home environments, the team plans to construct an automated sim-to-real parameter calibration workflow and combine it with domain adaptation algorithms to narrow the distribution gap between simulation and reality, quantitatively validating key transfer metrics such as emergency response latency, manipulation success rate and navigation accuracy.
  • Constructing a dedicated multi-robot collaborative care dataset. This involves designing collaborative task paradigms in which quadruped robots perform whole-home inspection and emergency localization while bipedal robots perform fine manipulation and close-range care, collecting full-process data on multi-robot perception sharing, dynamic task allocation, interactive obstacle avoidance and communication synchronization, and establishing annotation rules and evaluation benchmarks for multi-agent collaborative data to support the development of multi-robot elderly care algorithms.
  • Promoting open sharing and industrial deployment of embodied intelligence datasets. The plan is to release core data subsets and a simulation deployment toolchain, provide baseline algorithms and data loading interfaces, and work with elderly care service institutions and robotics enterprises on real-scenario validation to form industry application standards for home-based elderly care embodied intelligence datasets.
  • Continuously expanding the coverage of home-based elderly care scenarios and tasks. Targeting older adults with different health conditions including semi-disabled, disabled and advanced-age individuals living alone, new sub-tasks are to be added such as home rehabilitation training, chronic disease health monitoring, cognitive impairment companionship, secondary rescue after a fall, precise medication and feeding assistance, and intelligent regulation of age-adapted environments, while covering multiple housing types such as small apartments, age-adapted renovated residences and community care homes, generating vast scene variants through parametric domain randomization to fill data gaps in long-tail risk scenarios.

11. Conclusion

The unstructured characteristics, high safety interaction requirements and specialized service tasks of home-based elderly care determine that data construction for embodied intelligence in this domain must break through the adaptability limits of general-purpose datasets. Simulation datasets, with their advantages of low generation cost and high scene controllability, have become the core path for resolving the shortage of embodied intelligence data resources in elderly care.

Addressing the three core technical challenges, the proposed method adopts scenario-specific customization and dual-engine collaboration, and through the four standardized layers of scene generation, physical modeling, control execution and data acquisition, achieves a fully standardized construction process running from high-fidelity three-dimensional home scene generation to multimodal interaction data output. Supported by the Marble world model and the NVIDIA Isaac Sim physics simulation engine, the approach completes scene modeling aligned with the spatial characteristics of home-based elderly care, human-robot interaction modeling adapted to the physiological characteristics of older adults, and differentiated control architecture design for core tasks including emergency response and daily care. The resulting specialized dataset covers cross-embodiment data as well as multiple dimensions of robot perception, decision-making and control, filling a gap in embodied intelligence simulation data for home-based elderly care and compensating for the adaptability deficiencies of general-purpose embodied intelligence datasets with respect to elderly care scene characteristics, task settings and human-robot safety interaction constraints.

Future work will extend scene and task coverage to form a full-scenario, full-task home-based elderly care embodied intelligence simulation data system, deepen sim-to-real transfer research by optimizing and correcting simulation datasets with measured data from real home environments, and pursue dataset construction for multi-robot collaborative elderly care to match the requirements of coordinated service delivery. As embodied intelligence continues its migration from the laboratory into the homes of older adults, the availability of physically faithful, safety-constrained and domain-specific simulation data will determine how quickly that transition can occur, and how safely it can be accomplished.

Scroll to Top