High-Quality Embodied Intelligence Simulation Datasets for Residential Elderly Care: A New Four-Tier Framework

Population aging is reshaping the demand for care, assistance, and emergency response in the home. As the share of older adults rises, residential elderly care has become a dominant care model, yet the shortage of caregiving labor and the lag in emergency response remain persistent problems. Embodied intelligence, which integrates artificial intelligence with robotics, is widely viewed as a key technology for addressing these gaps. Unlike purely digital artificial intelligence, embodied intelligence requires an agent to perceive, reason, decide, and act inside a physical environment. For residential elderly care, that means a robot must navigate compact rooms, recognize falls, deliver objects, support standing, manage medication, and interact safely with older adults whose bodies may be fragile. The training of embodied intelligence models depends on large-scale, high-fidelity, interactive data. However, collecting such data in real homes is difficult because of safety risks, ethical constraints, high costs, long collection cycles, and the difficulty of reproducing dangerous events. Simulation datasets have therefore become a central solution for embodied intelligence research, because they offer controllable scenes, repeatable interactions, and scalable task generation.

A newly released study proposes a systematic method for constructing high-quality embodied intelligence simulation datasets specifically for residential elderly care. The work brings together the Marble world model and the NVIDIA Isaac Sim physics engine in a dual-engine architecture. It builds a four-tier pipeline covering scene generation, physical modeling, control execution, and data acquisition. It then designs specialized workflows for quadruped robots in emergency response and bipedal humanoid robots in daily care. The result is a domain-specific embodied intelligence data framework that aims to support vision-language navigation, vision-language-action, physical intelligence, spatial intelligence, and sim-to-real transfer for elderly care robotics. The study was accepted on July 27, 2026, and released online on September 10, 2026. Its central claim is that generic simulation datasets are not sufficient for residential elderly care, and that embodied intelligence in this domain requires customized scene generation, safety-aware physical constraints, and task-specific closed-loop data collection.

1. Why Embodied Intelligence Needs a Dedicated Data Foundation for Aging in Place

The seventh national population census reported that 18.7 percent of the population was aged 60 or above, an increase of 5.44 percentage points from the previous census. This demographic shift has made aging in place a mainstream care model, but it has also exposed structural weaknesses in caregiving capacity and emergency response. Older adults often live in non-standardized homes with irregular furniture placement, narrow passages, low lighting, and age-friendly additions such as handrails, walkers, nursing beds, and anti-slip mats. These environments are not the same as standard residential or laboratory spaces. They are compact, personalized, and highly variable. For embodied intelligence, this means that navigation, perception, manipulation, and human-robot interaction policies trained in generic scenes may fail when deployed in real elder-care homes.

Residential elderly care also involves a special class of interaction objects. In general manipulation datasets, robots mostly interact with rigid or semi-rigid objects. In elder care, the robot must interact with the older adult’s body. That body is both highly individual and relatively fragile. Physical contact must respect force limits, collision thresholds, and compliant contact models. A robot that performs well in a warehouse or a kitchen may not be safe in a bedroom where it must help an older adult stand up, deliver medication, or respond to a fall. Embodied intelligence in this setting therefore requires data that captures not only geometry and appearance but also contact mechanics, human pose, safety constraints, and task semantics. The absence of such data is a major bottleneck for embodied intelligence deployment in elderly care.

2. The Data Bottleneck in Embodied Intelligence for Elderly Care

Real-world data collection for embodied intelligence in elderly care faces several hard constraints. First, safety risks are high. A robot learning near an older adult can cause falls, collisions, or distress. Second, ethical constraints are strict. Older adults may not consent to repeated data collection, especially during emergencies or intimate care tasks. Third, scenes are highly personalized. Every home has a different layout, different furniture, different lighting, and different care needs. Fourth, sample diversity is insufficient. Rare events such as falls, choking, or sudden illness are difficult to capture on demand. These factors make it impractical to rely only on real-world data for embodied intelligence model training.

Simulation offers an alternative. A simulation dataset can generate large numbers of scenes, randomize object positions, control lighting, reproduce dangerous events safely, and provide precise ground truth. For embodied intelligence, simulation is not merely a supplement; it is a necessary infrastructure for training, testing, and validating policies before real-world deployment. However, generic simulation datasets have their own limitations. They may cover tabletop grasping, object rearrangement, or standard household navigation, but they often do not cover the long-horizon, multi-step, safety-critical tasks that define residential elderly care. They may provide visual realism but lack accurate physical properties. They may include human avatars but not the limited joint range, fragile body mechanics, or fall postures of older adults. The result is a domain gap that prevents direct transfer to elderly care embodied intelligence.

3. From Static Scenes to Generative World Models: The Evolution of Embodied Intelligence Simulation

The development of embodied intelligence simulation datasets can be understood as a transition from static scene reconstruction to dynamic task generation, and then to world-model-driven, on-demand simulation. The study identifies five broad stages. In the first stage, basic simulation environments focused on digitizing real-world scenes. Representative datasets include SUN3D, Matterport3D, and Gibson. These works used RGB-D scanning and multi-view reconstruction to obtain high-fidelity indoor 3D structures. Matterport3D provided detailed models of large-scale real indoor spaces and became a source for many later simulation platforms.

In the second stage, navigation and instruction datasets shifted the focus from scene modeling to language-guided behavior. Representative datasets include R2R, CVDN, and REVERIE. Their key innovation was to bind natural language instructions to visual observations and navigation trajectories, creating a three-way coupling among instruction, perception, and action. This made embodied intelligence research more task-oriented and more closely tied to human commands.

In the third stage, manipulation and interaction datasets systematically introduced physical interaction and long-horizon tasks. Representative datasets include ALFRED, iGibson, and ManiSkill. The main breakthrough was the expansion from navigation goals to long-horizon task completion. However, this stage still faced limited data scale, weak task generalization, and a lack of unified representation across different robots. These problems pushed the field toward scale and generalization.

In the fourth stage, large-scale and foundation-model-oriented datasets began to emerge. Representative datasets include ManiSkill2, BEHAVIOR-1K, and Open X. Data scale and task complexity increased significantly. ManiSkill2 expanded manipulation diversity and physical accuracy. BEHAVIOR-1K extended household activities to thousands of scenes and complex behavior combinations. At the same time, the fusion of simulation data and real data became an important research direction. Yet core bottlenecks remained: high data acquisition cost, insufficient long-tail task coverage, and static data distributions that were difficult to extend.

In the fifth stage, world models and generative simulation have driven rapid evolution. Genie 2 and Genie 3 function as interactive foundation world models that can autoregressively generate operable, spatiotemporally consistent 3D interactive environments from a single image or text prompt. Marble is oriented toward persistent 3D world construction. It can generate structurally complete 3D scenes from reference images or text, allow free exploration, and export scenes in standard formats for downstream physical simulation. V-JEPA 2 is a video-based self-supervised world model that learns physical dynamics in latent space through joint embedding prediction. It has future-state prediction capability and zero-shot robot planning potential. These world models push embodied intelligence simulation data from collection and reconstruction toward on-demand generation. However, their generated content still faces challenges in physical consistency, long-horizon controllability, and reliable simulation-to-real fusion. This is precisely why the proposed framework uses Marble for scene generation and introduces NVIDIA Isaac Sim for physical constraint validation in a dual-engine design.

4. Core Technical Challenges for Embodied Intelligence in Home Care

The study identifies three core technical challenges. The first is the high-fidelity automated generation of residential elderly care scenes. Such scenes must be realistic, reasonable, and physically interactive. They must include age-friendly facilities and compact layouts, not generic young-family templates. The second is closed-loop control and data collection for safety-critical tasks such as emergency response. This requires event triggering, robot control, and data annotation to be automated end to end. The third is safe modeling of care actions. Force feedback, collision constraints, and other physical rules must be embedded into action representation and data generation. These challenges cannot be solved by simply scaling up generic embodied intelligence datasets.

There is also a semantic distribution shift between general datasets and elderly care. Current mainstream datasets focus on short, standardized actions such as tabletop grasping and object rearrangement. Elderly care tasks are often long-horizon, multi-step, and highly context-dependent. They include helping an older adult sit up, delivering medicine at the right time, guiding a person to a safe place, or responding to a fall. Existing datasets lack demonstrations and annotations for these tasks. In addition, the interaction object is different. In general manipulation, the object is usually rigid or semi-rigid. In elderly care, the core interaction object is the older adult’s body. Finally, existing indoor scenes are mostly standardized residences or laboratory environments. They do not fully consider the special characteristics of age-friendly living spaces, such as compact layouts, care-specific facilities, highly personalized item placement, and low-light conditions. These features directly affect robot navigation, perception robustness, and manipulation space planning.

5. How Existing Embodied Intelligence Simulation Datasets Compare

The study systematically compares mainstream indoor embodied intelligence simulation datasets across simulation space type, scenario, core dataset, scene scale, task complexity, and robot configuration. The comparison shows that general residential datasets have matured around different robot hardware configurations, including fixed-base single-arm robots, mobile-base single-arm robots, pure mobile platforms, and dual-arm or mobile dual-arm robots. However, they are designed for ordinary household environments and do not cover age-friendly facilities, older-adult behavior patterns, or elderly-care service tasks. Age-friendly residential datasets remain in an early exploratory stage. They have begun to address elderly care but still suffer from limited scene coverage, coarse older-adult behavior modeling, single-dimensional care tasks, and missing multi-agent interaction.

Table 1. Comparison of embodied intelligence simulation datasets for residential and healthcare spaces
Simulation space Scenario Core dataset Scene scale Task complexity Robot configuration Core strength Limitation
Residential space General residential simulation RLBench Configurable manipulation scenes 100 graded-difficulty tasks Fixed-base single arm Standard benchmark, task grading Single scene type, no mobility
Residential space General residential simulation ManiSkill Configurable manipulation scenes 100+ multi-type tasks Fixed-base single arm Diverse tasks, configurable scenes Lacks long-horizon tasks
Residential space General residential simulation ProcTHOR 10,000 scenes Generalization tasks Mobile-base single arm Procedural generation, large scale Limited object interaction precision
Residential space General residential simulation ALFRED 120+ scenes 25,000+ instruction trajectories Mobile-base single arm Large instruction-trajectory scale Discrete actions
Residential space General residential simulation Gibson 572 scanned spaces Generalization tasks Pure mobile platform Real scans, high fidelity Navigation only, no manipulation
Residential space General residential simulation BEHAVIOR-1K 50 complete residences 1,000+ complex activities Dual-arm configuration Broad coverage: complete homes, long-horizon complexity High computational cost
Residential space General residential simulation AgiBot World Alpha 100+ scenes 92,000+ real robot trajectories Mobile dual-arm Large real-robot trajectory scale High collection cost
Residential space Age-friendly residential simulation ElderSim 4 background scenes 55 tasks Mobile or humanoid Focus on older-adult behavior and age-friendly environments Scarce background scenes, single task dimension
Residential space Age-friendly residential simulation KIST SynADL 5 core scenes 55 activity categories Mobile-base single arm Scalable synthesis Lacks physical interaction and care manipulation
Healthcare space Care-area simulation Isaac for Healthcare Generalizable scenes Generalization tasks Fixed-base single arm High-fidelity care manipulation simulation Fixed base, no mobility
Healthcare space Care-area simulation Syn Mediverse 13 types, 31 semantic categories Mobile-base single arm Rich multimodal semantics, mobile care scene understanding Limited interaction tasks, not residential elderly care

The table shows that general residential datasets have achieved maturity in scene scale, task diversity, and hardware adaptation, but they model ordinary young-family environments. They do not cover age-friendly facilities, older-adult behavior patterns, or elderly-care service tasks. Age-friendly datasets such as ElderSim and KIST SynADL have made initial progress, but they still lack scene coverage, fine-grained older-adult behavior modeling, multi-dimensional care tasks, and multi-agent interaction. They are also often focused on action recognition or activity synthesis rather than closed-loop embodied intelligence. This leaves a significant gap for embodied intelligence in residential elderly care.

6. Four Main Generation Paradigms for Embodied Intelligence Datasets

The study also reviews four mainstream generation paradigms for embodied intelligence datasets. Each paradigm has different trade-offs in scalability, physical fidelity, and task controllability. The first is physics-simulation-based procedural generation. This paradigm uses high-fidelity physics engines and large-scale parallel computation. It randomizes scene layout, object pose, and physical properties, then drives agents to execute interaction tasks automatically. Representative works include RLBench and ManiSkill. Its strengths are high throughput, perfect ground-truth annotation, and reproducibility. Its main limitation is the sim-to-real gap, especially in contact modeling and high-frequency physical effects.

The second paradigm is teleoperation-based data collection. Human operators use VR or haptic devices to control real or virtual robots. This captures natural interaction trajectories and expert demonstrations. Representative works include RoboCasa and MimicGen. The data naturally satisfies real physical constraints and carries high behavioral rationality. However, it is labor-intensive, low-throughput, and subject to operator differences and communication latency. It is difficult to scale to millions of episodes.

The third paradigm is 3D reconstruction and digital twins. It uses NeRF, 3D Gaussian Splatting, and related techniques to scan and model real indoor scenes with high visual fidelity. Representative works include RoboTwin and D3Fields. It provides photorealistic rendering and real-world alignment, which is useful for safe policy testing. However, reconstruction is time-consuming, and the resulting scenes often lack accurate physical properties such as mass, friction, and joint damping. Interaction reproducibility and diversity are limited by reconstruction accuracy.

The fourth paradigm is generative-model-driven data synthesis. It uses large language models, multimodal large models, or diffusion models to generate tasks, scenes, or trajectories. Representative works include GenSim, RoboGen, UniSim, and Genie. One path combines large models with simulators for semi-supervised generation. Another path directly generates visual frames and action trajectories without an explicit simulator. The direct path has high scalability and visual diversity but often violates physical laws, lacks precise state annotation, and offers limited controllability. For residential elderly care, where safety and physical fidelity are critical, multi-paradigm fusion is becoming a trend. The proposed method follows this direction by combining Marble world-model scene generation with Isaac Sim physical validation.

Table 2. Multi-dimensional comparison of mainstream generation paradigms for embodied intelligence datasets
Generation method Core mechanism Representative work Scalability Physical fidelity Main advantage Key limitation
Physics-simulation-based procedural generation Simulation randomization and execution RLBench, ManiSkill High High Perfect ground-truth annotation, reproducible experiments, large-scale parallelization Sim-to-real gap, limited generalization
Teleoperation-based data collection Human teleoperates robot RoboCasa, MimicGen Low High Captures human strategies, high semantic quality, real action demonstrations Labor-intensive, low throughput, large operator differences
3D reconstruction and digital twin NeRF, 3DGS, and related techniques RoboTwin, D3Fields Medium Medium Photorealistic rendering, real-world correspondence, safety policy testing Time-consuming reconstruction, missing physical properties, interaction limits
Generative model synthesis Large language or vision models drive generation GenSim, RoboGen High Medium Task diversity, minimal human design, leverages world knowledge Requires verification pipeline, depends on LLM capability, prone to hallucinated tasks
Diffusion or video model synthesis Diffusion or video models UniSim, Genie High Low No explicit simulator required, scalable pretraining, strong visual diversity Often violates physics, no precise state annotation, limited controllability

7. A Four-Tier Construction Architecture for Embodied Intelligence in Elderly Care

The proposed method addresses the three core challenges through a four-tier architecture. The architecture is driven by two top-level service requirements: emergency response and daily care. All layers are designed around the actual interaction logic of elderly care. The architecture uses a bottom-up progressive construction model. Lower layers provide foundational support, upper layers define task orientation, and all layers exchange data and control signals through standardized interfaces. This creates a complete data flow and control flow from virtual home scene reconstruction to multimodal data annotation. The final output is a high-quality, domain-specific embodied intelligence simulation dataset that covers robot environmental perception, behavioral decision-making, and motion control.

Table 3. Four-tier architecture for the proposed embodied intelligence simulation dataset
Layer Primary function Key technologies Output
Scene generation layer Generate high-fidelity, diverse 3D home environments Marble world model, text and image prompts, OpenUSD Age-friendly home layouts, furniture, appliances, assistive devices, scene variants
Physical modeling layer Build physically accurate scenes and robot systems NVIDIA Isaac Sim, URDF, rigid body dynamics, contact mechanics Physical properties, collision meshes, robot models, older-adult human models, safety constraints
Control execution layer Drive robots to perform elderly-care tasks High-level planning, low-level control, domain randomization Navigation, manipulation, emergency response, daily care behaviors
Data acquisition layer Extract, align, annotate, and store multimodal data Time synchronization, multimodal recording, elderly-care labels RGB, depth, proprioception, scene state, fall labels, contact force, care semantics

The scene generation layer is the foundation. It uses the Marble world model pipeline to automatically generate complete 3D home scenes from 2D reference images or text descriptions. The output includes floor plans that match real Chinese residential elderly-care spaces, indoor objects with physically realistic materials and textures, and scene variants for different tasks. All generated scenes are exported in OpenUSD format to ensure seamless integration with downstream physics engines. The layer also supports parametric control of scene elements for later large-scale domain randomization. On top of Marble’s generation capability, the method injects age-friendly spatial priors. It builds an asset library that includes handrails, nursing beds, walkers, and anti-slip mats. It incorporates generation constraints such as compact floor plans, barrier-free circulation, and low illumination. This makes the generated scene distribution closer to real residential elderly care in China, rather than a generic young-family template.

The physical modeling layer receives the 3D scene assets and performs physics-level modeling in NVIDIA Isaac Sim. It assigns accurate physical properties to all objects, including mass, friction coefficient, restitution coefficient, and collision volume. It builds rigid body dynamics and contact mechanics that follow real physical laws. For high-frequency interaction objects in elderly care, it performs refined collision mesh modeling and physical parameter calibration. The layer also builds robot systems. Depending on the task, it imports URDF models of Unitree robots and configures joint degrees of freedom, actuator parameters, and sensors. It also builds human models and posture models. It creates parametric older-adult human models with body dimensions, limited joint ranges, and typical fall postures. It introduces human-robot safety constraints, including contact force limits, collision thresholds, and compliant contact models, to adapt to older adults as fragile interaction partners.

The control execution layer connects the simulation environment with data acquisition. It binds general motion and manipulation strategies to elderly-care tasks. For emergency response, it introduces fall event triggering and safe navigation constraints. For daily care, it introduces force control and obstacle avoidance constraints for operations near the human body. The layer uses a hierarchical control architecture with high-level decision planning and low-level motion control. It configures different algorithm combinations for different tasks. It also applies domain randomization, such as randomizing object positions, lighting conditions, robot initial states, and terrain parameters, to enhance data diversity and generalization.

The data acquisition layer systematically extracts, aligns, annotates, and stores multimodal data. It synchronously records RGB images, depth maps, proprioceptive data, and other sensor outputs from the robot. It covers both first-person robot views and global monitoring views. Beyond general multimodal records, it adds elderly-care-specific annotation fields, including fall posture category, rescue target point, human-robot relative distance and bearing, care task semantics, and contact force. This layer is essential for producing embodied intelligence datasets that can support model training, policy validation, and safety analysis.

8. Emergency Response with Quadruped Robots

To support emergency response in residential elderly care, the study designs a dedicated data construction pipeline for quadruped robots. The pipeline uses the Marble-generated 3D home scenes and the Unitree Go1 as an example. The dataset focuses on high-frequency emergency events, especially falls among older adults. The quadruped robot is responsible for whole-home mobility and environmental monitoring. This ensures that the dataset covers perception, motion, and interaction data for emergency scenarios.

Data collection uses a simulation-generation fusion mode. It relies on the standardized Marble-USD scene and the IsaacLab simulation engine to achieve scalable, high-fidelity generation. Through the OpenUSD data exchange channel, it outputs high-resolution multimodal visual data, including RGB images, depth maps, semantic segmentation maps, and 3D point clouds. The data covers different lighting conditions and different terrains, including flat and uneven surfaces. The physical modeling layer builds the interaction system between the robot and the emergency scene. It loads the Unitree Go1 quadruped robot model and configures perception sensors and physics engine parameters. At the perception level, RGB cameras and depth cameras are mounted to output multimodal visual data:

DSensor = {IRGB, Idepth}

This provides visual input for fall target recognition and navigation. At the physical level, the method calibrates joint damping, stiffness, and collision detection rules to ensure that the robot’s kinematics and dynamics in simulation are highly consistent with the physical robot. It also builds a physical model of the older-adult fall emergency scene. The state and spatial features of the fallen older adult are defined as:

Oelder = {χelder, σfall, ptarget}

Here, χelder is the 3D pose of the older adult in the fall posture. σfall ∈ {0,1} is the fall state label. ptarget is the target point where the older adult is located, which is the endpoint of the robot’s rescue navigation. After configuration, the physical modeling layer can receive rescue task instructions such as “go to the fallen older adult’s position.” It provides a credible robot-environment interaction interface for the control execution layer and provides multimodal perception and state data sources for the data acquisition layer.

The control execution layer in this scenario uses a three-level architecture: high-level perception and planning, mid-level command conversion, and low-level motion control. The high-level perception and planning module is based on the NaVILA vision-language navigation model. It takes the multimodal visual data and rescue task instructions from the physical modeling layer as input. It extracts visual features from RGB and depth images using a convolutional neural network. It combines the target point and state of the fallen older adult. It maps natural language rescue instructions to a global path point sequence W through a vision-language alignment mechanism. The path planning process is formalized as a maximum a posteriori probability estimate:

W* = argmaxW P(W | Irgb, Idepth, Itask, τgoal)

The mid-level command conversion module receives the path point sequence and task constraints from the high-level module. It decouples the abstract path planning result into a standardized command set that can be executed by low-level motion control:

U = {vctrl(t), ωctrl(t) | t ∈ [0, Tnav]}

Here, vctrl(t) is the desired linear velocity command at time t. ωctrl(t) is the desired angular velocity command. Tnav is the total navigation time to the older adult’s position. Smooth interpolation and obstacle avoidance optimization ensure command continuity and safety. The low-level control uses the RSL_RL framework for Unitree Go1. It uses an Actor-Critic neural network policy trained with proximal policy optimization, or PPO. The policy converts mid-level velocity commands into torque control signals for the robot’s 12 joints. This enables stable and agile quadruped locomotion on complex home terrain. The low-level motion policy uses an independently parameterized Actor-Critic dual-network architecture. The policy network, or Actor, maps observations to a probability distribution over joint actions. The model uses a three-layer fully connected neural network to output action means, while the standard deviation parameter is independently learnable. The value network, or Critic, uses the same network topology but independent parameters. It outputs a single scalar representing the state value estimate, which is used for advantage function calculation. Training uses synchronous batch mode. A single iteration can execute policy sampling in thousands of parallel simulation instances simultaneously, significantly accelerating data collection efficiency.

The data acquisition layer synchronously records heterogeneous data streams at each simulation time step. It uses the global clock provided by the simulation engine for time alignment. The visual perception data stream records first-person image data from the RGB camera and depth camera mounted on the robot head at 20 Hz. It also configures a global bird’s-eye view monitoring camera to record scene overview images at the same frequency. This is used for data annotation verification and algorithm behavior visualization. The proprioceptive data stream synchronously records the robot’s complete proprioceptive state vector at a 50 Hz control frequency. This includes position, velocity, and torque measurements of the 12 joints, as well as the base three-axis angular velocity and acceleration from the IMU sensor. The scene state data stream records the time-varying states of all dynamic objects in the simulation environment. Its core content includes the 3D pose of the fallen older adult, the fall state label, and the Euclidean distance and relative bearing between the robot base and the target older adult. The fall state label includes side-lying, supine, or prone.

Table 4. Emergency response data streams for the quadruped robot pipeline
Data stream Frequency Content
Visual perception 20 Hz First-person RGB and depth images, global bird’s-eye view
Proprioception 50 Hz 12 joint positions, velocities, torques, IMU base angular velocity and acceleration
Scene state Synchronized with simulation clock 3D pose of fallen older adult, fall state label, robot-to-target distance and bearing

9. Daily Care with Bipedal Humanoid Robots

For daily care, the study designs a data construction pipeline for bipedal humanoid robots. The pipeline uses the Unitree G1 as an example. The dataset focuses on high-frequency daily assistance needs in residential elderly care, including meal assistance, medication management, item organization, and delivery. The construction flow begins with the scene generation layer, which also uses the Marble pipeline. It focuses on home environments with dense functional manipulation spaces, including kitchen areas, bedrooms, and living rooms. During scene generation, the method pays special attention to fine-grained modeling of manipulable objects. Each graspable object is configured with accurate geometric dimensions, mass, center of mass, surface friction coefficient, and grasp keypoint annotations. This ensures the physical realism of subsequent robot manipulation interactions.

The generated care scenes are imported into Isaac Sim in OpenUSD format for full-scene physical property configuration. For daily care scenarios, the physical modeling layer focuses on rigid body dynamics modeling of interactive objects. This includes accurate collision mesh generation, friction cone parameter calibration, and contact stiffness settings between objects. These ensure the physical realism of grasping, placing, pushing, and pulling interactions. The physical modeling layer also builds the Unitree G1 humanoid robot system. It imports the G1 URDF model, configures joint drive parameters for 23 degrees of freedom across the body, and mounts a head RGB-D camera for scene perception and object detection. It also mounts wrist RGB cameras for close-range manipulation guidance, torque sensors for grasping force perception, and whole-body joint encoders.

The control execution layer in this scenario combines the Ψ0 (Psi-Zero) end-to-end learning framework with large language model-driven task planning. This enables end-to-end execution from natural language care instructions to coordinated whole-body robot operations. Ψ0 is a general control policy specifically designed for humanoid robots. It can uniformly handle coupled locomotion and manipulation control. In daily care scenarios, it completes fine-grained dexterous operations such as grasping, carrying, and placing. The Ψ0 architecture follows a layered design consisting of a VLM backbone, an action expert, and a lower-limb controller. The three systems have distinct roles and work together. System-2, the VLM backbone, is responsible for high-level perception and semantic understanding. System-1, the action expert, is responsible for whole-body action generation. System-0, the lower-limb controller, is responsible for low-level mobile stability. This architecture is suitable for complex mobile manipulation tasks in home care. It can maintain whole-body balance and stable locomotion while ensuring fine upper-limb manipulation.

System-2 is the highest-level decision module of Ψ0. It is built on the Qwen3-VL-2B-Instruct pretrained vision-language model. It handles scene understanding, object recognition, and task semantic parsing. The module receives multi-view visual observations from the physical modeling layer and natural language care instructions Tcare. It outputs a unified multimodal feature representation Zvlmc to provide semantically rich input for downstream action generation. The head RGB-D camera outputs RGB images and the two wrist RGB cameras output images. These are divided into fixed-size image patches, linearly embedded, positionally encoded, and fed into a multi-layer Transformer encoder. The visual encoder inherits Qwen3-VL pretrained weights on large-scale image-text pairs. It has open-world object recognition capability and can accurately identify daily objects in home environments, such as medicine bottles, water cups, and tableware. The language encoding process tokenizes care instructions and generates text features through the Qwen3-VL language encoder:

Zclang = TextEncoderQwen(Tcare)

System-1 is the mid-level execution module of Ψ0. It uses a multimodal diffusion Transformer, or MM-DiT, architecture. It maps the multimodal features output by System-2 into whole-body action sequences for the robot. System-0 is the low-level execution module of Ψ0. It is built on an AMO, or adaptive model-based optimization, reinforcement learning tracking policy. It maps the 8-degree-of-freedom mobile commands output by System-1 into 15-degree-of-freedom lower-limb joint angles. This ensures whole-body balance and mobile stability while the humanoid robot performs upper-limb manipulation tasks.

The data acquisition layer synchronously records heterogeneous modal data streams during simulation. It uses the Isaac Sim global simulation clock for strict timestamp alignment, forming a standardized care task dataset. Because of the three-level Ψ0 architecture, the data acquisition layer must synchronously record the input and output states of each system to support model training, inference analysis, and system-level debugging. The visual perception data stream records image data from the multi-view camera system mounted on the G1 robot at 10 Hz. The proprioceptive data stream records the complete proprioceptive state of the G1 robot at a 100 Hz control frequency. This includes joint states, IMU data, end-effector states, and foot states.

Table 5. Daily care data streams for the humanoid robot pipeline
Data stream Frequency Content
Visual perception 10 Hz Multi-view camera images from head RGB-D and wrist RGB cameras
Proprioception 100 Hz Joint states, IMU data, end-effector states, foot states
System-level states Synchronized with simulation clock Inputs and outputs of System-2, System-1, and System-0

After this pipeline, the bipedal robot can perform daily care tasks in the home scene. The dataset covers diverse care task instructions, high-fidelity multi-view visual observations, precise whole-body robot states and action sequences, complete internal state records of the three-level Ψ0 system, and detailed object interaction states. This makes it suitable for training and evaluating embodied intelligence policies for humanoid care robots.

10. How the Proposed Embodied Intelligence Dataset Compares with Existing Work

The study emphasizes that its current contribution is mainly the methodology and architecture design rather than a fully mass-produced dataset. It has completed the overall technical plan, the four-tier layered framework, the dual-engine collaboration logic, and the task pipelines for two types of robots. Because of real-world safety risks and challenges, large-scale dataset production has not yet been completed. The evaluation therefore focuses on theoretical qualitative comparison and the rationality of the architecture design.

Compared with existing age-friendly datasets, the proposed method has three main advantages. First, in scene generation, existing age-friendly datasets rely on manual modeling, which makes scene expansion costly. The proposed method uses the Marble world model to achieve automated, parametric scene generation driven by text and images. This addresses the root cause of limited sample diversity in elderly-care scenes. Second, in physical constraints, existing methods mostly perform human motion animation rendering and lack human-robot contact mechanics and older-adult protection thresholds. The proposed method uses Isaac Sim to calibrate physical parameters for all objects and embeds human-robot safety limits into the modeling process. This matches the safety requirements of close-range interaction in elderly care. Third, in task design, current age-friendly datasets lack closed-loop robot interaction workflows. The proposed method separates quadruped emergency rescue and bipedal daily care into two independent pipelines. It designs control chains around real elderly-care business tasks and achieves a full closed loop from scene to modeling, control, and data acquisition.

From a top-level design perspective, the proposed construction approach addresses three major shortcomings of existing elderly-care datasets. It offers advantages in scene scalability, physical realism, and task practicality. At the same time, the entire generation pipeline uses OpenUSD, a common industrial format, for data interoperability. The generated data format is compatible with mainstream humanoid and quadruped robot simulation development frameworks and physical robot driver interfaces. This means that after connecting to physical prototypes, sim-to-real transfer testing can be carried out. For embodied intelligence, this interoperability is important because it reduces the cost of moving from simulation to real robots and supports reproducible benchmarking.

11. Limitations, Safety Considerations, and Research Outlook

The study acknowledges limitations. Because of real-world safety risks, the research and current related work on residential elderly-care embodied intelligence simulation datasets cannot directly apply simulation-trained robots to real elderly-care scenarios for large-scale validation. As a result, key sim-to-real transfer indicators cannot be fully verified in real scenes. This is a common challenge for embodied intelligence in safety-critical care domains. It also means that the proposed dataset construction method should be understood as a foundation for future validation rather than a finished deployment solution.

Future work is planned across several dimensions. The first is deeper sim-to-real transfer calibration. The plan is to collect visual, force, motion trajectory, and interaction data from real quadruped and bipedal robots in real home environments. An automated sim-to-real parameter calibration workflow will be built. Domain adaptation algorithms will be used to reduce the distribution gap between simulation and reality. Key transfer indicators such as emergency response latency, manipulation success rate, and navigation accuracy will be quantitatively verified.

The second direction is multi-robot collaborative care datasets. The plan is to design collaborative task paradigms in which quadruped robots perform whole-home inspection, emergency positioning, and bipedal robots perform fine manipulation and close-range care. The data collection will cover multi-robot perception sharing, dynamic task allocation, interactive obstacle avoidance, and communication synchronization. Annotation rules and evaluation benchmarks for multi-agent collaborative data will be developed to support multi-robot elderly-care service algorithms. At the same time, the work aims to promote open-source sharing and industrial deployment of datasets. It plans to release core data subsets and simulation deployment toolchains, provide baseline algorithms and data loading interfaces, and collaborate with elderly-care institutions and robot companies for real-world validation. The goal is to form industry application standards for residential elderly-care embodied intelligence datasets and accelerate the transition from laboratory research to scaled elderly-care services.

The third direction is to expand the coverage of residential elderly-care scenes and tasks. The plan includes tasks for older adults with different health conditions, such as semi-disabled, disabled, and advanced-age older adults living alone. New tasks include home rehabilitation training, chronic disease health monitoring, cognitive impairment companionship, secondary rescue after a fall, precise medication and feeding assistance, and intelligent regulation of age-friendly environments. The work will also cover different housing types, including small apartments, age-friendly renovated residences, and community-based elderly care homes. Parametric domain randomization will be used to generate massive scene variants and fill the data gap for long-tail risk scenarios.

12. Conclusion

Residential elderly care is non-structured, safety-critical, and task-specific. These characteristics mean that embodied intelligence data construction cannot simply reuse generic datasets. Simulation datasets offer low generation cost and high scene controllability, making them a core path for solving the shortage of embodied intelligence data in elderly care. The proposed method addresses three core technical challenges through scenario-specific customization and dual-engine collaboration. It uses a four-tier standardized architecture covering scene generation, physical modeling, control execution, and data acquisition. This enables a complete standardized pipeline from high-fidelity 3D home scene generation to multimodal interactive data output.

The method is supported by the Marble world model and the NVIDIA Isaac Sim physics engine. It achieves scene modeling that fits the spatial characteristics of residential elderly care, human-robot interaction modeling that adapts to older-adult physiological characteristics, and differentiated control architecture design for core tasks such as emergency response and daily care. The resulting specialized dataset covers cross-robot embodiments and multiple dimensions of robot perception, decision-making, and control. It helps fill the gap in embodied intelligence simulation data for residential elderly care. It also addresses the adaptability limitations of general embodied intelligence datasets in elderly-care scene features, task settings, and human-robot safety interaction constraints.

Future work can further expand scene and task coverage to form a full-scene, full-task residential elderly-care embodied intelligence simulation data system. It can deepen sim-to-real transfer research by combining real measured data from home scenes with simulation datasets. Through virtual-real fusion, the transfer performance of embodied intelligence models from simulation to real scenes can be improved. At the same time, multi-robot collaborative elderly-care dataset construction can be carried out to support multi-robot service needs. For embodied intelligence, the ultimate measure of success will not be the size of a dataset alone, but whether that data enables safe, reliable, and useful robots in the homes of older adults.

Scroll to Top