Expert Strategy Data Collection and Environmental Impacts in Embodied Intelligence: Dual ViA Segmented Sampling Under Controlled Illumination, Background and Tabletop Conditions

A newly reported study in the field of embodied intelligence has demonstrated that the performance of embodied intelligence models depends not only on model architecture and training algorithms but also, to a decisive degree, on how demonstration data is collected and on how carefully the environmental conditions during collection are controlled. The research, conducted by a team from the School of Information Engineering at Shenyang University of Chemical Technology, the Shenyang Institute of Automation of the Chinese Academy of Sciences, and the Innovation Institute of Intelligent Robotics Shenyang Co., Ltd., proposes a demonstration data collection strategy based on dual ViA (Vision in Action) segmented sampling and combines it with controllable settings for illumination, background and tabletop conditions to construct comparative datasets. Evaluated on an industrial raw-material grasping and placing task executed by a dual-arm collaborative robotic platform with clutter interference, the combination of dual ViA segmented sampling and environmental condition control raised the model grasp success rate by 52 percent and the placement success rate by 61 percent, while reducing task execution position deviation by 70 percent during the validation stage compared with data collected using the dual ViA segmented sampling strategy alone under conventional environmental conditions. The findings offer a structured methodology for building high-quality demonstration datasets for embodied intelligence and provide quantitative evidence that expert strategy organization and environmental condition control jointly improve data quality and validation performance in complex manipulation scenarios.

  • 1. Why embodied intelligence depends on high-quality demonstration data
  • 2. Multimodal datasets for embodied intelligence and their limitations
  • 3. Existing data collection methods for embodied intelligence
  • 4. The dual ViA segmented sampling strategy
  • 5. Illumination condition control for stable visual input
  • 6. Background condition control for target discriminability
  • 7. Tabletop condition control for boundary clarity
  • 8. Experimental setup on a dual-arm collaborative platform
  • 9. Results: expert strategy versus conventional data collection
  • 10. Results: illumination condition control
  • 11. Results: background condition control
  • 12. Results: tabletop condition control
  • 13. Results: joint environmental condition control
  • 14. Conclusions and future work for embodied intelligence

1. Why embodied intelligence depends on high-quality demonstration data

As embodied intelligence technology advances rapidly, robotic systems are being asked to deliver ever higher performance in flexible operation, autonomous task planning and adaptation to complex environments. To meet these demands, robotic systems are shifting from automation based on programmed control toward embodied intelligence agents that possess perception, cognition and decision-making capabilities. However, the performance of embodied intelligence models hinges not only on model structures and training methods but also on the support of high-quality demonstration datasets. Such datasets must not only cover rich multimodal information but also satisfy the requirements of model training and validation in terms of task coverage, sampling precision and data consistency.

At present, existing embodied intelligence datasets are concentrated mainly in household kitchen scenes, daily life activities, industrial production and healthcare. Yet demonstration data for long-sequence, high-precision dual-arm operations in specific industrial tasks remains relatively scarce, making it difficult to directly satisfy the training and validation needs of models in complex manipulation scenarios. As the complexity of embodied intelligence tasks increases, existing standard datasets struggle to cover the long-sequence, high-precision operational requirements found in particular industrial settings. Building high-quality demonstration datasets for specific tasks has therefore become an important route to improving the adaptability and reliability of robotic systems working within embodied intelligence frameworks.

Prior research on visual imitation learning has shown that the generalization performance of robot policies is affected by changes in environmental conditions such as illumination and camera position. At the same time, existing data collection methods for embodied intelligence still fall short in the organization of expert strategies. In particular, the explicit division of key stages within long-sequence tasks is often insufficient, so critical operation samples are easily diluted by transitional segments, which in turn affects the inference accuracy and execution reliability of models on complex tasks. Moreover, variation in illumination, background and tabletop conditions changes the distribution of visual input, which then influences target recognition, action alignment and placement localization. It is therefore necessary to set environmental conditions in a controllable manner and to analyze their impact on model validation performance.

To address these problems, the reported research takes an industrial raw-material grasping and placing task as its object and builds a demonstration data collection workflow from two directions: expert strategy organization and controllable environmental condition acquisition. The study then analyzes how model validation performance changes under different collection conditions. During the data collection stage, a dual ViA segmented sampling strategy is used to organize the task flow into stages and to achieve temporal synchronization of visual and action information so as to improve the validity of data in key phases. At the same time, illumination, background and tabletop conditions are set in a controllable way and comparative datasets are constructed to compare model performance under different environmental conditions. On a dual-arm collaborative robotic platform performing the industrial raw-material grasping and placing task, grasp success rate, placement success rate and task execution position deviation are adopted as evaluation metrics to compare model behavior under different collection strategies and environmental conditions, providing experimental evidence for analyzing how expert strategy organization and environmental condition control affect demonstration data quality and model validation performance within embodied intelligence.

2. Multimodal datasets for embodied intelligence and their limitations

In recent years, multimodal datasets for embodied intelligence have played a key role in the perception, decision-making and task execution of intelligent agents. These datasets fuse data from robotic arms, dexterous hands, vision, audition and other sensors together with language descriptions, providing a data foundation for agents to acquire environmental representations and semantic information, and helping to improve the generalization ability and interactive adaptability of policy models. To advance cross-modal perception and unified evaluation, several research teams have constructed representative datasets.

Open X-Embodiment, built jointly by multiple research institutions, covers 22 embodied platforms and provides 160,000 task demonstration episodes under a unified RLDS standard, supporting cross-platform training and evaluation, promoting data standardization and facilitating the development of general-purpose robot models. RoboMIND, proposed by Wu and colleagues, collected 107,000 demonstration trajectories through human teleoperation, covering 479 tasks and 96 object categories, and includes 5,000 failure cases that support task reflection and correction, thereby promoting learning of multi-object interaction and long-horizon tasks. ARIO, proposed by Wang and colleagues, integrates real and simulated data from 258 robot platforms, provides three million demonstration episodes and 321,064 tasks, and adopts a timestamp mechanism to improve the diversity and adaptability of task execution. A kitchen task dataset proposed by Ren and colleagues records tasks performed by 20 participants using 17 tools, providing 680 data segments and 56,000 annotations covering tactile, electromyographic, audio and eye-tracking data, which promotes robot skill learning in fine manipulation tasks. The Kaiwu dataset proposed by Jiang and colleagues contains 11,664 multimodal demonstration instances covering hand motion, manipulation pressure and sound, and is designed specifically for complex assembly tasks, supporting research on robot learning and hand-eye coordination.

Although existing multimodal datasets have made significant progress in improving the perception and task execution capabilities of embodied robots, their research focus is mainly on standardized scenes and simplified tasks, with limited support for specific tasks and scenarios. This restricts the adaptability and responsiveness of embodied robots in complex environments. Consequently, when addressing specific industrial manipulation tasks, it remains necessary to build targeted demonstration datasets that combine task objectives and environmental conditions in order to improve execution stability and the adaptability of models in complex environments.

3. Existing data collection methods for embodied intelligence

Beyond dataset construction, data collection methods for embodied intelligence also concern how demonstration workflows are organized, how task environments are controlled and how collection efficiency affects policy learning. Building data collection workflows for specific tasks and specific environments has become an important route to improving the generalization ability of robot policies. Unlike existing datasets that concentrate on standardized scenes and simplified tasks, the core objective of data collection for embodied intelligence is to provide the agent with comprehensive and accurate environmental perception information, helping it make reasonable decisions and execute tasks effectively in complex environments.

Around these issues, a variety of data collection methods have been proposed. DART, proposed by Park and colleagues, combines cloud-based simulation and augmented reality technology to further improve data collection efficiency. FastUMI, proposed by Liu and colleagues, provides 10,000 demonstration trajectories through hardware decoupling and a simplified algorithmic workflow, supports imitation learning algorithm training and reduces dependence on specific hardware platforms. The ViA (Vision in Action) system proposed by Xiong and colleagues imitates the active perception behavior of humans, adopting a six-degree-of-freedom robot neck and a virtual reality interface to synchronize perception and action, optimizing dual-arm robot manipulation performance, significantly improving the completion rate of visually occluded tasks, and training an active perception policy model through imitation learning. Adversarial data collection, proposed by Huang and colleagues, introduces visual, language and physical perturbations to improve data diversity and information density, reduces reliance on large-scale datasets, and significantly improves task execution performance when using only 20 percent of demonstration data, especially in understanding new task instructions, responding to environmental perturbations and recovering from errors. AIRSPEED, developed by Xia and colleagues, is an open-source general-purpose data production platform that supports multiple devices and simulation platform interfaces, reducing data collection cost and improving efficiency. DABI, proposed by Kobayashi and colleagues, is a data augmentation method for imitation learning based on bilateral control that downsamples image and robot state data to achieve roughly tenfold data expansion, significantly improving task success rates and validating its effectiveness in dual-arm control imitation learning.

In summary, existing data collection methods for embodied intelligence focus mainly on collection efficiency, task coverage and hardware generality, but they still have shortcomings when applied to high-precision dual-arm manipulation tasks. First, the explicit organization of key stages in long-sequence tasks is insufficient, so key operation samples are easily diluted by transitional segments. Second, the coordination between active viewpoint control on the observation side and action constraints on the execution side receives insufficient consideration. Third, there is a lack of controllable comparative analysis of environmental conditions such as illumination, background and tabletop. It is therefore necessary to design a structured demonstration collection method for specific industrial manipulation tasks and to systematically analyze the influence of environmental conditions on model performance.

4. The dual ViA segmented sampling strategy

The reported work proposes a method for expert strategy data collection and environmental condition control oriented toward dual-arm robot manipulation tasks, focusing on the dual ViA segmented sampling strategy used during demonstration data collection and on how illumination, background and tabletop conditions affect model training and inference performance. The overall workflow begins with the dual ViA segmented sampling strategy, which improves the validity of data in key operational stages through task staging and visual-action synchronization. Next, illumination, background and tabletop condition control reduces interference from environmental variation in the visual input. Finally, comparative datasets are built from different collection conditions, and the influence of expert strategy and environmental conditions on task performance is analyzed through model training and inference validation.

To address the problem that key operational stages in long-sequence demonstration tasks are easily diluted by transitional segments and that observation information and action targets are easily mixed, the study proposes a dual ViA segmented sampling strategy. Without changing the continuous execution flow of the robot, this strategy organizes the task process into stages at the data collection level and jointly introduces observation-side viewpoint control and execution-side action constraints, synchronously recording observations, states, gripper states, master-side control inputs, executed actions and stage labels, thereby forming structured demonstration data with clear boundaries and unambiguous semantics.

In terms of segmented sampling, the complete demonstration process is divided according to the semantic goals of the grasping and placing task into eight stages: target localization, approach and alignment, gripper closure, lift and transfer, descent and alignment, gripper release, lift and retreat, and manipulator reset. The robot execution process remains continuous, but the data recording process is organized according to stage labels, so that key operational processes such as grasping alignment, stable clamping and placement alignment are explicitly marked within the long sequence, thereby weakening the interference of invalid transitional segments on policy learning.

On the observation side, the system uses the positional deviation of the target in the image plane as feedback to drive the wrist camera to make small viewpoint adjustments, keeping the target as close as possible to the central region of the image, and suppresses abrupt viewpoint changes through smoothness constraints, reducing the effect of image jitter on the temporal continuity of observations. At the same time, the system constrains the camera observation distance and end-effector posture to ensure that the target and key operational regions remain within the effective field of view. Observation-side control acts mainly during target localization, lift and transfer, descent and alignment, gripper release, and manipulator reset, and is used to improve the stability of visual input in key stages.

On the execution side, the system organizes the key actions in the grasping and placing process step by step. The grasping stage sequentially completes approach and alignment and gripper closure, and briefly pauses after the gripper closes in order to stabilize clamping depth and contact state. The placing stage sequentially completes descent and alignment, gripper release, lift and retreat, and manipulator reset in order to clarify the action target of each stage. During execution, controllable motion patterns such as vertical descent, vertical lift and smooth transitions are preferred in order to reduce the influence of velocity discontinuities, posture deviation and contact disturbance on the demonstration trajectory.

To describe the dual ViA segmented sampling strategy in a unified way, the task stage state set is defined as a collection of eight elements denoted s1 through s8, corresponding to target localization, approach and alignment, gripper closure, lift and transfer, descent and alignment, gripper release, lift and retreat, and manipulator reset. At time t, the multi-view observation, robot state, gripper state, master-side robot control input and master-side gripper control input are recorded, with the master-side control input set formed by the robot arm control input and the gripper control input, which together constitute the demonstration control information at the current moment. Given the previous stage label, the current stage label is obtained through a stage decision function that updates the stage by rule-based judgement according to target visibility, the relative position between the end effector and the target, the gripper closure state and the current task completion status. Through this mechanism, the system can perform online judgement of task stages during continuous execution and ensure that stage boundaries remain consistent with the actual operation process.

During stage scheduling, the system uses stage labels as the main thread to trigger the corresponding control. According to the specific implementation of this study, the observation side participates mainly in scheduling during target localization, lift and transfer, descent and alignment, gripper release, and manipulator reset, so as to ensure that the target remains within the effective field of view during key stages. The execution side participates in scheduling during approach and alignment, gripper closure, lift and transfer, descent and alignment, gripper release, lift and retreat, and manipulator reset, so as to ensure the smoothness and consistency of key action processes. The control trigger variables of the observation side and the execution side are defined separately, with the observation-side trigger activated for the relevant stages and the execution-side trigger activated for its own set of stages. On this basis, the control outputs of the observation side and execution side are expressed by multiplying the trigger variable with the corresponding control generation function. The observation-side control output corresponds to the pose adjustment command of the observation-side manipulator end effector, used to drive the camera fixed on the wrist to complete viewpoint adjustment, while the execution-side control output corresponds to the end-effector motion and gripper control commands. Through this stage-driven approach, observation-side and execution-side control can work cooperatively around key stages, ensuring a stable correspondence between visual input, robot state and control output.

In terms of data organization, the algorithm synchronously records observations, states, gripper states, master-side control inputs, actions and stage labels under a unified timestamp basis, and represents the structured sampling unit at time t as a tuple of these elements. The complete episode corresponding to a single demonstration task is then represented as a sequence of sampling units, where the total sampling length determines the size of the episode. By introducing structured sampling units and stage labels, the system not only preserves the complete task flow but also explicitly marks the positions of key stages in the sequence, so that subsequent models can more clearly identify key operational processes and reduce interference from invalid transitional segments.

The dual ViA segmented sampling scheduling algorithm takes as input the multi-view observation stream, the robot state stream, the gripper state stream, the master-side robot control stream and the master-side gripper control stream, and outputs a structured episode with stage labels. The algorithm initializes the stage label and the episode, then loops while the demonstration task is not finished: it reads the current observation, robot state, gripper state and master-side control inputs, constructs the master-side control input, issues an error warning and moves to the next moment if the current observation or state data is abnormal, updates the current stage label according to the stage decision function, computes the observation-side trigger variable and the execution-side trigger variable according to the stage, executes observation-side control when the observation-side trigger is active, and executes execution-side control when the execution-side trigger is active. The algorithm records the action information at the current moment, constructs the structured sampling unit, appends it to the current episode, and repeats the process until the task ends, finally outputting the structured episode.

In summary, the role of the dual ViA segmented sampling scheduling algorithm is not to change the order of task execution but to establish a unified stage-driven scheduling mechanism for long-sequence demonstration collection. The algorithm can coordinate observation-side viewpoint control and execution-side action constraints around stage boundaries and complete the structured organization of multi-source data under continuous execution conditions, thereby translating the dual ViA segmented sampling strategy into an executable and synchronizable unified collection process that improves the validity, consistency and reproducibility of demonstration data for embodied intelligence.

5. Illumination condition control for stable visual input

During demonstration data collection, illumination conditions directly affect camera imaging quality and the stability of depth observation. If the illumination of the work area fluctuates with time, weather and external environment, it is easy to introduce interference such as shadows, reflections and uneven brightness distribution, causing differences in visual input between different collection batches and thereby affecting the comparability of subsequent model training and inference validation. To address this, the study designs an illumination condition control method whose core idea is to transform the originally randomly fluctuating illumination factor into a controllable, quantifiable and reproducible environmental variable, maintaining a stable illumination distribution in the work area and providing a consistent visual observation basis for demonstration collection.

In implementation, illumination condition control focuses on illumination intensity, irradiation range and shadow interference so that the target region maintains a relatively uniform and stable illumination state during collection. On the one hand, by optimizing the position and direction of light sources, interference from local shadows and highly reflective regions on target boundaries, surface texture and depth observation is reduced. On the other hand, by constraining the overall brightness distribution of the work area, observation deviations caused by illumination fluctuations between different collection batches are reduced, thereby improving the consistency of visual input at stage granularity.

To quantitatively evaluate the uniformity of illumination distribution in the work area, the coefficient of variation is adopted as a measure of illumination intensity fluctuation. The coefficient of variation is defined as the ratio of the standard deviation of illumination values in the sampling region to their mean, where the mean and standard deviation are computed over all pixels in the sampling region and each pixel value represents the illumination value of that pixel. This indicator characterizes the degree of dispersion of the illumination distribution relative to the average brightness. A smaller value indicates a more uniform illumination distribution in the work area and more stable visual observation. The indicator provides a unified quantitative basis for comparing lighting states under different illumination conditions.

6. Background condition control for target discriminability

During demonstration data collection, background conditions directly affect the distinguishability of the target object and the stability of visual observation. If background color, texture or complexity changes between different collection batches, it easily causes blurred target boundaries, reduced regional contrast and a shift in the distribution of visual input, which in turn affects the comparability of subsequent model training and inference validation. To address this, the study designs a background condition control method that constrains background color, texture and complexity, enhances the visual distinction between the target region and the background region, and reduces the influence of background variation on the distribution of demonstration data.

In implementation, background condition control focuses on background color, texture complexity and the degree of visual interference so that the target region maintains stable identifiability during collection. On the one hand, by constraining background color and texture features, interference from the background region on target boundaries is weakened. On the other hand, by controlling the visual complexity of the background, the target boundary remains relatively stable across different collection batches, reducing interference from the background region on target recognition.

To quantitatively evaluate the color similarity between the target region and the background region, a matching degree indicator is introduced as a measure of the effectiveness of background condition control. The matching degree is computed from the histogram values of the background region image and the target region image across color components, normalized by the sum of the target region histogram values over all dimensions. This indicator characterizes the degree of similarity between the target region and the background region in terms of color distribution. A larger value indicates that a higher proportion of the background region shares the same color as the target region, meaning the target and background are more similar and the target boundary is harder to distinguish. A smaller value indicates a more obvious color difference between the target region and the background region, meaning the target is more distinguishable. The indicator provides a quantitative basis for comparing the color similarity between target and background under different background conditions.

7. Tabletop condition control for boundary clarity

During demonstration data collection, tabletop conditions directly affect the distinguishability of the target object and the stability of visual observation. If tabletop material, color or texture changes between different collection batches, it easily causes local reflection, boundary noise and brightness contrast fluctuation, which then leads to blurred target boundaries, a shift in the distribution of visual input and inconsistent spatial references during the placement stage. To address this, the study designs a tabletop condition control method that constrains tabletop material, color and texture complexity, reducing the influence of reflection, boundary noise and brightness contrast fluctuation on visual observation.

In implementation, tabletop condition control focuses on constraining tabletop material, surface color and texture complexity so that the target object maintains stable identifiability during collection. On the one hand, by selecting tabletop materials with weak reflection and simple texture, interference from local highlight regions and surface noise on target boundaries is reduced. On the other hand, by constraining tabletop color and pattern complexity, the spatial reference during the placement stage becomes more stable, reducing positioning deviation caused by tabletop variation.

To quantitatively evaluate the brightness contrast relationship between the target and the tabletop region, Weber contrast is introduced as a measure of the effectiveness of tabletop condition control. Weber contrast is defined as the difference between the mean brightness of the target object region and the mean brightness of the tabletop region, divided by the mean brightness of the tabletop region. When the tabletop region brightness is approximately constant, this expression can be used to characterize the brightness contrast relationship of the target relative to the tabletop. A positive value indicates that the target is brighter than the tabletop, while a negative value indicates that the target is darker than the tabletop. A larger absolute value indicates a more obvious brightness difference between the target and the tabletop, a clearer target boundary, and tabletop conditions more favorable to visual observation. A smaller absolute value indicates that the brightness of the target and the tabletop are closer, making the target boundary harder to distinguish. This definition is a specific expression of the general Weber contrast definition in the tabletop scenario of this study and is consistent with its basic idea of describing the brightness difference between a target and a background under approximately constant background brightness conditions. The indicator provides a quantitative basis for comparing the brightness contrast relationship between target and tabletop under different tabletop conditions.

To address the influence of tabletop material, reflection and texture complexity on visual observation, a tabletop condition control algorithm is proposed. The algorithm takes as input the target image, the tabletop region image, an initial tabletop configuration, a Weber contrast threshold, a reflection constraint threshold and a texture complexity threshold, and outputs a tabletop configuration that satisfies the collection requirements. The algorithm initializes the tabletop configuration, collects the target image and tabletop region image under the current tabletop condition, computes the mean brightness of the target and of the tabletop region, computes the current tabletop contrast indicator from these mean brightness values, and evaluates the current tabletop reflection level and texture complexity. If the reflection level exceeds the reflection constraint threshold, the tabletop material is adjusted to reduce local reflection and the configuration is updated. If the texture complexity exceeds the texture complexity threshold, the tabletop texture complexity is reduced or a low-texture tabletop is substituted and the configuration is updated. If the absolute value of the Weber contrast falls below the contrast threshold, the tabletop color or brightness is adjusted to enhance the brightness contrast between the target and the tabletop region, and the configuration is updated. The target image and tabletop region image are then re-collected, and the mean brightness, contrast, reflection level and texture complexity are recomputed. The adjustment steps repeat until the reflection constraint, texture complexity constraint and contrast constraint are all satisfied, at which point the final tabletop configuration is output.

8. Experimental setup on a dual-arm collaborative platform

To investigate in depth how expert strategies and environmental conditions during demonstration data collection affect the training and inference performance of embodied intelligence models, a series of comparative experiments was carried out on a Realman dual-arm collaborative robotic platform. The experiments were designed to study the role of expert strategy data collection methods and of environmental conditions such as illumination, background and tabletop in the data collection process, comparing data collection methods with and without expert strategy and adjusting different environmental conditions to study their specific effects on model inference performance.

All experiments used the ACT (Action Chunking with Transformers) model for performance evaluation. ACT models the continuous manipulation process by predicting multi-step action chunks and is suitable for long-sequence control tasks such as dual-arm manipulation. Therefore, ACT was selected as the unified baseline model to compare the effects of datasets collected under different demonstration collection strategies and environmental conditions on model validation performance. The Realman platform is equipped with four cameras as visual input devices, including left and right wrist cameras, a top camera and a bottom camera.

To verify the effectiveness of the proposed method in complex scenarios, an industrial raw-material grasping and placing task was designed, with clutter interference introduced to simulate complexity in real environments. The data collection task was performed by the same experienced expert operator. The experimental accuracy requirement was set to 2.5 mm, and the industrial raw material was required to undergo positional offset and angular rotation within a specified range, while the vise maintained a fixed size, position and orientation to ensure experimental rigor and data accuracy. Experiments were conducted under two conditions. Under condition 1, the initial position and orientation of the industrial raw material were randomly placed within a 5 cm radius around the center. Under condition 2, the initial position and orientation were randomly placed within a 10 cm radius around the center. This design was intended to test the grasping and placing performance of the robot under different spatial constraints. Fifty online inference trials were performed under condition 1 and condition 2 respectively, and after averaging the results of condition 1 and condition 2, the grasp success rate, placement success rate and task execution accuracy were recorded. Task execution accuracy is expressed as the target placement position deviation, where a smaller value indicates higher task execution accuracy. By comparing task performance under different conditions, model performance and the stability of the method on the target task were evaluated.

For clutter settings, three red building blocks and two yellow building blocks were added as clutter. The red blocks measured 40 mm by 20 mm by 50 mm and the yellow blocks measured 40 mm by 20 mm by 40 mm. This setting was intended to increase task difficulty and to verify the influence of expert strategy and environmental conditions on model performance. The placement positions of the clutter were generated through a stratified random sampling mechanism, ensuring random distribution of items on the tabletop while maintaining a certain distance between the clutter and the raw material, so that the gripper would not contact the clutter during grasping. The clutter therefore mainly formed visual interference. The specific steps were, first, to randomly select the raw material position, and second, to use stratified random sampling to generate clutter placement positions while satisfying the minimum distance constraint from the raw material, ensuring both randomness and reasonableness of the positional distribution.

As for the experimental condition settings, conventional illumination conditions relied only on indoor lighting and natural light sources, simulating natural illumination conditions in common environments and ensuring that the experiment was conducted without additional artificial light sources, reflecting the illumination environment in daily operations. Conventional background conditions meant that there were no obstructions around the work area and that the workbench was rotated after collecting a certain amount of data, so as to ensure inconsistency of background conditions. Conventional tabletop conditions used a metal workbench surface, simulating the common workbench material and color found in actual environments. Controlled illumination conditions placed shadowless lamps and top LED array lamps above the work area to reduce interference from natural light fluctuation and local shadows on visual input, combined with reflective cloth to uniformly control the irradiation range and brightness distribution. The shadowless lamps were used to weaken local shadows and high-reflection interference in the target region, while the LED array lamps were used to improve the uniformity of overall illumination in the work area. Through these settings, the illumination distribution of the work area could be kept relatively stable and the value of the coefficient of variation could be reduced, thereby improving illumination uniformity and visual observation stability.

Controlled background conditions installed enclosures on three sides of the work area and covered them with background cloth of uniform color and material to constrain background color and texture features and reduce the visual complexity of the background region, which reduced observation deviation caused by background variation between collection batches and reduced the matching degree value, that is, reduced the color distribution similarity between the target region and the background region, thereby enhancing the distinguishability of the target object. Controlled tabletop conditions laid a covering layer of uniform material and color on the work surface to constrain the surface characteristics of the tabletop. This setting reduced local highlight regions and surface noise produced by the metal tabletop and improved the visual distinction between the target object and the tabletop region. According to the Weber contrast formulation, unifying the tabletop material and surface color can increase the absolute value of the Weber contrast, thereby improving the brightness contrast level of the target region relative to the tabletop region, making the target boundary clearer and enhancing the stability of visual observation.

Six datasets were used for embodied intelligence model training and inference validation, each containing 100 episodes. The first dataset consisted of data collected with the conventional data collection method under conventional illumination, conventional background and conventional tabletop conditions, serving as the baseline dataset under a conventional environment. The second dataset consisted of data collected with the dual ViA segmented sampling strategy under conventional illumination, conventional background and conventional tabletop conditions, used to analyze the influence of the collection strategy on model training and inference results. The third dataset consisted of data collected with the dual ViA segmented sampling strategy under controlled illumination, conventional background and conventional tabletop conditions, used to analyze the influence of illumination conditions on model training and inference results. The fourth dataset consisted of data collected with the dual ViA segmented sampling strategy under conventional illumination, controlled background and conventional tabletop conditions, used to analyze the influence of background conditions on model training and inference results. The fifth dataset consisted of data collected with the dual ViA segmented sampling strategy under conventional illumination, conventional background and controlled tabletop conditions, used to analyze the influence of tabletop conditions on model training and inference results. The sixth dataset consisted of data collected with the dual ViA segmented sampling strategy under controlled illumination, controlled background and controlled tabletop conditions, used to comprehensively analyze the influence of environmental condition control on model training and inference results.

The data collection workflow was designed to verify the effectiveness of the dual ViA segmented sampling strategy in dual-arm collaborative operation, particularly in improving the precision and stability of data collection. In the experimental design, the left and right robotic arms were each equipped with a ViA system. The ViA on the right arm was mainly responsible for visual tracking and localization of the target raw material workpiece, optimizing visual input quality by dynamically adjusting the viewpoint. The ViA on the left arm was responsible for grasping and placing operations and adjusted its actions in time according to changes in the target object’s posture, so as to ensure precise grasping and placement of the raw material at the designated position. The segmented sampling strategy divides the long-sequence task into a number of small intervals and uses a precise timestamp synchronization and alignment mechanism to ensure consistency of data in each stage along the time dimension, thereby improving data quality. If the time taken for data collection exceeds the predetermined step length, the data is regarded as invalid and is no longer used for subsequent model training.

In the dual ViA segmented sampling data collection process, from step 0 to step 80 the experiment begins and the start button is pressed to start data collection, while the wrist camera of the right robotic arm adjusts its viewpoint so that the image center is precisely aligned with the raw material workpiece, providing stable and accurate visual feedback and ensuring that the robot can accurately recognize and localize the target raw material. From step 80 to step 180 the left robotic arm performs the grasping task: the arm translates above the raw material workpiece, adjusts the gripper angle to ensure a suitable grasping posture, and descends vertically to about 1 cm above the workpiece, ensuring that the front and rear distances between the gripper and the raw material are consistent and avoiding jitter. From step 180 to step 240 the gripper closes completely, ensuring that the clamping depth meets the requirement, with an appropriate delay set to avoid unstable gripper motion. From step 240 to step 300 the wrist camera of the right robotic arm continues to follow the raw material workpiece while the left robotic arm rises vertically about 30 cm and moves horizontally at a constant speed to above the vise groove, adjusting the raw material orientation in preparation for precise placement. From step 300 to step 370, while the wrist camera of the right arm continues to follow the workpiece, the left arm slowly lowers the raw material accurately into the vise groove, ensuring that the accuracy reaches 2.5 mm. From step 370 to step 420, while the wrist camera continues to follow the workpiece, the left arm releases the gripper, ensuring that the raw material is precisely placed in the vise groove and avoiding contact between the raw material and the inner wall of the vise when the gripper opens. From step 420 to step 450, while the wrist camera continues to follow the workpiece, the left arm slowly rises about 15 cm in preparation for subsequent tasks. From step 450 to step 500, both arms return to the initial position at constant speed, completing one full data collection cycle.

In the conventional data collection workflow, from step 0 to step 500 the experiment begins and the start button is pressed to start data collection. The right robotic arm remains stationary while the left robotic arm performs the grasping operation. First, the left arm moves down to above the raw material and adjusts the gripper angle for precise grasping. Then the arm raises the raw material vertically and moves it smoothly to above the vise groove. The left arm opens the gripper to place the raw material precisely, ensuring that the raw material is correctly placed in the groove when the gripper opens. After placement is complete, the left arm returns to the initial position, completing one full data collection cycle.

9. Results: expert strategy versus conventional data collection

The first comparison experiment was designed to study the influence of expert strategy data collection methods on the training and inference performance of embodied intelligence models. Two data collection methods were compared: conventional data collection and dual ViA segmented sampling strategy data collection, with illumination, background and tabletop conditions kept conventional and only the data collection method varied. The evaluation results show that the model trained on data collected with the dual ViA segmented sampling strategy achieved an average grasp success rate of 34 percent, an average placement success rate of 18 percent and an average accuracy of 2.5 mm, and compared with the conventional data collection method, the grasp success rate increased by 22 percent, the placement success rate increased by 13 percent and task execution position deviation remained unchanged.

Method Condition 1 grasp success rate (%) Condition 1 placement success rate (%) Condition 1 accuracy (mm) Condition 2 grasp success rate (%) Condition 2 placement success rate (%) Condition 2 accuracy (mm)
Conventional method 22 10 2.5 2 0 2.5
Proposed method 46 28 2.5 22 8 2.5

To evaluate the reliability of the experimental results, a two-proportion Z-test was used to analyze the statistical significance of the differences in grasp and placement success rates. The results indicate that the performance improvement obtained with the dual ViA segmented sampling strategy is statistically significant. Specifically, for the grasp success rate the p-value was 0.0113 under condition 1 and 0.0021 under condition 2, while for the placement success rate the p-value was 0.0218 under condition 1 and 0.0412 under condition 2. All p-values were below 0.05, indicating that these performance improvements are statistically significant.

The results show that in the conventional data collection method, the right robotic arm remained stationary while the left arm independently completed the grasping and placing task. Because the collection process did not divide the long-sequence task into key stages and lacked stable multi-view observation, the observation information and action information in key operational processes were mixed, and the collected data had insufficient representational capability for key stages, which in turn made the model’s execution stability weaker during the task inference stage. In the dual ViA segmented sampling strategy, the grasping and placing process is organized into stages according to task semantics and operational goals, and through observation-side viewpoint control, execution-side action constraints and stage-driven scheduling, the observation, state and action information in key stages maintain a stable correspondence. This expands the effective visual coverage during task execution, strengthens the representational capability of key stage samples, improves the consistency and validity of demonstration data, and enables the model to show better task execution stability during the inference stage of embodied intelligence workflows.

10. Results: illumination condition control

The second experiment was designed to study the influence of illumination condition changes on the training and inference performance of embodied intelligence models. Two illumination conditions were set: conventional illumination and controlled illumination, with background and tabletop conditions kept conventional and the dual ViA segmented sampling data collection method used, so that only the illumination condition varied. Under the controlled illumination condition, the model trained on data collected with the dual ViA segmented sampling strategy achieved an average grasp success rate of 65 percent, an average placement success rate of 46 percent and an average accuracy of 1.75 mm. Compared with the model trained under conventional conditions, the grasp success rate increased by 31 percent, the placement success rate increased by 28 percent and task execution position deviation decreased by 30 percent.

Illumination condition Condition 1 grasp success rate (%) Condition 1 placement success rate (%) Condition 1 accuracy (mm) Condition 2 grasp success rate (%) Condition 2 placement success rate (%) Condition 2 accuracy (mm)
Conventional illumination 46 28 2.5 22 8 2.5
Controlled illumination 78 52 1.5 52 40 2.0

The same two-proportion Z-test was used to analyze the differences in grasp and placement success rates. The results indicate that the performance improvement under controlled illumination is statistically significant. Specifically, for the grasp success rate the p-value was 0.00098 under condition 1 and 0.0019 under condition 2, while for the placement success rate the p-value was 0.0143 under condition 1 and 0.00018 under condition 2. All p-values were below 0.05, indicating that these performance improvements are statistically significant.

The results show that under conventional environmental conditions the illumination in the work area is affected by natural light variation, and shadows caused by robotic arm occlusion also appear during the grasping and placing process, resulting in uneven illumination distribution and a relatively large value of the coefficient of variation, which reduces the stability of visual input. Under controlled illumination conditions, by uniformly configuring white reflective cloth, shadowless lamps and top LED array lamps, the illumination distribution of the work area becomes more uniform and the value of the coefficient of variation is clearly reduced, indicating that illumination fluctuation and local shadow interference are effectively suppressed. The imaging quality of the target region and the stability of temporal observation are accordingly improved, showing that illumination uniformity is an important factor affecting the execution stability of grasping and placing tasks in embodied intelligence systems.

11. Results: background condition control

The third experiment was designed to study the influence of background condition changes on the training and inference performance of embodied intelligence models. Two background conditions were set: conventional background and controlled background, with illumination and tabletop conditions kept conventional and the dual ViA segmented sampling data collection method used, so that only the background condition varied. Under the controlled background condition, the model trained on data collected with the dual ViA segmented sampling strategy achieved an average grasp success rate of 59 percent, an average placement success rate of 39 percent and an average accuracy of 1.75 mm. Compared with the model trained under conventional conditions, the grasp success rate increased by 25 percent, the placement success rate increased by 21 percent and task execution position deviation decreased by 30 percent.

Background condition Condition 1 grasp success rate (%) Condition 1 placement success rate (%) Condition 1 accuracy (mm) Condition 2 grasp success rate (%) Condition 2 placement success rate (%) Condition 2 accuracy (mm)
Conventional background 46 28 2.5 22 8 2.5
Controlled background 72 46 1.5 46 32 2.0

The same two-proportion Z-test was used to analyze the differences in grasp and placement success rates. The results indicate that under the controlled background condition, the performance improvement is statistically significant in some respects. Specifically, for the grasp success rate the p-value was 0.0082 under condition 1 and 0.0113 under condition 2. For the placement success rate the p-value was 0.0623 under condition 1 and 0.0027 under condition 2. The grasp success rate was statistically significant in both groups, with p-values below 0.05. The placement success rate was statistically significant under condition 2 with a p-value below 0.05, but was not statistically significant under condition 1 with a p-value above 0.05.

The results show that under conventional background conditions the background color and texture lack unified constraints, so the degree of similarity between the target region and the background region in color distribution is relatively high and the matching degree value is relatively large, meaning the target boundary is not prominent enough and model recognition of the target position is affected. Under controlled background conditions, by unifying background color and texture features, the color similarity between the target region and the background region is reduced and the matching degree value is clearly reduced, indicating that interference from the background region on target boundary recognition is effectively suppressed. The contour and position features of the target object become more prominent, showing that background consistency helps reduce visual interference and improve target recognition stability in embodied intelligence tasks.

12. Results: tabletop condition control

The fourth experiment was designed to study the influence of tabletop condition changes on the training and inference performance of embodied intelligence models. Two tabletop conditions were set: conventional tabletop and controlled tabletop, with illumination and background conditions kept conventional and the dual ViA segmented sampling data collection method used, so that only the tabletop condition varied. Under the controlled tabletop condition, the model trained on data collected with the dual ViA segmented sampling strategy achieved an average grasp success rate of 51 percent, an average placement success rate of 27 percent and an average accuracy of 2.25 mm. Compared with the model trained under conventional conditions, the grasp success rate increased by 17 percent, the placement success rate increased by 9 percent and task execution position deviation decreased by 10 percent.

Tabletop condition Condition 1 grasp success rate (%) Condition 1 placement success rate (%) Condition 1 accuracy (mm) Condition 2 grasp success rate (%) Condition 2 placement success rate (%) Condition 2 accuracy (mm)
Conventional tabletop 46 28 2.5 22 8 2.5
Controlled tabletop 60 34 2.0 42 20 2.5

The same two-proportion Z-test was used to analyze the differences in grasp and placement success rates. The results indicate that under the controlled tabletop condition, the performance improvement is statistically significant in some respects. Specifically, for the grasp success rate the p-value was 0.161 under condition 1 and 0.0321 under condition 2. For the placement success rate the p-value was 0.517 under condition 1 and 0.0838 under condition 2. The grasp success rate was statistically significant under condition 2 with a p-value below 0.05, but was not statistically significant under condition 1 with a p-value above 0.05. The placement success rate was not statistically significant in either group, with p-values above 0.05.

The results show that under conventional tabletop conditions the metal tabletop easily produces local reflection and boundary noise, and the brightness contrast between the target object and the tabletop region is insufficient, so the absolute value of the Weber contrast is relatively small and the clarity of the target boundary decreases. Under controlled tabletop conditions, by unifying the tabletop material and surface color, the brightness contrast between the target object and the tabletop region is enhanced and the absolute value of the Weber contrast is clearly increased, indicating that local reflection and boundary noise are effectively suppressed. The target boundary becomes clearer during the placement stage, showing that tabletop condition control improves positioning stability mainly by improving the visual contrast relationship between target and tabletop in embodied intelligence data collection.

13. Results: joint environmental condition control

The fifth experiment was designed to study the influence of combining the three environmental condition control methods, namely illumination condition control, background condition control and tabletop condition control, on the training and inference performance of embodied intelligence models. Two environmental conditions were set: conventional conditions and controlled conditions. Under conventional conditions illumination, background and tabletop conditions were not controlled, while under controlled conditions illumination, background and tabletop conditions were all controlled. The dual ViA segmented sampling data collection method was used. By jointly applying the three environmental condition control methods, their influence on task execution success rates and task execution accuracy on the robotic platform was evaluated. Under joint control of the three environmental conditions, the model trained on data collected with the dual ViA segmented sampling strategy achieved an average grasp success rate of 86 percent, an average placement success rate of 79 percent and an average accuracy of 0.75 mm. Compared with the model trained under conventional conditions, the grasp success rate increased by 52 percent, the placement success rate increased by 61 percent and task execution position deviation decreased by 70 percent.

Condition Condition 1 grasp success rate (%) Condition 1 placement success rate (%) Condition 1 accuracy (mm) Condition 2 grasp success rate (%) Condition 2 placement success rate (%) Condition 2 accuracy (mm)
Conventional condition 46 28 2.5 22 8 2.5
Controlled condition 90 84 0.5 82 74 1.0

The same two-proportion Z-test was used to analyze the differences in grasp and placement success rates. The results indicate that the performance improvement obtained with the combination of the three environmental condition control methods is statistically significant. Specifically, for the grasp success rate the p-value was 2.40 multiplied by 10 to the power of minus 6 under condition 1 and 1.92 multiplied by 10 to the power of minus 9 under condition 2. For the placement success rate the p-value was 1.69 multiplied by 10 to the power of minus 8 under condition 1 and 1.95 multiplied by 10 to the power of minus 11 under condition 2. All p-values were below 0.05, indicating that these performance improvements are statistically significant.

The results show that under conventional conditions, where illumination, background and tabletop are all uncontrolled, illumination intensity is uneven and shadows are produced by the robotic arm during grasping and placing, so illumination fluctuation interferes with model execution. The visual difference between background and object is small, which affects the visual prominence of the object and thus reduces model recognition accuracy. The visual difference between tabletop and object is also small, so object edges are unclear, which affects recognition accuracy and positioning accuracy and ultimately causes unstable task execution. After jointly controlling illumination, background and tabletop conditions, the originally randomly varying environmental conditions are transformed into controllable, quantifiable and reproducible environmental variables, the illumination distribution of the work area becomes more stable, and the visual distinction between target and background as well as between target and tabletop improves simultaneously, so interference from the three environmental conditions on visual input is suppressed as a whole. Compared with single environmental condition control, the model performs better in both grasp and placement success rates, showing that joint environmental condition control can more effectively improve the consistency and stability of visual input during the demonstration data collection stage, thereby further improving inference stability in the raw-material grasping and placing task within an embodied intelligence framework.

14. Conclusions and future work for embodied intelligence

The reported study addresses the problem of demonstration data collection in grasping and placing tasks performed by dual-arm collaborative robots, and proposes a method for expert strategy data collection and environmental condition control. Through the dual ViA segmented sampling strategy, key stages in the task process are explicitly marked and observation-side viewpoint control and execution-side action constraints can be recorded cooperatively under a unified timestamp, thereby improving the consistency and validity of demonstration data. On this basis, illumination, background and tabletop conditions are set in a controllable manner and comparative datasets are constructed to analyze the influence of environmental conditions on model validation performance.

The experimental results show that compared with data collected using the dual ViA segmented sampling strategy under conventional environmental conditions, jointly adopting dual ViA segmented sampling and environmental condition control raised the model grasp success rate by 52 percent, raised the placement success rate by 61 percent and reduced task execution position deviation by 70 percent. These results indicate that expert strategy organization and environmental condition control can strengthen the representational capability of key stage samples, reduce interference from environmental variation on visual input, and improve the stability of model inference in complex scenarios, offering a practical contribution to the wider development of embodied intelligence.

Although the proposed method achieved favorable experimental results in the current industrial raw-material grasping and placing task, there remains room for further optimization. Future work will focus on adaptive modeling and compensation for environmental uncertainty factors such as illumination variation, background disturbance and tabletop material change, in order to improve system stability under unstructured collection conditions. At the same time, validation will be carried out in more complex long-sequence tasks such as assembly and insertion, to further evaluate the applicability of the method to broader embodied intelligence applications.

Scroll to Top