The rapid expansion of the humanoid robot industry has created an unprecedented demand for data. As these robots move from factory floors to homes, hospitals, and public spaces, they continuously collect, process, and learn from vast amounts of personal information. In this evolving landscape, the traditional purpose principle in personal data protection law faces serious challenges. My aim in this study is to systematically re-examine the purpose principle and propose a path that balances industrial development with the protection of personal data in the age of humanoid robots. The central question is not whether we should keep the purpose principle, but how we can reinterpret and operationalize it in a way that is both legally sound and technologically adaptive.
Humanoid robots are complex systems integrating artificial intelligence, advanced manufacturing, new materials, and multimodal sensing. They interact with humans through vision, voice, touch, and even emotional cues, which requires the collection of biometric data, behavioral patterns, preferences, health data, and location information. Unlike traditional data processing systems, humanoid robots often operate in dynamic, unstructured environments where the specific purpose of each data collection instance may not be fully known in advance. This inherent uncertainty collides with the purpose principle’s demand that purposes be specified, explicit, and legitimate prior to collection. My analysis unfolds in several steps: first, I trace the origin and internal structure of the purpose principle; second, I identify the practical dilemmas that humanoid robots pose; third, I argue for scenario-based standardization as a bridge between abstract legal principles and concrete technical operations; fourth, I critically evaluate the “compatibility” standard borrowed from EU law and propose a more localized alternative; and finally, I emphasize the need to keep research purposes open to foster innovation.

1. The Purpose Principle: Origin and Internal Architecture
The purpose principle, also known as purpose limitation, originated from privacy protection in European human rights law. Article 8 of the European Convention on Human Rights prohibits interference with privacy unless such interference is in accordance with law and necessary in pursuit of a legitimate aim. This idea was carried into data protection law by the Council of Europe’s Convention No. 108, which required that personal data be processed fairly and lawfully, and that data collection be limited to the minimum necessary for a predetermined purpose. Later, the OECD Guidelines in 1980 added the notion that data should be collected for a specified purpose, and subsequent uses that are compatible with the original purpose are permitted. This gave rise to a two-stage structure: a narrow entry at the collection stage and a broader exit for secondary uses.
In the General Data Protection Regulation (GDPR), purpose limitation and data minimization are listed as separate but intertwined principles. The collection phase must satisfy specific, explicit, and legitimate purposes; the processing phase must be adequate, relevant, and limited to what is necessary in relation to those purposes. The minimization principle has three components: relevance, sufficiency, and necessity. These components are balanced internally to achieve regulatory goals. Sufficiency allows more data to be collected when needed, relevance imposes a quantitative limit, and necessity examines the specific scenario to determine the types and amounts of data that are truly required.
The second layer of analysis concerns the use of collected data for new purposes. In the GDPR, new purposes require a separate legal basis or an assessment of whether the new purpose is compatible with the original purpose. This “compatibility” test has been interpreted differently across Member States. Belgium uses a reasonable expectations test, Germany and the Netherlands use a balancing test, and the United Kingdom and Greece apply transparency, legality, and fairness standards. This divergence demonstrates that the compatibility standard is not a clear rule but a flexible framework that relies on contextual judgment.
In Chinese law, the Personal Information Protection Law (PIPL) merges purpose limitation and minimization into one principle in Article 6. Article 6 requires that processing purposes be clear and reasonable, and that the processing be directly related to those purposes. It also demands that the means used achieve the purpose with the least infringement on individual rights and interests. For sensitive personal information, a specific purpose and sufficient necessity are required. If a processor intends to use already collected data for a different purpose, Article 22 requires a new consent from the data subject. Thus, China’s legal framework appears stricter than the GDPR: it uses a “directly related” standard rather than a “compatible” standard, and it does not carve out a general exemption for compatible secondary uses.
My understanding of the purpose principle is that it serves as a gatekeeper at the front end of data processing. It limits the volume and variety of data that can be lawfully collected, thereby reducing the overall risk of harm. It also provides a foundation for transparency and user trust. However, the principle is not self-executing. Its abstract language leaves many questions open: What does “directly related” mean? How do we measure “minimum necessary” in a machine learning system? How should we treat data that is collected for one purpose but later becomes useful for a related or even unrelated purpose? These questions become particularly acute in the context of humanoid robots, because their data acquisition and processing are continuous, multimodal, and often unpredictable.
2. The Humanoid Robot Era and the Value Conflict
Humanoid robots represent the latest stage in the evolution of AI interaction systems. The progression moves from cloud-based virtual assistants to on-device assistants and finally to embodied robots that mimic human form and movement. Cloud assistants rely on remote servers to process voice or text commands. On-device assistants use local sensors and processors to provide personalized services. Humanoid robots go further: they combine the capabilities of virtual assistants with physical actuators, enabling them to move, manipulate objects, and interact with humans in a natural and emotionally engaging manner. This embodiment necessitates the simultaneous operation of multiple sensing modalities: cameras for visual data, microphones for audio data, force sensors for touch, inertial sensors for balance, and thermal sensors for temperature. The resulting data streams are fused to enable navigation, object recognition, natural language understanding, and social behavior.
The development of humanoid robots depends on training sophisticated AI models. These models require massive datasets that capture the diversity of human environments, behaviors, and interactions. For example, a humanoid robot designed for elderly care must learn how to recognize falls, interpret emotional states, and understand speech in noisy environments. To achieve this, it needs not only the personal data of the care recipient but also environmental data, data about other individuals, and even simulated data. Moreover, the robot continuously improves its performance by collecting feedback and new data during operation. This creates a positive feedback loop: more data leads to better models, better models lead to more natural interactions, which in turn generate more data. The purpose principle, by restricting data collection to pre-defined purposes, disrupts this loop and may slow down technological advancement.
At the same time, I recognize that the purpose principle serves a crucial function in preventing data monopolies and protecting individual autonomy. In the absence of purpose limitation, companies can silently accumulate data beyond what is needed for any stated function. This can lead to data hegemony: super-platforms use their massive data advantages to reinforce market dominance, engage in price discrimination, and manipulate user behavior. The purpose principle acts as a counterweight to these tendencies. It forces processors to articulate why they need specific data points and to limit their collection accordingly. For humanoid robots, which are often deployed in intimate settings such as homes and hospitals, the potential for invasive data collection is enormous. A robot that monitors all audio and video in a household could reveal far more than the user intended to share. Therefore, the purpose principle remains normatively attractive, even though its implementation is far from straightforward.
The tension between data-driven innovation and the purpose principle is not new, but humanoid robots intensify it in three ways. First, the multi-purpose nature of humanoid robots makes it difficult to specify a single purpose at the time of data collection. A single robot may be asked to perform a complex task such as “assist with rehabilitation,” which involves a wide range of sub-tasks: tracking patient movement, adjusting resistance, reminding about medication, monitoring vital signs, and even distracting the patient from pain. Each sub-task may require different data types and quantities. Second, the overall purpose may evolve over time as the robot’s algorithms are updated. A purpose that is valid in version 1.0 may no longer be sufficient in version 2.0. Third, humanoid robots often generate data that are not directly related to any human user, but may still affect personal data, such as incidental images of neighbors or voices in the background. This complicates the application of the purpose principle, which assumes a clear relationship between a processor and a data subject.
To illustrate the different challenges, I summarize them in the following table:
| Dimension | Traditional Data Processing | Humanoid Robot Processing |
|---|---|---|
| Purpose specificity | Relatively fixed, often defined by a service contract | Dynamic, evolving with robot learning |
| Data types | Structured, mostly textual or numeric | Multimodal: images, audio, touch, motion, biometrics |
| Data source | Directly provided by user | Collected passively through sensors in real-time |
| Processing environment | Controlled, e.g., server rooms | Uncontrolled, in homes and public spaces |
| Secondary use risk | Moderate, limited by commercial purpose | High, because data are rich and continuous |
| Regulatory challenge | Requires auditing consent forms and data flows | Requires real-time assessment of purpose adherence |
3. Why the Purpose Principle Struggles in Practice
3.1 The internal ambiguity of “purpose”
I have observed that the term “purpose” is notoriously vague. A purpose can be expressed at different levels of abstraction. For example, “improve user experience” can be broken down into “reduce response time,” “increase recommendation accuracy,” or “understand user sentiment.” These more concrete purposes may themselves be divided into even more specific tasks. Without a clear hierarchy of purposes, it is impossible to determine what data are necessary. Moreover, the same purpose may justify different data collections depending on the context. Suppose that a humanoid robot’s purpose is “provide companionship to an elderly person.” To fulfill this purpose, the robot may need to collect a daily schedule, medication reminders, family photos, and emotional cues. Does it also need to collect location data? Possibly yes, for safety reasons. Does it need to analyze the tone of voice? Probably yes, to detect loneliness. The list can go on indefinitely. The principle of purpose limitation does not tell us where to stop.
Furthermore, the legitimacy of a purpose can change over time as social norms evolve. What was considered acceptable in 2020 may be deemed intrusive in 2030. But the purpose principle requires that purposes be fixed at the outset. This static nature clashes with the dynamic reality of humanoid robot development. I argue that the law needs to incorporate adaptive mechanisms that allow purposes to evolve within certain boundaries, while still protecting individuals from arbitrary or malicious uses.
3.2 The problem of data minimization in machine learning
Data minimization is the operational core of the purpose principle. In traditional data processing, it is relatively easy to identify the minimum data needed for a task. For example, to send a package, the courier needs the recipient’s name, address, and phone number. No other personal data are necessary. In machine learning, the situation is fundamentally different. A model’s performance depends on the quantity and diversity of training data. For a humanoid robot that must recognize novel objects, navigate unpredictable terrains, and understand regional accents, the optimal dataset cannot be known in advance. The feature space is too large, and the model’s nonlinearities make it impossible to analytically determine the minimal sufficient dataset. Therefore, the concept of “minimum necessary” becomes a moving target.
To make this concrete, let us model a simplified learning task. Suppose we have a dataset \(D = \{(\mathbf{x}_i, y_i)\}_{i=1}^{N}\), where \(\mathbf{x}_i\) is the feature vector and \(y_i\) is the target label. The goal is to train a model f that minimizes the expected loss:
$$\min_f \mathbb{E}_{(\mathbf{x}, y) \sim P} [\mathcal{L}(f(\mathbf{x}), y)]$$
In practice, N is not known in advance; the modal architecture may require more data to improve generalization. If we impose a constraint that |D| cannot exceed a predetermined bound B, we obtain:
$$|D| \le B \quad \text{subject to} \quad \mathbb{E}[\mathcal{L}] \le \epsilon$$
But the relationship between |D| and \(\mathcal{L}\) is usually nontrivial. For a deep neural network, the generalization error decreases with N in a way that depends on the underlying data distribution. No analytical formula exists. Thus, setting B without considering the model and the task is impossible. This explains why many developers resist strict data minimization rules: they are forced to collect as much data as possible to hedge against future uncertainties.
Another technical problem is that the minimization of one individual’s data may degrade the system’s performance for other individuals. If a user opts out of data collection, the model loses the ability to learn patterns that would have benefited that user. In a humanoid robot serving multiple members in a household, removing one person’s data affects the robot’s interactions with everyone else. Therefore, the unit of minimization is not the individual, but the whole system. This systemic nature of machine learning calls for a more nuanced approach than simply reducing the number of data points.
3.3 The missing rules for secondary use
I have found that Chinese law lacks clear rules for the secondary use of already collected personal data. Under Article 22 of the PIPL, if a processor intends to use personal data for a purpose other than the original one, the processor must obtain new consent. This consent requirement applies to any purpose change, regardless of whether the new purpose is closely related to the original. In contrast, the GDPR permits secondary uses without consent if they are compatible with the original purpose. China’s “directly related” standard is stricter and more rigid. The consequence is that any innovative use of data for a new purpose—such as using health monitoring data to develop a new diagnostic algorithm—would require fresh consent from every data subject. In practice, this makes it extremely difficult to conduct large-scale research or to improve products based on user feedback.
Some Chinese scholars argue that the “compatibility” standard should be imported to relax the grip of purpose limitation. Others suggest that “directly related” can be interpreted broadly to encompass some secondary uses. But I believe that this legal transplant may not work well in China’s regulatory environment. The compatibility standard is highly context-sensitive and unpredictable. It requires case-by-case weighing of several factors, such as the relationship between purposes, the position of the data subject, the data sensitivity, the potential consequences, and the safeguards in place. This is burdensome for courts and creates uncertainty for businesses. In contrast, a rule-based approach that enumerates permissible secondary uses in specific scenarios might provide greater clarity and stability.
3.4 Inconsistent judicial practice
I have examined Chinese court decisions involving the direct-relatedness standard and found considerable inconsistency. In one case, a court held that using personal data to push personalized course recommendations was not directly related to the original purpose of providing course services. In another case, a court held that a company could reasonably reuse user relationship information in its associated products because it matched the product positioning of a social networking app. These opposite outcomes reflect the difficulty of defining “directly related” in concrete disputes. The concept is inherently gradational; there is no bright line between direct and remote. Judges must rely on their intuitions about what is reasonable, which vary depending on the facts and their own attitudes toward privacy.
The European Court of Justice’s judgment in Digi (Case C-77/21) provides some guidance. The court ruled that creating a sub-database for testing and error correction was closely linked to the original purpose of fulfilling a service contract, and that such processing did not violate the purpose limitation principle. This is a relatively permissive interpretation, emphasizing the connection between the original purpose and the secondary use, as well as the reasonable expectations of the data subject. However, the court did not systematically apply all the factors set out in the Article 29 Working Party guidelines. This shows that even within the EU, the assessment of compatibility is ad hoc and evolving.
| Standard | Scope | Flexibility | Predictability | Cost |
|---|---|---|---|---|
| Directly related (China) | Narrow | Low | High (in theory) | Low for courts, high for industry |
| Compatible use (EU) | Broader | High | Low | High for courts, lower for industry |
| Enumerated exceptions (proposed) | Scenario-based | Medium | Medium | Medium |
4. Scenario-Based Standardization: Making Abstract Principles Concrete
4.1 Why quantification fails
I have seen attempts to implement the purpose principle through technical tools such as purpose-bound access control (PBAC) and expression languages like ODRL. These tools aim to encode purposes as machine-readable constraints. For example, an ODRL policy might state that a data resource can be accessed only for a purpose equal to “treatment” and not for “marketing.” This works well in static settings. However, humanoid robots are not static. Their data requirements change as they move through different contexts. An ODRL policy defined at the time of data collection cannot anticipate the dynamic needs of a robot that must respond to an emergency, such as detecting a fall and contacting a healthcare provider. In such cases, the immediate purpose is situational and cannot be pre-defined.
PBAC systems typically rely on a user-declared purpose for each access request. This is insufficient for multi-step, high-level purposes such as “conduct scientific research on gait analysis,” which involves multiple sequential activities: collecting motion data, preprocessing, training a model, validating, and publishing results. Mapping these activities onto a single purpose declaration is not feasible. Moreover, PBAC assumes that the purpose of data access is known at the time of access. In machine learning pipelines, data are often reused for purposes that emerge later, such as debugging a system after a failure. The technical tools cannot handle this open-endedness.
The deeper reason why the purpose principle cannot be fully quantified is that language is ambiguous and context is fluid. The word “improvement” means different things in different scenarios. If a robot is designed to improve customer satisfaction, the metric might be the number of positive reviews. If the goal is to improve purchase conversion, the metric is the click-through rate. Different metrics require different data, and hence different minimization assessments. In my view, this contextual dependence cannot be resolved by a global formula. Instead, we need to construct a family of standards that specify what purposes are legitimate and what data are necessary for each major application scenario of humanoid robots.
4.2 The role of standards in data protection
Standards are voluntary technical documents that provide specifications, guidelines, and characteristics for repeated use. In China, standards have played an important role in implementing data protection law. The Personal Information Security Specification (GB/T 35273) is a notable example. Although not legally binding, it is widely adopted by industry and referenced by courts and regulators. Its recommendations on minimizing data collection, obtaining consent, and securing data have effectively become de facto rules. The standard also introduces the concept of “basic business functions” and “extended business functions” in mobile apps, which helps to categorize data collection purposes.
For humanoid robots, I propose that we develop a set of scenario-specific standards that define:
- which purposes are permissible in each scenario;
- which categories of personal data are necessary for those purposes;
- how long data may be retained;
- which safeguards must be in place;
- which secondary uses are deemed acceptable without further consent.
These standards should be developed in collaboration with engineers, ethicists, lawyers, and industry representatives. They should be updated regularly to reflect technological progress. Because standards are more flexible than legislation, they can adapt more easily to changes in humanoid robot capabilities and social expectations. At the same time, standards can provide a safe harbor for companies that comply with them: if a processor follows the scenario standard, its data processing is presumed to satisfy the purpose principle.
4.3 A scenario taxonomy for humanoid robots
I propose to divide humanoid robot applications into several broad scenarios based on their data sensitivity and the degree of control over the environment. The scenarios are:
(1) Domestic companionship and care. This includes elderly care, child companionship, and assistance for people with disabilities. Data are highly sensitive, involving health conditions, emotional states, and private conversations. The primary purpose is to ensure safety and well-being. Secondary purposes such as improving the robot’s communication algorithms may be permitted only if data are effectively anonymized or if the user has an explicit opt-out.
(2) Medical and rehabilitation. This includes surgical assistance robots, rehabilitation exoskeletons, and telepresence health robots. The purpose is strictly therapeutic. Data collection must be limited to what is clinically necessary, such as joint angles, electromyography signals, gait parameters, and vital signs. Research using these data may be allowed with ethical review and patient consent.
(3) Public service and logistics. This includes concierge robots in malls, delivery robots, and guide robots in museums. Data are less sensitive but may include location, preferences, and interactions with the public. The purpose is to provide a specific service. The robot should not continuously record audio/video unless necessary for the service.
(4) Industrial and professional. This includes robots in factories, construction sites, and hazardous environments. Data may include workers’ locations and biometrics for safety purposes. The purpose is operational efficiency and accident prevention. Monitoring of an intrusive nature, such as continuous emotional analysis, is prohibited.
(5) Security and emergency response. This includes rescue robots, bomb disposal robots, and surveillance robots. The purpose is to protect life and property. The use of personal data is highly justified, but the collection should be limited in time and scope to the emergency at hand.
For each scenario, a separate standardization committee should draft a technical code of practice. The code should include a data dictionary that specifies the categories and granularity of data allowed. For example, in domestic care, instead of collecting raw audio, the robot might be required to process audio locally and only send extracted features such as voice activity to the cloud. This design principle, often called “edge processing,” can significantly reduce the amount of personal data leaving the home.
The following table summarizes the recommended data minimization levels across scenarios:
| Scenario | Data Sensitivity | Purpose Clarity | Permitted Data Categories | Secondary Use |
|---|---|---|---|---|
| Domestic care | High | Moderate | Health, presence, voice, images (for safety) | Only with consent or anonymization |
| Medical rehab | Very high | High | Clinical measurements, motion, vital signs | Research with ethics approval |
| Public service | Low | High | Location, preferences, interaction logs | Aggregated statistics only |
| Industrial | Medium | High | Location, safety metrics, work performance | Prohibited except for safety audits |
| Emergency | High | High | Situational awareness, biometrics, rescue images | Limited to the emergency operation |
In addition to the scenario standards, I recommend the creation of a technical committee that reviews the standards at least every two years, considering the rate of change in humanoid robot technology. If the review indicates that a standard is too restrictive or too lax, the standard should be amended. This dynamic process will help maintain the balance between supporting structures and adaptive changes.
4.4 The problem of “non-necessary but related” data
China’s national standard GB/T 41391-2022 distinguishes between necessary personal information, non-necessary but related personal information, and unrelated personal information. Necessary personal information is required for the basic function of an app to work. Non-necessary but related personal information is typically used for extended functions, and users may choose to provide or withhold consent. Unrelated personal information cannot be collected. I find this categorization useful but imperfect for humanoid robots. The boundary between “basic” and “extended” functions is blurred in an embodied robot. For example, a home robot’s basic function might be to clean the floor, but its extended function might be to provide security monitoring. Both functions require different data, and users may have different expectations. A better approach is to define “core purpose” and “permissible secondary purposes” for each robot deployment, rather than to rely on the changing notions of “basic” and “extended.”
5. Against the Transplant of the Compatibility Standard
Many legal scholars have called for China to adopt the EU’s “compatibility” standard to allow secondary uses of personal data without re-obtaining consent. I disagree. The compatibility standard is not a silver bullet; it is a black box. GDPR Article 6(4) lists several factors to consider when assessing whether a new purpose is compatible with the original one, but these factors are open-ended and require context-specific balancing. The Article 29 Working Party’s Opinion 03/2013 further elaborates these factors, but does not provide a clear algorithm. In practice, the application of the standard varies widely across EU member states. This unpredictability is problematic for businesses that need to know in advance whether their planned data processing is lawful. For humanoid robot companies, which operate in multiple jurisdictions and process large-scale real-time data, such uncertainty is expensive.
Moreover, the compatibility standard can be abused. A company could easily justify a broad array of secondary uses by arguing that they are “compatible” because they are intended to improve the user experience or to innovate. Without strong oversight, the purpose limitation principle could be eviscerated. In the EU, the GDPR has been criticized for its formalistic and bureaucratic approach to data protection. The compatibility standard adds another layer of paperwork and interpretation, without necessarily giving individuals more control.
I remember a specific case in Hungary, Digi, where the European Court of Justice allowed the creation of a sub-database for testing and error correction. The court reasoned that this secondary use was closely linked to the original purpose of providing subscription services. This outcome might seem reasonable, but it sets a dangerous precedent. If error correction is deemed compatible, then many other maintenance and development activities could also be considered compatible. For instance, using data to train a new model that provides the same service in a different language might be seen as compatible. Or using data to profile users for pricing decisions could also be framed as “performance improvement.” The boundaries become extremely fluid.
Instead of the compatibility standard, I advocate a rule-based approach with enumerated exceptions. The law (or a standard) should explicitly list the situations in which personal data can be used for a new purpose without fresh consent. These exceptions should be limited to situations that are clearly beneficial to the data subject and to society, such as:
- preventing or detecting cyberattacks;
- ensuring the safety and reliability of products (e.g., software updates, security patches);
- complying with legal obligations (e.g., tax reporting, court orders);
- conducting scientific research with independent ethics review;
- providing services explicitly requested by the data subject; and
- fulfilling tasks in the public interest, with legal authorization.
All other secondary uses would require new consent. This approach gives clear guidance to humanoid robot companies. They know exactly which secondary uses are forbidden and which are permitted. It also reduces the burden on courts, as they do not need to conduct a multi-factor compatibility analysis in every case. Of course, any enumeration will have gaps. New use cases may emerge that were not anticipated. To address this, I propose that the list be reviewed and expanded periodically through a transparent process involving regulators, industry, and civil society.
6. Keeping Research Purposes Open
Scientific research is one of the most important reasons to relax the purpose principle. Humanoid robot research requires continuous data collection for algorithm training, model validation, and system improvement. The knowledge gained from analyzing real-world interactions cannot be fully predicted in advance. Therefore, if the purpose principle is applied too strictly, it will impede scientific progress. The GDPR recognizes this by exempting research processing from certain compliance requirements, such as the need to specify a precise purpose, provided that appropriate safeguards exist. However, the definition of “scientific research” is not always clear.
In my opinion, the concept of scientific research should be interpreted broadly to include technological development and demonstration, basic research, applied research, and privately funded research. The UK Information Commissioner’s Office has taken this view, stating that research-related processing includes technology development and demonstration, as well as private sector research. This is a welcome interpretation because many breakthroughs in humanoid robotics come from industrial R&D labs, not just universities. If we restrict research exceptions to academic institutions, we would create an uneven playing field and force private companies to artificially separate their research and development departments from their business operations, which is impractical.
However, a broad research exception also carries risks. A company could label any data processing as “research” and bypass the purpose principle. To prevent this, I propose that the research purpose be subject to two constraints. First, the research must have a genuine scientific objective, not a purely commercial one. This does not mean that the research cannot have commercial value; it must have an underlying question that contributes to general knowledge, such as “how to improve gait recognition accuracy.” Second, the research must be accompanied by safeguards, including data minimization, pseudonymization, and internal access controls. If a company uses personal data to train a model that will be deployed for commercial profiling, that activity should not be considered research. The line is sometimes difficult to draw, but it can be made operational through ethical review and independent oversight.
China has already embraced a differentiated approach to development and use in the field of generative AI. The Interim Measures for the Management of Generative AI Services focus on regulating the provision of generative AI services to the public, while not restricting the internal development and testing phases. This “loose development, strict application” philosophy is appropriate for humanoid robots as well. In the early stages of a humanoid robot’s lifecycle, when it is still being trained and validated in a controlled environment, the need for data is high and the risks to individuals are lower because the robot is not yet in public use. Once the robot is deployed for commercial services, the purpose limitation principle should apply with full force. Therefore, I recommend that the law explicitly state that the development and testing of humanoid robot technology for the purpose of scientific research shall be exempt from the strict purpose specification requirements, provided that adequate safeguards are in place.
Let me express this as a formal condition. Let \(R\) be the set of processing activities characterized as research. Let \(P\) be the set of principles that apply to normal processing. Then for any processing activity \(r \in R\), the exemption condition is:
$$r \in R \iff \exists Q \ (Q \text{ is a legitimate scientific question}) \land (r \text{ is necessary to answer } Q) \land (r \text{ includes safeguards })$$
This formulation forces us to articulate the scientific question, the necessity, and the safeguards. It is not a blank check; it is a structured exemption.
The question of international collaboration also demands an open interpretation of research. Humanoid robot development is a global effort. Scientists in different countries share datasets to improve model robustness across cultures and environments. If the purpose principle is interpreted in a way that prevents cross-border data sharing for research, humanoid robot development will be fragmented and slower. I urge regulators to recognize that the societal benefits of humanoid robots, such as reducing the burden of elderly care and improving disaster response, outweigh the risks of research data processing, as long as the processing is transparent and subject to oversight.
7. Towards an Adaptive Legal Framework
I return to the central theme: how to balance supportive structures and adaptive changes in a society undergoing rapid technological transformation. The purpose principle is a supportive structure that preserves individual dignity and prevents power asymmetries. But it must adapt to the reality of humanoid robots. My proposed path is not to abandon the purpose principle, but to operationalize it through scenario-specific standards and to broaden the research exemption while narrowing the secondary use exception to a set of clearly defined cases.
The law itself should be written in a way that does not impose onto the world a specific technology that may become obsolete. Instead, it should set out general principles and delegate the details to standards. This is already common in areas like food safety and environmental protection. In data protection, we need to ensure that the standards are legitimately adopted and revised. The drafting of standards must be transparent and Inclusive, with input from all stakeholders, including those who are not professional data handlers. The enforcement of standards should be coordinated with regulatory agencies. A company that complies with the applicable standard should be treated as having fulfilled its legal duty under the purpose principle. Such a safe harbor will encourage companies to develop and adopt the standards.
I propose the following multi-layer governance model for humanoid robot data processing:
| Layer | Instrument | Content | Updating Mechanism |
|---|---|---|---|
| 1. Constitutional | National legislation | General principles, rights, and obligations | Rarely updated |
| 2. Sectoral | Regulations and judicial interpretations | Specific rules for humanoid robots | Periodic (5–10 years) |
| 3. Technical | Standards, guidance, and best practices | Data categories, minimization metrics, security measures | Annual or biennial review |
| 4. Organizational | Internal compliance programs and codes of conduct | Training, audit, and accountability mechanisms | Continuous |
At the first layer, the legislator should affirm that the purpose principle applies to humanoid robots, but with a presumption of validity for processing that follows an approved standard. At the second layer, regulators should issue special rules for high-risk applications, such as humanoid robots in medical or educational settings. At the third layer, standard-setting bodies should produce detailed technical specifications. At the fourth layer, companies should implement the standards through their internal governance, including privacy impact assessments, data protection officers, and audit trails.
One of the most difficult tasks is to define “minimum necessary” in a way that is both measurable and flexible. I suggest a “cost-benefit” framing, where the benefit is the improvement in task performance and the cost is the privacy risk. Let \(U(D)\) be the utility of using dataset D for a given task, and \(C(D)\) be the cumulative privacy cost. The optimal dataset D* maximizes the net benefit:
$$D^* = \arg\max_{D \subseteq \mathcal{D}} [ U(D) – \lambda C(D) ]$$
where \(\lambda\) is a weight reflecting the social value of privacy. In practice, U and C are hard to specify, but the standard can provide approximate formulas or thresholds. For example, in a navigation task, the utility might be the success rate, while the cost could be the number of identifiable locations stored. The standard can set a cap on the allowable cost, e.g., “the robot shall not store images of individuals outside a radius of 10 meters.” This is more concrete than a general statement about minimization.
Another important element is the “safe harbor” for edge computing. If a humanoid robot processes data locally and transmits only aggregated or encrypted features to the cloud, the data protection risk is significantly reduced. I propose that standards treat local processing as an automatic mitigation measure, thereby allowing a larger volume of raw data to be captured initially, as long as the raw data remain on the device and are deleted after a short period. Formally, let \(D_{raw}\) be raw sensor data, and let \(E\) be local extraction function. The cloud receives \(E(D_{raw})\), not \(D_{raw}\). If \(E\) is designed to remove personal identifiers, then the data minimization principle is satisfied, even though \(D_{raw}\) was temporarily collected. This interpretation is essential for humanoid robots, because they need raw data to perform tasks like obstacle detection, which may inadvertently capture people in the background. The law should not prevent the robot from seeing, but it should prevent the robot from recording and uploading everything.
I also advocate for “privacy by design” as a mandatory requirement for humanoid robot manufacturers. This means that the data architecture must be built with minimization and purpose limitation in mind from the initial design stage. For example, a robot’s microphone array can be designed to filter out silence, or a camera can be equipped with on-chip facial blurring. These design choices reduce the risk of unintentional collection. The standard should list such technical measures and require their inclusion in the product certification.
8. The Future of the Purpose Principle in the Age of Humanoid Robots
I have argued that the purpose principle cannot be simply translated into code, nor should it be abandoned. It is a foundational norm that ensures accountability, transparency, and fairness in the processing of personal data. In the era of humanoid robots, the principle must be reinterpreted in a way that acknowledges the dynamic, multi-purpose, and context-dependent nature of data processing. The most promising way forward is to combine scenario-based standardization with a narrow set of enumerated exceptions for secondary uses, while keeping a broad and open-ended research purpose.
Let me outline a concrete legal framework for the purpose principle in humanoid robot processing:
Article 6 revised (proposal). Processing of personal data by humanoid robots shall be based on a clear and reasonable purpose. The purpose may be specified at the level of the robot’s deployment scenario, as defined in the relevant national standard. Processing shall be limited to the data necessary to achieve that purpose, taking into account the robot’s specific functions and the technical safeguards adopted. Secondary uses of personal data for a new purpose shall require new consent, except for the following categories: (a) security and reliability maintenance; (b) scientific research, subject to ethical review; (c) law enforcement and public safety; (d) public interest statistics; and (e) other uses expressly permitted by law or by an approved standard.
This revision is not a radical departure from the existing law, but a clarification that aligns the law with technological realities. It recognizes that the purpose of a humanoid robot is not a single fixed statement but a description of its role in a particular context. The standard will provide the granularity that the law cannot.
I also believe that the time has come to rethink the role of consent in secondary uses. In many cases, consent is not meaningful because the user cannot foresee all potential future uses. Instead, we should rely on a combination of ex-ante governance (standards and certification) and ex-post accountability (oversight and liability). For low-risk secondary uses, we can assume openness. For high-risk secondary uses, we require a legal basis beyond consent. This is the path taken by many mature data protection regimes.
The concept of “new quality productive forces” is central to China’s development strategy. Humanoid robots are a perfect embodiment of this concept. They integrate AI, advanced manufacturing, and new materials to create products and services that were previously unimaginable. To foster these new productive forces, the legal environment must be enabling, not constricting. A overly rigid purpose principle would starve humanoid robots of the data they need to improve and evolve. At the same time, a lax purpose principle would erode public trust, leading to a backlash that harms the industry in the long run. Therefore, the optimal path is one of adaptive equilibrium.
I recall the wisdom of Lon Fuller, who said that the central problem of social design is to maintain the balance between supporting structures and adaptive flows. The purpose principle is a supporting structure. The technological and social changes brought by humanoid robots are adaptive flows. We need laws that are strong enough to hold us together, yet flexible enough to let us move forward.
9. Conclusion
In conclusion, my analysis leads to the following recommendations. First, the purpose principle remains an essential element of personal data protection, but it must be adapted to the unique characteristics of humanoid robots. Second, the adaptation should be achieved through scenario-based standards that specify permissible purposes, necessary data types, retention limits, and safeguards for each major application area. Third, the “compatibility” standard from EU law is not suitable for China because it is too unpredictable and could weaken the protection of individuals. Instead, we should adopt a rule-based list of exceptions for secondary uses, which is more transparent and comprehensible. Fourth, the exception for scientific research must be interpreted broadly, to include both academic and industrial research, while upholding ethical safeguards. Fifth, the law should provide a safe harbor for processors that comply with the approved standards, thereby creating a virtuous cycle of innovation and protection.
I have used several formulas and tables to illustrate my points. The most important equation is the one that balances utility and privacy:
$$D^* = \arg\max_{D \subseteq \mathcal{D}} [ U(D) – \lambda C(D) ]$$
This equation reminds us that data minimization cannot be viewed in isolation. It involves a trade-off. The role of standards is to provide a socially accepted value of \(\lambda\) for each scenario, and to define practical measures for U and C. As humanoid robots become more integrated into our lives, these standards will need continuous refinement. The law should not treat the purpose principle as a static rule, but as a living principle that evolves with technology and society.
Humanoid robot technology is still in its early stages, but its potential is enormous. I am convinced that with the right legal framework, we can enjoy the benefits of humanoid robots without sacrificing our privacy and autonomy. The key is not to choose between innovation and protection, but to design a system that makes them mutually reinforcing. My proposal for scenario-based standardization, a limited set of enumerated exceptions, and an open research purpose provides a viable path toward this goal. It respects the core values of the purpose principle while accommodating the demands of a data-driven world. I hope this analysis will inspire further discussion and action among legislators, standard-setters, companies, and civil society.
Ultimately, the question is not whether humanoid robots will change our society—they will. The question is whether our laws will be ready to guide that change in a way that upholds human dignity. By re-examining the purpose principle, we take a significant step toward that readiness.
