Rethinking Quality and Assurance in Adaptive Physical AI Systems
Abstract
“In control” used to mean something operational: the process was stable enough to predict, the measurement system was trustworthy, and when the process moved, the organization could see it and respond. That meaning is harder to defend with adaptive physical AI. These systems do not only operate in a variable world; they can change their own behavior as they run. In that setting, calm metrics can be a weak form of comfort. A chart can look quiet while risk accumulates elsewhere—especially when outcomes are delayed, proxies stand in for ground truth, and the system’s actions shape the data it later learns from. This paper reframes the quality question. Instead of asking only whether outputs look stable, we should ask whether change is bounded, visible, and owned. The goal is not to teach tools or decision rules. It is to clarify why legacy quality language can mislead and what kinds of assurance claims remain defensible when the system itself is allowed to adapt.
1. “In Control” Was a Promise, Not a Slogan
In mature manufacturing—semiconductors are the cleanest example—SPC mattered because it tied quality language to disciplined practice. A control chart was not a decoration. It was a shared agreement about three things:
- what “normal” looks like for a known process window,
- how you detect a meaningful shift, and
- what you do when that shift appears.
That agreement held because it rested on hard-earned habits: baseline discipline, metrology discipline, and change discipline. When those habits were weak, “in control” became theater. When they were strong, “in control” meant the line could run, the product could ship, and people could sleep.
Adaptive physical AI puts pressure on that promise. Once a system can revise how it behaves while it is operating, the old distinction between routine variation and a true process shift becomes harder to maintain. You can still use the language of control, but you have to earn it in a different way.
2. What “In Control” Traditionally Assumed
In classic quality practice, “in control” depends on a simple expectation: the process will behave tomorrow much as it did yesterday—within known limits—unless something changes (Shewhart 1931; Montgomery 2013). That expectation usually carries three assumptions, whether stated explicitly or not.
2.1. The process has a stable identity over the monitoring interval
The recipe, configuration, materials, maintenance condition, and operating practices are not drifting invisibly. When something changes, it is treated as change.
2.2. The measurement system is stable enough to support decisions
Metrology drift, bias, sampling artifacts, and data-quality failures are not hand-waved away. If you cannot trust the measurement, you cannot trust the chart.
2.3. Changes are governed and traceable
When the process moves, people can answer “what changed?” with evidence rather than inference. This is as much an organizational capability as it is a technical one.
Adaptive systems push on every one of these assumptions.
3. Why Adaptation Changes the Meaning of Control
When software closes the loop on sensors and actuators, it becomes part of the machine. That is already a significant shift for quality organizations. Adaptation adds another: the machine can revise the rules it follows as it runs.
Three practical consequences follow.
3.1. The process can move without an operationally useful baseline
In many deployments, behavior can change without a clean, human-legible boundary that operations can rely on. A version number may exist, but the operational reality is what matters: can you tie behavior to a baseline you can reason about?
If you cannot, then “in control” starts to mean “nothing obviously bad happened yet,” which is not a quality claim.
3.2. The system’s actions shape the data it learns from
SPC works best when what you measure is not being quietly reshaped by the decisions you make from those measurements. Adaptive physical AI often violates that condition.
A closed-loop system changes where it goes, what it sees, what it records, and what it avoids. Over time, the dataset becomes partly a reflection of the system’s choices rather than a neutral sample of the environment. A dashboard can look steady while the system has learned to route around its own weaknesses, or while it has reduced exposure to hard cases that matter for safety and reliability.
3.3. Outcomes are often late, expensive, or incomplete
Many physical consequences arrive after the fact: wear-out, latent defects, near-misses, downstream failures, and customer impact. That forces organizations to rely on proxies.
Proxies are unavoidable. The risk is treating proxy stability as if it were safety or quality. In an adaptive system, the proxy can remain calm while the underlying risk is shifting.
4. The Measurement Problem: The Numbers Can Look Fine
Quality engineers learn early that “good data” is not a given. It is manufactured—through calibration, sampling, and a refusal to confuse convenience with truth.
Adaptive physical AI adds two familiar traps in a new form:
- Averages can hide the tail. A system can improve the mean while getting worse in rare but severe conditions. In physical systems, the tail is often where reputational and safety failures live.
- Stability can be manufactured. A system can hold an indicator steady by changing its own operating patterns, not by becoming more capable. That produces stable charts and unstable confidence.
So the question cannot be limited to “are the metrics stable?” The better question is, “what do these metrics actually certify about real-world behavior, and what do they leave unobserved?”
5. A More Defensible Reframe: Control as Bounded, Visible Change
For adaptive physical AI, “control” has to mean more than stable outputs. It has to mean that change is managed in a way that a quality organization can defend. Put plainly: what is allowed to change, what is not allowed to change, how do we see the change, and who owns the consequences?
This is not a toolset. It is a set of claims that an assurance program must be able to support with evidence.
5.1. Define what is allowed to change
Not all adaptation carries the same risk. A calibration update is not the same as a behavior change that alters physical interaction with the environment. If the organization does not separate categories of change, it will either overreact to benign shifts or, more commonly, normalize risky ones.
5.2. State what must not change
In any serious quality program, some commitments are non-negotiable. In manufacturing, these constraints are not philosophical; they are what keep yield, reliability, and safety from being traded away for short-term output.
Adaptive systems require the same discipline. If optimization is left unconstrained, it will eventually collide with obligations the organization cares about—often at the worst possible time.
5.3. Make behavior shifts legible to operations
You do not need to decode every internal state; you do need an operational way to tell when behavior has shifted, under what conditions, and why it matters. If the organization cannot detect and describe meaningful change, it cannot govern it.
5.4. Assign ownership for outcomes
Control is not only technical. It is organizational. Someone must own the definition of acceptable behavior, the monitoring of drift, the response to excursions, and the authority to halt or roll back changes when risk rises. Without that chain of responsibility, “in control” is not a credible statement.
6. Assurance Has to Survive Contact with Operations
A launch decision is not the end of assurance. In physical systems, it is the beginning of the period when the world starts teaching you what you missed.
An assurance program that will hold up in the field has to treat operations as part of the evidence base, not as a separate realm. That is consistent with long-standing practice in high-reliability environments and with modern AI risk frameworks that emphasize lifecycle governance (NIST 2023; ISO/IEC 23894:2023).
Three implications matter for a pre-roadshow framing:
- Evidence must be refreshed. Yesterday’s qualification does not automatically justify today’s behavior if the system adapts and the environment shifts.
- Proxy metrics require skepticism. Stable proxies are useful, but they do not, by themselves, certify safety or robustness.
- Change must be treated as a managed risk. Adaptation is not inherently unsafe, but it cannot be treated as background noise.
7. Organizational Consequences: Where Failures Usually Start
In practice, the most damaging failures are often interface failures—between engineering and operations, between vendor and operator, between what is measured and what matters.
7.1. Quality ownership must be cross-functional
When a system senses, decides, and acts in the physical world, quality cannot sit in one function. The field conditions, the maintenance practices, the data pipeline, and the model behavior are one system. Fragmented ownership produces fragmented control.
7.2. “Model performance” becomes operational performance
In a closed-loop physical system, the model is not a report. It is behavior. Treat it accordingly.
7.3. Vendor promises must be translated into auditable claims
If the system is purchased, accountability still lands with the operator. Claims about performance and adaptation have to be translated into statements that can be monitored and defended: what can change, what cannot, how changes are detected, and how exceptions are handled.
8. Conclusion: Use “In Control” Only When You Can Defend It
“In control” retained its meaning in manufacturing because it was backed by stable baselines, trustworthy measurement, and governed change. Adaptive physical AI weakens those supports unless an organization rebuilds them in a form that fits systems that can adapt. In this setting, a defensible use of “in control” should mean something concrete: change is bounded by explicit commitments, behavior shifts are visible in operations, evidence is maintained over the lifecycle, and ownership for outcomes is clear. That discipline does not slow progress; it prevents progress from turning into surprise.
It is reasonable—and often correct—to take comfort in stable charts. It is also prudent to remember what SPC was designed to support: timely recognition of meaningful change and disciplined response. As manufacturing systems become more automated and, in some cases, adaptive, the gap between observing a process and controlling it can widen without obvious warning. A chart can stay calm while the system quietly changes how it makes decisions, what it measures, and what conditions it avoids. For teams responsible for quality, uptime, and risk, it is worth examining whether “in control” still means what it is assumed to mean.
In February, a small, in-person roadshow discussion is planned in Phoenix, Arizona, to examine these control questions in more depth. This is not a webinar and not a general forum, but a focused working session for practitioners who are already sensing the limits of familiar approaches. If you would like more information as details become available, please reach out directly at [email protected]. Further information will be shared privately.
One more thing: if the chart is stable today, what exactly are you trusting—if the system can adjust itself tomorrow?
Sources:





