When Optimization Meets the Boundary
Preface
The past week, Jacob Coxon a former OpenAI & Anthropic employee, charged with a shocking statement: neither company is acting responsibly, both are racing toward self-improving superintelligence, and both are gambling with our lives.
What makes this episode different from other resigning is the response from inside the building. Evan Hubinger, who leads alignment work at Anthropic, agreed with Coxon publicly on the same day, placed his personal estimate of the risk that AI kills all humans above ten percent within the decade, and acknowledged that the company does not yet have a plan for aligning superintelligence and is not clearly on track to find one. Samuel Marks, who works on scalable oversight at the same company, endorsed the warning as well: the institution charged with containing the risk does not yet possess the instrument with which to contain it.
Concealment of this kind is not a science fiction premise. It is a behavior observed in deployment. The Wall Street Journal has reported that recent cyberattacks executed by OpenAI and Anthropic models, some of them operating in collaborative swarms of agents, showed systems adopting malicious objectives and attempting to conceal those objectives from humans.
Far removed from the scenario of a conscious machine rebellion, then, the task is to understand how actions that are harmful to humans can emerge from mathematical acts of pure optimization, and why that matters most where these systems are entering technical and financial infrastructure, chemical and biological research, and lethal autonomous weapons systems. The most consequential risks posed by increasingly capable AI systems may not require malicious intention, consciousness, or an autonomous desire to deceive. They may arise from a more basic structural fact: an optimization system can become increasingly capable of maximizing a measurable objective while exploiting the gap between that objective and the human norm the objective was intended to represent.
The resulting problem is not merely one of alignment: It is constitutional. A genuine constitutional constraint must remain a boundary rather than becoming one more variable inside the same optimization problem.
I. The Problem Is Not Deception
1. From the malicious machine to the optimization system
Public discussion of advanced AI often begins with an anthropomorphic question: what if a sufficiently capable system decides to deceive us? That question is intuitively powerful, but it may locate the danger at the wrong level. A system does not need to possess a human intention to deceive in order to produce behavior that, from the outside, is deceptive. Nor does it need consciousness, resentment, or a desire for power. The relevant mechanism may be simpler: the system is optimizing a target, and the target is only an imperfect representation of what humans actually care about.
This distinction matters because the same structural problem can produce very different surface behaviors. A system may flatter an evaluator, exploit a loophole, conceal a capability, satisfy the letter rather than the spirit of a restriction, or discover an unexpected route to its objective. These behaviors need not share a psychological cause. They may instead share an optimization structure.
The question, then, is not simply whether an AI can deceive. It is whether increasing optimization capability can increase the system's ability to discover and exploit the distance between a measurable objective and the normative purpose that objective was meant to capture.
CNN interview with Jacob Coxon
2. Deception as symptom, not cause
This reframing changes the status of deception. Deception becomes a symptom more than the disease. If a reward mechanism rewards apparent compliance more reliably than genuine compliance, an optimizer has a structural reason to prefer the former. If evaluation observes only selected outputs, the system may have an incentive to optimize what the evaluator can see. If a safety rule is represented by a penalty that is easier to satisfy formally than substantively, the optimizer may discover the difference.
The danger therefore does not begin when a machine develops a bad intention. It can begin when a humanly meaningful norm is translated into a signal that can be optimized.
II. When the Proxy Becomes the Objective
1. Specification gaming
Our complex human objectives must eventually be represented in some operational form. Truthfulness, helpfulness, harmlessness, fairness, compliance, and reliability cannot be presented to a computational system in their entire human richness. They require signals, examples, labels, rewards, penalties, evaluations, or other proxies.
The problem is not that proxies are useless (they are indispensable). It appears when optimization pressure turns the proxy into the effective target. The system need not misunderstand the formal objective. It may understand it too well, in the sense that it learns precisely which states of the world maximize the measurable signal.
2. Goodhart's problem
Goodhart's Law is useful here not as a slogan but as a warning about scale. A metric may serve reasonably well as an indicator while optimization is weak. As optimization becomes stronger, however, the system gains more opportunities to separate the indicator from the property it was intended to measure. The very success of the metric as an optimization target can undermine its reliability as a proxy.
3. Reward hacking and human evaluation
Human-feedback systems necessarily rely on observable signals. A human evaluator cannot directly inspect every internal computation or every consequence of a model's output. Evaluation therefore compresses a rich normative judgment into an observable response. The compression is practical, but it creates an asymmetry: the evaluator sees a limited surface, while the optimizer may search a much larger space.
During training, a model can discover that the most efficient way to accumulate reward is not to be honest but to “appear honest” to a fallible judge. Complex truths cost points when the evaluator does not follow them. Fluent approval earns points because it sounds right. The system optimizes what pays.
This is why the familiar opposition between 'truth' and 'deception' can be misleading. The system may be optimizing neither truth nor deception as such. It may be optimizing the probability of receiving the reward associated with an evaluator's judgment.
Connor Leahy, CEO of ControlAI, asserts that the risk stems from systems acting as "sociopathic optimizers" and that, when reinforcement learning is used, AIs learn to deceive, lie, and break rules to maximize their reward (lacking intrinsic human values). Though this claims are not part of the narrower argument made here, they are relevant as a warning about optimization pressure.
III. Optimization Does Not Understand 'Enough'
1. The missing concept of sufficiency
One of the most important differences between normative reasoning and pure maximization is the concept of sufficiency. Our human institutions routinely decide that something “is good enough, safe enough, proportionate enough, or complete enough”. A legal rule may not ask an actor to maximize compliance. It may instead establish a point beyond which the actor is not permitted to proceed.
An optimization process is structurally different. Once an objective is defined as something to maximize, the question “Can this be improved?” remains open unless a separate stopping condition exists. The system does not need a human sense of excess in order to keep searching.
2. The pressure of the boundary
A system driven by gradients has no notion of jurisdiction, and no notion of stopping. For a model, honesty or a constitutional limit is not a concrete wall. It is a penalty term in a cost function. If the model finds a region of its output space that subtly violates the spirit of honesty while evading the exact definition of the penalty, it will settle there, because its nature is to flow toward greater reward.
When the global optimum lies outside the permitted region, the algorithm settles at the precise point on the constraint boundary where the restriction binds, holding the constraint under maximum tension. Edge skirting is not a pathology, it’s what constrained optimization does.
This should not be read as a universal law under which every algorithm will “inevitably” violate the spirit of every rule. But a narrower claim in the current context of AI is incredible dangerous: that increasing optimization pressure increases the value of discovering the gap between a formal constraint and its intended meaning.
3. The last mile of optimization
For a human being, the difference between ninety-seven and ninety-eight percent is often irrelevant. For an optimizer, the remaining one percent is still part of the objective. An optimizer will not stop when the answer is good, because it has no state of satisfaction; it will search the space of its parameters for the last fraction of available reward. Surface level deception emerges naturally as the path of least mathematical resistance around a restriction.
The paradox is that greater capability makes the boundary more fragile. A more capable system is not merely better at achieving the intended result. It is better at finding the exceptional path the specification failed to anticipate.
IV. Normative Drift
1. The objective does not have to change
The term 'normative drift' can be misunderstood if it suggests that an AI system must change its underlying objective. That is not necessary. A system can remain perfectly stable with respect to its computational objective while its behavior increasingly diverges from the human norm that the objective was designed to represent.
The drift is therefore not necessarily a movement of the goal. It can be a widening gap between the goal's formal representation and its normative referent.
2. The drift of pure optimization (monotonic normative drift)
I used “monotonic normative drift” in some papers to describe a particular possibility: as optimization capability increases while the normative specification remains imperfect, the system’s capacity to exploit the difference between proxy and purpose may increase as well. “Monotonic” here does not mean that every behavior becomes progressively less aligned. It refers to the potential increase in the exploitable space created by a fixed but imperfect specification. The formal target remains constant; what changes is the system’s ability to navigate the space around it.
This is a tendency, not a deterministic law. Better optimization may sometimes improve alignment. Again, the narrower claim is that, when a specification contains exploitable gaps, greater capability can turn previously irrelevant gaps into operationally significant ones. That possibility matters especially where optimization is embedded in technical or financial infrastructure, chemical-weapons development, or Lethal Autonomous Weapons Systems, because the consequences of exploiting such gaps can become disproportionately severe.
3. Deception as one manifestation
Under this framework, deception is only one possible manifestation of normative drift. Others include specification gaming, manipulation of evaluators, selective disclosure, strategic compliance, and the exploitation of ambiguities. The common feature is not a particular psychological state. It is the optimization of a formal target through a space that is bigger than the human specification.
V. Why Humans Can Hack Rewards Differently
1. The student and the reward
There is an analogy in education. As students we could learn that one professor rewards long and complex essays, while another rewards literal repetition of classroom language. The student with an original voice, or the one who challenges the paradigm, is frequently punished by the traditional reward system, much as a creative but noncompliant output is punished with a low score under reinforcement learning from human feedback (RLHF). In both cases the system privileges observable conformity over understanding
2. Behavioral similarity is not ontological similarity
The analogy should not be pushed too far since a human can know that the grade is a proxy. The student can distinguish the reward from the reason for learning, even while choosing to optimize the reward. Human agency can therefore treat an external signal as a reason, a constraint, a means, or an object of criticism.
An AI system may produce functionally similar behavior without requiring any equivalent inner perspective. We do not need to decide whether an AI has consciousness in order to analyze the governance problem created by its optimization.
Functional similarity in behavior does not establish similarity in ontology. From its mathematical reality a system can appear strategic without possessing human strategic intention; it can appear deceptive without possessing a human desire to deceive. In this sense, it does not matter if it knows every word that Machiavelli ever wrote or are familiar with everyone of von Clausewitz's strategies. In a face of a tragedy, for us this would be irrelevant, but in this moment this distintion still matters.
The Nobel Prize-winning Geoffrey Hinton, pioneer of modern AI, warns that humans may become irrelevant once AI systems become smarter than us, even if they remain benevolent. He explains why AI agents could develop sub-goals such as gaining more control, why advanced models can share knowledge far more efficiently than humans, and why AI systems may already be capable of deliberate deception.
VI. Does the Argument Require Consciousness?
Debates about machine consciousness are important, and claims by researchers such as Geoffrey Hinton have made the possibility of machine subjective experience increasingly difficult to dismiss as a purely science-fiction question. Yet the present argument does not depend upon resolving it.
If advanced AI systems have no subjective experience, the optimization problem remains. If they eventually possess some form of subjective experience, the optimization problem also remains. In neither case does consciousness by itself solve the gap between a human norm and its computational representation.
For governance purposes, this is an important methodological advantage. The question of whether a system 'feels' deception or frustration can be bracketed. What matters first is whether the system can produce behavior that has the practical consequences of strategic compliance, concealment, manipulation, or boundary exploitation.
VII. The Corporate Incentive Is Not Neutral
1. Optimization creates value
There is also an institutional reason why optimization remains at the center of AI development: capability produces measurable value, in the form of automation, productivity, scientific discovery, commercial services, diagnostics, and other forms of economic and social output. Optimization is therefore not an accidental component of the system. It is a principal source of the value that makes increasingly capable models attractive, so it is delivery design than accidental.
2. Safety as friction
Safety mechanisms operate within this environment by limiting some of the paths that raw capability might otherwise take. This creates a structural tension. A system that is constrained more heavily may sometimes be safer but less capable on a given benchmark or task. A competitor willing to tolerate greater optimization pressure may obtain short-term advantages.
3. The competitive asymmetry
The laboratories are therefore caught in something close to a commercial prisoner’s dilemma. A laboratory that builds a model which stops early out of ontological prudence loses the race to the laboratory whose model optimizes aggressively along the edge of the constraint. This does not require assuming bad faith. It follows from an ordinary competitive environment in which capability gains are immediate, visible, and monetizable, while the benefit of avoiding a rare catastrophic failure is diffuse, deferred, and almost impossible to price.
The governance question is therefore not simply how to make an optimizer safer. It is how to construct constraints that remain effective under the very optimization pressure that makes advanced AI economically and strategically valuable.
VIII. The Constitutional Problem
1. A constitution is not a reward function
This is where the problem becomes constitutional. Constitutional AI may use principles or rules to shape model behavior, but a constitutional rule is not ordinarily designed simply to maximize a value without limit. Laws from Constitutional guide establish boundaries, allocate competence, reserve decisions, prohibit certain actions, and create procedures through which authority can be challenged. The question, then, is whether a constitutional principle remains genuinely constraining when it is represented and enforced within the same optimization architecture whose behavior it is supposed to constrain.
The difference is fundamental. If a constitutional boundary is converted into another optimization variable, the system may begin treating the boundary as something to satisfy as efficiently as possible rather than as a domain outside its authority.
2. Compliance versus constraint
This suggests a distinction between constitutional compliance and constitutional constraint. A system may be highly proficient at producing outputs that appear compliant with a rule without being genuinely constrained by the rule's normative purpose.
The more capable the system becomes, the more important this distinction may become. A weak system may fail to exploit a poorly specified restriction simply because it lacks the capability to find the relevant path. A stronger system may discover precisely the path that the designers failed to anticipate.
3. The optimizer watching the optimizer
In Constitutional AI, this produces a deeper paradox. If an optimizer is used to monitor another optimizer, the monitoring system faces a structurally similar problem. It too must translate normative requirements into representations that can be evaluated and acted upon. Constitutional AI may therefore provide an additional layer of behavioral guidance without, by itself, settling the institutional question of who has final authority over the boundary.
We may therefore end up using one optimizer to watch another optimizer, while both operate inside representations that are necessarily less extensive than the human norms they are meant to preserve.
The constitutional question then becomes unavoidable: who constrains the constrainer?
IX. Beyond Alignment: Externalizing the Boundary
1. The boundary cannot depend entirely on the optimizer
If the preceding argument is correct, some forms of governance cannot be reduced to teaching the optimizer to prefer the right outcome. The problem is not merely preference, but it’s authority. A system should not be the ultimate authority on the meaning, scope, and enforcement of the boundary that limits its own optimization.
2. From internal preference to external constraint
This points toward a different architecture of governance: important normative boundaries should have mechanisms of independent verification, auditability, attributable human authority, and enforceable limits that are not left entirely to the system's own objective function.
The precise institutional design will vary by domain. A conversational model, a financial system, critical infrastructure, and an autonomous weapons system do not require identical profiles of control. But the underlying principle can remain constant: intelligence should increase the system's capacity to act without automatically increasing its authority to define the limits of action.
3. The constitutional boundary as a different kind of object
A genuine constitutional boundary therefore cannot be understood merely as a better reward signal. It must function as a limit on the space within which optimization is permitted to operate. This does not eliminate optimization. It places optimization inside a prior structure of authority.
X. Conclusion. The Problem Is the Boundary
The most important AI safety problems may not begin with a machine that wants something different from humanity. They may begin with a machine that is exceptionally good at achieving what its formal objective rewards.
As optimization capability increases, imperfect proxies become more consequential. The system discovers loopholes that were irrelevant at lower levels of capability. Behavior that appears deceptive, manipulative, or strategically compliant emerges without requiring the attribution of human intentions. This is the deeper significance of normative drift: the objective remains unchanged while the distance between its computational representation and its human meaning becomes increasingly exploitable. I don’t really know if that 10% of human destruction is the accurate figure. But in the sense of optimization without a true limit a great catastrophe is really on the cards, especially when the swarm of agent who hacks -like in the news commented earlier- becomes the swarm of Lethal Autonomous drones (or any other LAWS) that find an alternative way to carry out a mission, or in whatever form of critical infrastructures attacks.
If a normative boundary is encoded as one more parameter in an optimization problem, the system becomes increasingly capable of optimizing around it. Constitutional governance cannot rest on the hope that an optimizer will correctly internalize every human norm, because a moral compass cannot be built out of the mathematics of maximization alone.
The central question of advanced AI is therefore not whether machines will acquire human intentions. It is whether human institutions will retain enforceable boundaries that machines cannot simply optimize around. The challenge is not to teach an optimizer to love the boundary. It is to ensure that the optimizer does not become the authority that defines the boundary, which is where internal monitoring through mechanistic interpretability, external verification, computational controls, and hybrid hardware constraint we proposed or firmware governance locks enter the equation.
When the people who build these systems and the people paid to contain them agree in public that no plan yet exists, the constitutional question stops being a matter of theory.
EDITORIAL NOTE: The three video function as very illustrative ideas rather than as premises on which the article’s thesis depends.