Why Multi-Agent AI Requires Constitutional Control by Design

Share
Why Multi-Agent AI Requires Constitutional Control by Design

From Fable 5 to Astra

Introduction

On August 1, 2026, OpenAI announced what it describes as its next major model family without issuing a conventional release. Instead of a press statement and a benchmark table, it published ten machine-checkable proofs of open problems in mathematics and theoretical computer science, each formalized in Lean 4 and pushed to a public repository. The model behind them is called Astra. Days earlier, Sam Altman had demonstrated the system to legislators and regulators in Washington.

Astra is not a larger language model, it is a multi-agent system: a root agent creates subagents, distributes portions of a problem, waits for partial results, and synthesizes a final answer, a design built to hold a single objective for hours or days rather than to answer a bounded request.

Traditionally the public and normative debate has been conducted in the comparative register of incremental capability: “is it smarter, faster, more accurate than the previous version?” Astra forces the discussion up one level, to a question of institutional significance: how do you govern a system whose final output does not come from a single model but from the interaction and emergent behavior of multiple autonomous agents?

1. The structural difference

From the standpoint of tort law and traceability, the difference is precise. A conventional large language model produces its answer along a chain of reasoning contained within a single architecture. A multi-agent system involves coordination, delegation, internal negotiation, and cross-verification among specialized agents. To the individual black box of each model one must add a second, emergent opacity: how the agents interact with one another, what one agent told another, which partial results the root agent kept and which it silently discarded.

 

Figure 1. The Second Opacity of Multi-Agent Systems

It is tempting to present this as a novel puzzle of attribution, but tort law solved the impossibility of identifying the individual actor with enterprise liability, respondeat superior, the French faute anonyme, product liability, or in Latin-American legal systems, the objective liability of owner and keeper for risky activities (attribute responsibility without requiring anyone to determine which component failed). Under this legal reading a multi-agent system is, doctrinally, an unusually complicated thing whose owner is perfectly identifiable.

The point is that those doctrines answer who pays, but they do not answer what happened. And the two questions have been quietly conflated in most of the current debate. Traceability is not an input to the attribution of liability. It is an autonomous interest (a matter of epistemic due process) belonging to the person affected by a decision, independently of who ultimately indemnifies them. That’s what matters to a worker with a denied benefit, a patient given a recommendation, an applicant refused credit, a party on the receiving end of an automated administrative act: each has a claim to know how the decision was reached that is not extinguished by the existence of a solvent defendant. Multi-agent architectures do not create a liability gap. They can increase (they will) the knowledge gap, and the law currently has no instrument that addresses it directly.

Figure 2. Liability and Traceability Answer Different Questions

 

2. The Fable 5 & the location of the opacity

On June 9, 2026, Anthropic released Claude Fable 5 and Claude Mythos 5. Three days later, the Commerce Department directed the company to suspend all access to both models by any foreign national, whether inside or outside the United States (including Anthropic's own foreign-national employees) under the Export Administration Regulations. Because the platform cannot verify nationality in real time, Anthropic could not comply selectively and disabled both models for every user worldwide. The controls were lifted on June 30 and Fable 5 was redeployed on July 1.

The government's letter did not specify the national-security concern it was invoking. Anthropic disclosed the technical substance itself, at the end of the episode: Amazon researchers had found a method of bypassing Fable 5's safeguards, prompting it to identify software vulnerabilities, and in one case the model produced code demonstrating how the vulnerability could be exploited. Anthropic disclosed the technical rationale. The normative standard that triggered the intervention remains non-public.

This was the first use of export-control authority as a kill switch over a commercially deployed frontier model. An instrument designed for the movement of dual-use goods across borders was applied to a service already in the hands of users, with immediate global effect, on the basis of a criterion that was never disclosed.

Fable 5 represented the opacity structure in its purest observed form to date. If a model of Astra's class is ever suspended in this way, the recipient will have even less to work with, because the object of the intervention will not be a model but a topology.

3. EO 14409 and NSPM-11

The June ‘26 instruments are usually described as a step toward regulatory transparency. But they fall short of doing so. Executive Order 14409 (signed June 2, 2026) directs agencies to design a voluntary framework for pre-release engagement, a classified benchmarking process to designate "covered frontier models," up to thirty days of pre-release federal access to such models before they are shared with trusted partners, and collaboration between developers and the public authorities on selecting those partners. It expressly declines to establish a mandatory licensing or preclearance regime, and it shortened an earlier ninety-day access window to thirty.

NSPM-11, issued three days later, plays in the same league. It directs the national security enterprise to accelerate AI adoption and to work with industry to make the most advanced frontier models broadly available to national security professionals. The two documents pull in opposite directions (harden it, or ship it) and practitioners noticed the tension immediately: the same model may be expected to pass through pre-release review while simultaneously being pushed into national-security hands as fast as possible.

What the June instruments delivered, then, is procedure without publicly articulated substantive standards. They specify who decides and within what timeframes. They do not specify the standard under which the decision is made. The questions that determine rights remain unanswered in public:

What quantitative or qualitative metrics convert a model into a "covered frontier model"?
What specific technical threshold activates a restriction or a national-security alert? What objective criteria justify excluding or limiting access for foreign developers — or for foreign nationals generally, as in June?
What formal review procedure, and what due-process guarantee, allows a company to challenge an arbitrary restriction administratively or judicially?

As far as all the data is known, these parameters remain classified. The framework itself was finalized around the August 1 deadline set by the order, with reviews to be conducted by the Commerce Department's Center for AI Standards and Innovation and by the National Security Agency, and the White House has been briefing the major laboratories on the completed version. Astra is expected to be the first model submitted under it.

There is a further problem, and it is the one that bears directly on multi-agent systems. The classified benchmark designates covered frontier models on the basis of advanced cyber capabilities. But not general capability, not autonomy, nor architecture.

A system whose distinguishing feature is that it can coordinate subagents on a single objective for days is not obviously captured by a criterion calibrated to vulnerability discovery. The regulatory trigger is looking at offensive capacity while the governance problem has migrated to control architecture. That miscalibration is not a gap in enforcement, it is a category error at the level of the definition.

4.: The Lean certificate (the strongest objection)

Any argument for process traceability now has to answer the way Astra was actually announced, because that announcement is a rival solution to the same problem.

Every one of the ten proofs ships with a machine-checkable Lean 4 certificate in a public repository. Lean's kernel is small and trusted; a certificate either compiles or it does not. Anyone with the compiler can verify each proof independently, without trusting OpenAI, without institutional access, and without a doctorate in mathematics. The verification is binary, reproducible, and costless to the verifier.

This is a serious answer, and it deserves to be stated at full strength before it is qualified. If the output of an opaque multi-agent process can be formally verified, why does the process need to be auditable at all? The black box has been rendered irrelevant by a checkable artifact. Traceability, on this view, is a demand made by people who have not noticed that verification has solved their problem from the other end.

Figure 3. Verification Ends Where Constitutional Control Begins

4.1 Limits

Domain. Output verification works where a formal oracle exists; mathematics has one. In the legal view, the domains in which liability is actually litigated do not: there is no Lean kernel for a medical recommendation, a credit denial, a hiring decision, an occupational-disability assessment, or a legal opinion. In those domains the correctness of the output is contested precisely because it is not formally decidable. What can be established is how the conclusion was produced — which is to say, traceability is not an inferior substitute for verification, it is the only available guarantee wherever verification is unavailable. And it is unavailable almost everywhere that matters legally.

Mathematical. A Lean certificate confirms that a proof follows from the formal statement it was given. It does not confirm that the formal statement faithfully captures the open problem as the mathematical community understood it. That correspondence judgment remains human, requires domain expertise, and is not mechanizable because it is a judgment about meaning, not about derivation.

That gap is not a defect to be engineered away; it is the location of the third. What I have elsewhere called “cyber-human mediation” appears here in a more precise institutional form: “cyber-human constitutional control” at AI-level, a third that is constitutive to the relation itself and is not an abstract postulate about human oversight; it is the name for the irreducible position that survives even total formal verification. The strongest available verification technology, applied in the most formalizable domain in existence, still terminates in a human judgment of correspondence[1].

Constitutional control in this sense does not mean state supervision. It means the existence of an institutional guarantor capable of preserving traceability, contestability and due process independently of the developer and of any singular exercise of executive power.

Verification

Correct derivation

──────────────────────────────

Boundary of formal methods

──────────────────────────────

Human judgment of correspondence

Cyber-Human Constitutional Control

5. Leiden

On June 2, 2026 (the same day as the executive order) a working group of sixteen mathematicians from fifteen universities, convened at Leiden University's Lorentz Center, released the Leiden Declaration on Artificial Intelligence and Mathematics, endorsed by the International Mathematical Union and relevant figures[2], identified unreliable results, the erosion of attribution, dependence on closed commercial systems, and the substitution of corporate publicity for community verification, with DeepMind's AlphaProof as the emblematic case (a blog-post announcement in July 2024 with peer-reviewed methods in Nature only in November 2025).

Read the Astra announcement against that document and something remarkable appears. Sixty days after that community articulated a standard of care, a commercial laboratory announced its most significant result in a form that complies with the core of it: machine-verifiable certificates, published openly, checkable by anyone, released simultaneously with the claim rather than fifteen months after it. No statute compelled this. No agency reviewed it. This is the spontaneous market answer to Leiden. A professional order stated a norm, and the market moved to it faster than the classified benchmark was finalized.

This is a polycentric third functioning in real time: visible, plural, and, in its own domain, capable. It is also the strongest available evidence that they (polycentric subjects) do not require a territorial monopolist to resolve such a great question, at least where the community in question has genuine authority over the subject matter[3].

6. Constitutional Control by Design

The previous section identified a possible occupant of that position. This one specifies what such an occupant would have to require and the requirement has to survive an obvious objection first.

If a human control intervenes at every handoff between agents, it destroys the long-horizon autonomy that makes the architecture valuable in the first place. A system designed to run for days cannot pause for approval at each delegation. Any proposal that requires it is not a governance framework; it is a prohibition wearing one.

So that “third” must be architectural, not operational. It does not approve steps. It imposes attestation and recording at the points where the topology creates the opacity; which is to say at the delegation boundaries themselves:

  • which agent delegated what, to which subagent, under what instruction;
  • what each subagent returned, and against what verification, if any;
  • which partial results the root agent incorporated and which it discarded, and on what basis;
  • where in the chain a human judgment of correspondence was exercised, and by whom.

This is the “Minimum Viable Log” applied to a multi-agent topology rather than to a single inference chain. It costs the system nothing in autonomy; it runs alongside execution rather than gating it. What it produces is the artifact without which no subsequent review (administrative, judicial, or professional) has anything to examine.

That artifact is also what determines which of the candidate occupants can actually function. A professional order can measure conduct against a standard of care only if the conduct leaves a record. An attestation or underwriting market can price a capability only if the capability is legible. And the Judiciary can be activated only by a party able to plead facts. This is the practical answer to a Fable 5-type intervention: with a record, the recipient can build the case that reclassifies a singular opaque act for what it is, instead of being left facing a pole. The demand is modest in engineering terms and considerable in constitutional terms. It converts the traceability interest from an aspiration into a producible record, and it does so without installing a licensor.

The Multi-Agent Turn
Use the power of AI for quick summarization and note taking, Gemini Notebook is your powerful virtual research assistant rooted in information you can trust.

7. Conclusion

There remains a gray zone that the June instruments leave untouched, and it is the one that determines rights. Existing frameworks clarify which agencies act and what the formal stages are. They do not define the substantive criteria that condition the state's decision.

What I would call the constitutional layer of AI regulation — the body of rules governing citizens' elementary rights in relation to the use of AI — remains largely unwritten, and where it has been written it has been classified[4]. Developers, and more importantly the people affected by their systems, are left in a position of relative legal defenselessness before discretionary intervention.

If multi-agent architectures become the standard, the legal debate will no longer be about what a model can do. It will be about the institutional structure that guarantees that capacity remains controllable. The question moves from "What can AI do?" to "Who constitutionally controls the architecture that decides what AI can do?"

Rather than regulating each new capability as it arrives (in mathematics first, then programming, then cybersecurity, then scientific research) so regulation should concentrate on requiring a stable institutional third that preserves traceability, responsibility, and reversibility regardless of how powerful the system becomes. Not a licensor. A guarantor whose criteria are public, whose position is contestable, and whose function can be discharged by more than one occupant.

 Astra makes that requirement concrete rather than theoretical. When the final decision emerges from the interaction of several autonomous agents, and when the only formal verification available terminates in a human judgment of correspondence that no certificate can supply, the third is not an external imposition on the architecture, but inside of it. Cyber-human constitutional control cannot be "human in the loop." It is an institutional position that guarantees traceability, challenge, and accountability without intervening in every decision of the system.


Sources

  • OpenAI, Ten advances in mathematics (August 1, 2026) and the accompanying openai/ten-proofs repository.
  • Anthropic, Statement on the US government directive to suspend access to Fable 5 and Mythos 5 (June 12, 2026); Redeploying Claude Fable 5 (June 30, 2026).
  • Executive Order 14409, Promoting Advanced Artificial Intelligence Innovation and Security (June 2, 2026), 91 FR 34565.
  • NSPM-11 (June 5, 2026).
  • Congressional Research Service, IF13268, Controlling Advanced Artificial Intelligence: Executive Order 14409 Explained.
  • The Leiden Declaration on Artificial Intelligence and Mathematics (June 2, 2026), leidendeclaration.ai.
  • How the June 2026 AI Executive Order Hides Its Own Guarantor: "The Stealth Third", Digital Nomos (June 13, 2026).
  • The Leiden Declaration on AI & Mathematics: The End of Data as "Res Nullius"?, Digital Nomos (June 12, 2026).

 


[1] Even when everything verifiable has been verified, a judgment remains regarding the correspondence between the formalized problem and the real problem. That's not a political preference, it's an epistemological limit. That's where the mediating third naturally comes in.

[2] Figures including Peter Scholze, Terence Tao, Kevin Buzzard and Scott Aaronson.

[3] To be clear: Astra's compliance is partial, since the Declaration's concern about dependence on closed commercial systems is untouched; Astra itself is not public and its design has not been disclosed. And on attribution the two orders now openly collide; OpenAI has taken the position that authorship of a proof generated entirely by AI should be credited as such, on the ground that claiming human authorship would misrepresent both the system's contribution and the nature of human intellectual work, while the Declaration recommends that no authorship credit be granted to software agents. That is not a misunderstanding to be smoothed over. It is a genuine normative conflict between two thirds, and it will be resolved by whichever one journals, universities, and eventually courts elect to follow.

[4] A terminological note is necessary for readers coming from the technical literature. "Constitutional AI" is Anthropic's term of art for a training method in which a model is aligned against an explicit written constitution. I use constitutional here in its ordinary legal sense: the layer of fundamental rights and the institutional architecture that guarantees them and the argument concerns the governance of AI by public law, not any laboratory's alignment technique.

Read more