# Proof Over Promise: The Mathematics of Agentic Control

> Source: /resources/whitepapers/Enclavia%20Proof%20Over%20Promise.pdf

## Proof Over Promise The Mathematics of Agentic Control

Converting drift, hallucination, and context decay from qualitative risks into measured quantities that a regulator, an authorizing official, or an auditor can independently check.

## Your Data · Your AI · Your Way

## Dev Roy — Founder & Chief Executive Officer, Enclavia.ai

Enterprise & AI platform architect. 20+ years across federal civilian, DoD, and Fortune 500 programs. Active Top Secret clearance.

### August 2026 · Public release

Fairfax, Virginia · dev@enclavia.ai · 703-984-9981 · enclavia.ai Contents

## Executive summary 3

Key figures in this paper 3

## The August 2026 inflection, and what it did not change 5

### Where verification is cheap, AI compounds 5

### The economics moved at the same time 5

## What actually fails in agentic systems 7

### Three failure modes, three signatures 7

## The formal stack: seven quantities Enclavia measures 8

### Confident fabrication: entropy against groundedness 8

### Off-task drift: distance in activation space 9

### Repetition and looping: semantic autocorrelation 9

### The composite index and graduated response 10

### Context management as constrained optimisation 10

### Silent degradation: sequential change detection 11

### Trust and routing11 Calibrated abstention: conformal prediction 12

## From equations to executable agents 14

### Placement in the Shepherd hierarchy 14

### What the supervision layer costs 14

### The constraint that forces sovereignty 14

## Mapping measurement to control evidence 16

## The economics of owning the stack 17

## Limits, and what is not yet proven18

## Sourcing methodology 19 Conclusion and the ask 20

The ask 20 Enclavia.ai · v1.0 · August 20262

## Executive summary

On 1 August 2026, OpenAI published ten new results in mathematics and theoretical computer science produced by an unreleased model it calls Astra. Every problem had been open for at least a decade. The headline result is the first explicit construction of a non-sofic group, a question Mikhail Gromov posed in 1999 and nobody answered for twenty-seven years. Others disprove Connes's rigidity conjecture, prove Ehrhart's volume conjecture, and settle three problems from the Erdős catalogue. The token cost was about $2,000.

The figure worth studying is not $2,000. It is zero: the reported “sorry” count in the Lean 4 certificate repository OpenAI published alongside the manuscript. In Lean, a sorry marks a step the author asked the checker to take on faith. Zero of them means every logical step in all ten proofs compiled under a verifier that has no opinion about who wrote the argument, how fluent it sounded, or how confident the model was when it produced it.

The lesson enterprises should take from Astra The most credible AI output published this year did not earn its credibility from the model. It earned it from a cheap, external, adversarial verification layer that anyone can run on a laptop without trusting the vendor.

Frontier capability arrived. Enterprise-grade verification did not arrive with it. Mathematics has Lean. Chip design has formal equivalence checking. Software has test suites. These are domains where verification is cheap relative to generation, and they are exactly the domains where AI is compounding fastest. Clinical narrative summarisation, intelligence fusion, contracting, adverse-event triage, and regulatory submission drafting have no Lean. Verification in those domains costs more than generation, which is why pilots stall at the governance gate rather than the capability gate.

Enclavia.ai builds the missing layer. This paper sets out how. The argument is that the failure modes everyone describes qualitatively — hallucination, drift, context collapse, silent degradation — are all measurable quantities with established mathematical treatments, most of them decades old and none of them exotic.

Shannon entropy over the token distribution detects a model committing confidently to a fabrication. Mahalanobis distance in activation space detects a worker that has locked onto the wrong objective. Sequential change detection, in the form used for statistical process control since Page's 1954 CUSUM paper, detects a model that passed accreditation in January and has quietly degraded by June. Conformal prediction converts a raw model score into a coverage guarantee with a distribution-free finite-sample bound, which is the difference between a vendor's confidence number and a claim you can defend in a submission.

Enclavia's Shepherd-AI system computes these quantities continuously, acts on them through a graduated intervention policy, and emits the results as signed evidence. The paper closes with the constraint that shapes the entire architecture: most of these signals require access to the token distribution and the residual stream, and closed frontier APIs do not expose either. Sovereignty in this design is not a compliance preference. It is the precondition for measurement.

Key figures in this paper

## Quantity Value Tier

Astra open problems solved / token cost

## problems, each open ≥ 10 years; ≈ $2,000 in tokens; Lean 4

certificates published with a reported sorry count of zero

## Independently

reported Enclavia.ai · v1.0 · August 20263

## Quantity Value Tier

Alphabet Q2 2026 free cash flow −$5.9B, the first negative quarter since the 2004 IPO; capex $44.9B, double Q2 2025; FY26 guidance raised to $195–205B

## Company

reported Amazon FY2026 capex guidance Raised to ≈ $220B from ≈ $200B, with the increase attributed by CEO Andy Jassy to higher memory chip costs

## Company

reported Agentic project cancellation forecast Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing cost, unclear value, and inadequate risk controls

## Analyst forecast

Workflow reliability at 98% per step A 20-step agent completes all steps cleanly 66.8% of the time; a 50- step agent, 36.4% 2 — Arithmetic, independence assumed Trust-score noise floor σ ≈ 0.085 at α = 0.15 and p = 0.95, setting the minimum separation between dispatch thresholds

## Derived

Owned-cell break-even vs. cloud ≈ 1.6 / 2.7 / 8.1 months for Opus-, Sonnet-, and Haiku-class run-rates 3 — Internal cost model Sourcing tiers are defined in Section 11. Every quantitative claim in this paper carries one. Enclavia.ai · v1.0 · August 20264

## The August 2026 inflection, and what it did not change

Astra is a genuine capability event. The results span group theory, high-dimensional geometry, coding theory, arithmetic circuit complexity, quantum complexity, lattice cryptography, and extremal combinatorics. OpenAI describes the family as built to coordinate multiple agents over long-horizon tasks running hours to days, which makes it an agentic result rather than a single-shot one. Thomas Bloom, who curates the Erdős problems catalogue and who publicly dismantled an earlier and overstated OpenAI mathematics claim in October 2025, called these results big news.

The critical response is equally instructive. Gary Marcus noted that the 249-page manuscript says nothing about how the model works, how the proofs were verified, what role humans played, or whether any proposed proofs contained errors. Both things are true at once, and the tension between them is the design pattern this paper is about. The process was opaque. The product was verifiable. Verification did not require the process to be transparent, because the Lean kernel checks the artefact, not the author.

### Where verification is cheap, AI compounds

The domains where AI capability is compounding fastest share a property: an automated checker can accept or reject an output without a human reading it. Lean compiles a proof or it does not. A test suite passes or it does not. Formal equivalence checking in silicon design either establishes that two representations agree or produces a counterexample. In every case the cost of verification is a rounding error against the cost of generation, and the checker is adversarial by construction.

Now consider the workloads Enclavia's customers actually run. A clinical study report section drafted against source data. A safety narrative for an adverse event. A fused intelligence summary correlating SIGINT against a mobile SAM order of battle. A contracting document assembled from local templates and policy. None of these have a kernel. Verification means a qualified human reading the output against the source, which costs more than generation and does not scale. That gap, not model quality, is why so many programmes stop at pilot.

The design objective Manufacture cheap, continuous, machine-checkable verification for domains that have no Lean. Not a proof of correctness — that is unavailable for natural-language reasoning — but a calibrated, auditable statement about when the system is operating inside its validated envelope and when it is not, computed at runtime and signed as it goes.

### The economics moved at the same time

The same week Astra was published, the quarterly results made the cost structure underneath frontier inference impossible to ignore. Alphabet reported $44.9B of capital expenditure for Q2 2026, exactly double the year-ago quarter, against $39.1B of operating cash flow. Free cash flow came in at roughly −$5.9B, the first negative quarter since the company went public in 2004, and full-year guidance was raised to $195–205B with the CFO signalling a further significant increase in 2027. Revenue was not the problem; revenue grew 24% to $119.8B and Google Cloud grew 82%.

Amazon told a similar story with a more specific cause. Trailing-twelve-month free cash flow turned negative for the first time since 2023, and Andy Jassy raised the capex forecast to about $220B from roughly $200B, Enclavia.ai · v1.0 · August 20265 attributing the increase directly to the higher cost of memory chips. Reuters, working from LSEG consensus, has Microsoft, Alphabet, Amazon, Meta, and Oracle on pace to spend more on capex than they generate in free cash flow by 2027.

For a buyer signing a three-year agentic programme, the implication is narrow and practical. Per-token pricing is a claim on a cost base that is currently under upward pressure from a physical memory shortage, funded in part by debt and equity issuance, and priced by vendors who have not yet had to pass through the full cost of the buildout. Building a regulated, multi-year workflow on that price card without a hedge is an unhedged bet on someone else's capital structure. It is also, separately, a bet that the vendor will keep exposing the interfaces your controls depend on — which, as Section 6.3 argues, they largely do not.

Enclavia.ai · v1.0 · August 20266

## What actually fails in agentic systems

Agentic systems do not fail the way chatbots fail. A chatbot produces one bad answer and a human sees it. An agent chains reasoning steps, tool calls, and retrievals, and a small error at step four propagates silently into the artefact delivered at step forty. The arithmetic is unforgiving.

R(n) = pn Probability that all n steps of a workflow are correct, given per-step reliability p and independent errors. At 98% per-step reliability, which is a respectable figure for a well-prompted frontier model on a bounded task, a twenty-step workflow completes cleanly 66.8% of the time. A fifty-step workflow completes 36.4% of the time. Pushing the model from 98% to 99% moves the twenty-step figure to 81.8%, which is still not a system anyone would validate for clinical or mission use. Reaching a 95% floor on a fifty-step workflow requires perstep reliability of 99.9%, and no amount of prompt engineering on a single model gets there.

Exhibit 1. Per-step accuracy is not workflow reliability. The independence assumption is a simplification — correlated errors make the real curve worse in some regimes and better in others — but the qualitative conclusion is robust: long agentic workflows cannot be made reliable by improving the model alone.

The gap between 99% and 99.9% is not closed by a bigger model. It is closed by a supervision layer that detects a failing step and repairs it before the failure propagates. That is the entire function of Shepherd-AI, and it is why Gartner's forecast that more than 40% of agentic projects will be cancelled by the end of 2027 attributes the cancellations to cost, unclear value, and inadequate risk controls rather than to model capability.

### Three failure modes, three signatures

Failure mode What it looks like Why it is invisible in a cloud API Confident fabrication The model commits to specific, plausible, unsupported content — a citation, a dose, a coordinate, a clause — with no hedging in the surface text.

The tell is in the output distribution, not the text. A confident fabrication reads exactly like a confident fact. Off-task drift The worker gradually optimises for a nearby objective: summarising when asked to extract, describing when asked to compare, planning a route around one corridor when asked for options.

Drift appears in activation space several tokens before it surfaces in text. By the time it is readable it has already shaped the plan. Repetition and looping The agent re-derives the same intermediate result, re-queries the same source, or cycles between two partial plans until the budget is exhausted.

Each individual step looks reasonable. Only the sequence reveals the stall, and the sequence is what an API call does not give you. Enclavia.ai · v1.0 · August 20267

## The formal stack: seven quantities Enclavia measures

What follows is the mathematical core of the platform. Each subsection states the quantity, the estimator used at runtime, how the threshold is set, and what the resulting number is good for. None of the mathematics is novel; the contribution is the selection, the calibration discipline, and the fact that it runs inside the customer's trust boundary where the necessary internal state is actually available.

### Confident fabrication: entropy against groundedness

The naive approach to hallucination detection is to ask the model how confident it is. This fails because a fabricating model is confident. The useful signal is the joint behaviour of two independent quantities: how concentrated the output distribution is, and how well the output is supported by retrieved evidence.

The first is Shannon entropy over the next-token distribution, normalised by vocabulary size so thresholds transfer across models: H(pt) = − Σv V∈ pt(v) log pt(v) Ĥt = H(pt) / log|V| Normalised predictive entropy at generation step t.

Token-level entropy alone is a weak signal because low entropy is also the signature of a correct, wellsupported answer. The refinement that matters is semantic entropy, which samples K completions, clusters them into meaning classes by bidirectional entailment, and computes entropy over the cluster distribution rather than the token distribution. Farquhar and colleagues published this approach in Nature in 2024 and demonstrated that it detects confabulation substantially better than raw likelihood, precisely because it measures disagreement about meaning rather than disagreement about wording.

SE(x) = − Σc C∈ p(c|x) log p(c|x) Semantic entropy over meaning clusters C obtained by bidirectional-entailment clustering of K sampled completions. The second quantity is groundedness. For each generated claim, Enclavia scores the maximum entailment probability against the retrieved evidence set actually supplied to the worker, using a local natural-languageinference model: G(claim) = maxe E∈ PNLI(e claim)⊨ Groundedness: the strongest entailment available from the evidence set E that the worker was given.

Neither number is decisive alone. Their combination is. Enclavia computes a fabrication risk index that is high exactly when the model is certain and unsupported, which is the dangerous quadrant: = (1 − ĤΦ t) · (1 − G) Fabrication risk index. Bounded on [0,1]; maximised by confident, ungrounded generation.

The four-quadrant policy Grounded (G high) Ungrounded (G low) Confident (Ĥ low) Release. The system is certain and the evidence supports it. This is the intended operating quadrant. Highest risk. Confident fabrication. Block, force citation, or abstain. Never release unattended in a regulated workflow.

Uncertain (Ĥ high) Re-prompt or narrow. Evidence exists but the Abstain. Neither the model nor the corpus Enclavia.ai · v1.0 · August 20268 Grounded (G high) Ungrounded (G low) model has not converged on it. Usually a retrieval-framing problem, not a knowledge problem.

supports an answer. Escalate to a human with the retrieval trace attached. The lower-left quadrant is the one that causes regulatory findings, and it is the one that a self-reported confidence score cannot detect by construction, because in that quadrant the model's self-report is high and wrong.

### Off-task drift: distance in activation space

Off-task drift is a geometry problem. At the start of generation the system captures a baseline representation of the task — the mean-pooled residual stream at a chosen layer over the instruction tokens. As generation proceeds it measures how far the current representation has moved from that baseline.

Cosine drift of the layer-ℓ hidden state against the task baseline captured at generation start. Cosine distance is cheap but scale-blind and treats every direction in representation space as equally important, which it is not. Where a task class has enough historical traffic to estimate a covariance, Enclavia uses the Mahalanobis distance against the nominal task manifold instead: dM(ht)2 = (ht − )μ T Σ−1 (ht − )μ Mahalanobis distance from the nominal task distribution, with μ and Σ estimated from validated in-envelope runs.

The practical benefit is threshold setting. Under an approximate multivariate normal model, the squared Mahalanobis distance in an effective d-dimensional subspace is distributed as chi-squared with d degrees of freedom, so the alarm threshold can be set at the 1 − α quantile of that distribution for a target per-step falsealarm rate rather than by hand-tuning. Enclavia reduces dimensionality first, because the covariance estimate is unstable when the number of validated runs is small relative to the hidden width; this constraint governs how much historical traffic a task class needs before Mahalanobis mode is enabled at all.

The reason to instrument the hidden state rather than the emitted text is timing. A worker that has locked onto the wrong objective shows the shift in activation space before it shows it in output tokens. Detecting at the text layer means detecting after the plan has already been bent.

### Repetition and looping: semantic autocorrelation

Loop detection by exact string match fails immediately, because a looping agent paraphrases. The signal is semantic self-similarity within a sliding window. Enclavia embeds recent generation segments locally, takes the maximum pairwise similarity in a window of size w, and trips when it stays above threshold for k consecutive evaluations: R(t) = maxi<j W(t)∈ cos(ei, ej) trip R(t) >⇔ τR for k consecutive windows Semantic repetition index over a sliding window W(t) of embedded segments.

The consecutive-window requirement matters. Legitimate reasoning restates intermediate results; a single high-similarity window is normal. A sustained one is a stall. Setting k from the observed distribution of Enclavia.ai · v1.0 · August 20269 legitimate restatement, rather than from intuition, is the difference between a monitor that catches loops and one that fires on every well-structured answer.

### The composite index and graduated response

The three signals are combined into a single index whose weights are calibrated per task class, because a code-generation task and a narrative-summarisation task have different nominal entropy profiles: D(t) = w1Ĥt + w2Dcos(t) + w3R(t), wΣ i = 1 Composite drift index. Weights are configuration, versioned and signed, not model parameters.

Exhibit 2. The intervention bands are policy, not model behaviour. Below τ1 the worker runs unattended. Between τ1 and τ2 the Meta Agent reprompts with a tightened instruction. Above τ2 it reroutes to a different worker and benches the original for steering. A hysteresis margin on re-entry prevents a worker oscillating across a threshold from thrashing the dispatcher.

Two implementation details carry most of the operational value. The first is that the response is graduated rather than binary: a first trip is a cheap re-prompt, not an escalation, which keeps the false-positive cost low enough that thresholds can be set conservatively. The second is that thresholds live in signed, versioned configuration. A change to τ1 is a change-controlled event with an approver and a diff, which is what makes the monitor auditable rather than merely present.

### Context management as constrained optimisation

Context is a budget with a price, not a container to fill. Every token admitted to the window displaces another and shifts the position of everything downstream, and position matters: Liu and colleagues showed in 2023 that models attend most reliably to material at the beginning and end of a long context and degrade in the middle. Enclavia treats window assembly as a knapsack problem.

maximise Σi vixi subject to Σi τixi ≤ B, xi {0,1}∈ Context assembly: select evidence chunks of token cost τ within budget B to maximise total value.ᵢ The value term is not raw retrieval similarity. It composes relevance, an exponential recency decay, and an authority weight that encodes provenance — an approved protocol amendment outranks a meeting note regardless of embedding proximity: vi = s(q, di) · exp(− tλ Δ i) · a(di), = ln 2 / hλ Chunk value with half-life h on recency and provenance weight a(·).

Selecting purely on value packs the window with near-duplicates. Enclavia applies maximal marginal relevance, the Carbonell and Goldstein formulation from 1998, to trade relevance against redundancy, then places the highest-value chunks at the window boundaries to exploit the position effect: MMR = argmaxd R\S∈ [ λms(q,d) − (1−λm) maxdj S∈ s(d, dj) ] Maximal marginal relevance. λ near 1 favours relevance; near 0 favours diversity.ₘ The governance consequence is that context assembly becomes reproducible. Given the same query, corpus state, and configuration, the same window is assembled, and the window itself is logged. For a 21 CFR Part 11 environment this is the difference between an auditable retrieval and an unexplainable one.

Enclavia.ai · v1.0 · August 202610

### Silent degradation: sequential change detection

Everything above operates within a single generation. A different failure operates across months: a model, a corpus, or a workload that was validated in January and is no longer the same distribution in June. Nothing in a single request looks wrong. The aggregate has moved.

This is a solved problem in statistical process control, and Enclavia uses the standard machinery. The cumulative sum statistic accumulates deviations from the validated mean and alarms when the accumulation exceeds a decision interval: St = max(0, St−1 + (xt − μ0 − k)), alarm S⇔ t > h CUSUM (Page, 1954). The reference value k sets the smallest shift worth detecting; h sets the false-alarm rate.

The Page-Hinkley test provides a one-sided variant suited to monotone degradation, and ADWIN — the adaptive-windowing detector of Bifet and Gavaldà, 2007 — maintains a variable-length window with a Hoeffding-bound test on sub-window means, which removes the need to fix a window size in advance for workloads whose volume is uneven.

Tuning is expressed in average run length. The decision interval h is chosen so that the expected number of dispatches between false alarms, ARL0, matches the operational tolerance — a monitor that cries wolf weekly gets switched off, and a monitor tuned so loosely that it never fires provides no assurance. Selecting k and h against a stated ARL0 target turns that trade-off into a documented engineering decision rather than a default.

Why this is the control a regulator actually wants ISO/IEC 42001 requires monitoring of AI system performance over the lifecycle. 21 CFR Part 11.10(a) requires validation sufficient to ensure consistent intended performance and the ability to discern invalid or altered records. Neither is satisfied by a point-in-time validation report. Both are satisfied by a running changedetector with a documented false-alarm rate and a signed alarm history.

### Trust and routing

Every worker carries a trust score updated after each task by exponential moving average, with a reward of 1 on verified success and 0 otherwise: wt+1 = rα t + (1 − ) wα t, = 0.15α Trust update. Half-life ln(0.5)/ln(1−α) ≈ 4.3 tasks; effective memory ≈ 1/α ≈ 6.7 tasks.

The parameter α is a memory-versus-stability trade-off with a quantitative consequence that is easy to get wrong. For a Bernoulli reward with success probability p, the steady-state variance of the estimator is αp(1−p)/(2−α). At α = 0.15 and p = 0.95 the standard deviation is approximately 0.085. Any two dispatch thresholds separated by less than roughly 0.17 will therefore oscillate on noise alone, and the system will thrash between routing decisions for reasons that have nothing to do with worker quality. Deriving that bound before setting thresholds is the kind of detail that separates a monitor that works in production from one that works in a demo.

Exhibit 3. Trust is a state variable with known dynamics, not a static label. The shaded band is the ±1σ steady-state noise floor; thresholds inside that band produce routing oscillation independent of actual worker behaviour.

Enclavia.ai · v1.0 · August 202611 Dispatch is then a constrained assignment problem. Given m tasks and n eligible workers, minimise total cost subject to each task receiving exactly one worker whose trust exceeds the task's floor: min Σij cijxij s.t. Σj xij = 1 i, x∀ ij = 0 where wj < τi Worker assignment. Solved exactly in O(n³) by the Hungarian algorithm (Kuhn, 1955).

The cost term is where sovereignty enters the mathematics rather than the marketing. It composes compute cost, a latency penalty, and an egress penalty that is effectively infinite for any route that would send regulated data outside the trust boundary: cij = πj · tokens(i) + λL · latency(i,j) + λE · egress(i,j) Route cost. The egress term makes a policy-violating route unreachable by the optimiser, not merely discouraged.

Routing of this kind is well established in the literature. The RouteLLM work from LMSYS demonstrated cost reductions in the region of 85% while retaining roughly 95% of the stronger model's benchmark performance, by sending the large majority of traffic to a cheaper model and escalating only where the router predicts a quality gap. Enclavia's target is the same shape — 85–95% of inference served locally, with frontier fallback reserved for cases where the local trust estimate is genuinely insufficient and the data classification permits it.

### Calibrated abstention: conformal prediction

The final quantity converts everything above into a statement with a guarantee. Split conformal prediction takes a held-out calibration set of n exchangeable examples, computes a nonconformity score for each, and takes the empirical quantile: qD = (n+1)(1− ) / n empirical quantile of { s⌈ α ⌉ 1 … sn } Split-conformal calibration.

P( ytest C(x∈ test) ) ≥ 1 − α Marginal coverage guarantee. Distribution-free and finite-sample, requiring only exchangeability. Two properties make this worth the engineering. It is distribution-free, so it assumes nothing about the model that produced the score, and it holds in finite samples rather than asymptotically. The practical caveats are equally concrete. Coverage is marginal, not conditional, so it holds on average and not necessarily on every subgroup, which is why Enclavia calibrates per task class and per site rather than globally. Realised coverage on a finite calibration set is itself a random quantity with a Beta distribution, so a calibration set of a few dozen examples produces a guarantee too loose to be useful; the working floor is on the order of a thousand examples per class. And exchangeability is exactly the assumption that distribution shift breaks, which is why the change-detectors of Section 4.6 sit alongside the conformal layer rather than being replaced by it. When CUSUM alarms, calibration is stale and the guarantee must be re-earned.

What this buys in a submission “The system abstains rather than answers on the cases where it cannot meet a 95% coverage target, with coverage established on a per-indication calibration set of n = 1,200 and re-established on any changedetector alarm.” That is a defensible sentence in a regulatory filing or an authorisation package. “The model reports high confidence” is not.

Enclavia.ai · v1.0 · August 202612 Enclavia.ai · v1.0 · August 202613

## From equations to executable agents

Mathematics that cannot run inside the latency budget is a paper, not a product. This section is about where each quantity is computed, what it costs, and why the placement is what it is.

### Placement in the Shepherd hierarchy

Tier Responsibility Quantities computed here Meta-Meta Agent Allocates models to hardware, owns the retrain and steering bucket, runs the reinforcement pipeline, handles session cleanup. Trust EMA updates; CUSUM / Page-Hinkley / ADWIN across sessions; conformal calibration set management and staleness flags.

Meta Agent Decomposes the request into tasks, selects a worker per task, verifies output, re-prompts or reroutes on failure. Assignment optimisation; groundedness scoring; semantic entropy at high-stakes gates; composite index policy evaluation.

Worker pool Executes the bounded task on allocated hardware under continuous observation. Per-token normalised entropy; hidden-state pooling and cosine drift; sliding-window semantic repetition. Context service Assembles the window for each dispatch and logs what it assembled.

Knapsack selection; recency decay; MMR redundancy control; position placement.

### What the supervision layer costs

The signals differ by orders of magnitude in cost, and the architecture reflects that.

- Normalised token entropy is effectively free. The distribution is already computed to sample the next

token; entropy is a reduction over a vector the runtime already holds.

- Hidden-state pooling and cosine drift are linear in hidden width per evaluated step, and are evaluated on a

stride rather than every token.

- Semantic repetition requires an embedding pass over recent segments, so it runs on a window cadence,

not per token, using a small local embedding model.

- Semantic entropy is the expensive one: it requires K sampled completions plus pairwise entailment. It is

reserved for high-stakes verification gates — a released clinical claim, a dispatched recommendation — and is never run per token.

- Change detection is negligible. CUSUM is a running scalar update per dispatch.

The result is a monitoring layer whose steady-state overhead is dominated by the stride-sampled hiddenstate work, with the expensive semantic checks concentrated at the small number of points where an artefact actually leaves the system. Enclavia measures this overhead per deployment and reports it as a Tier 2 internal measurement rather than publishing a single headline number, because it varies substantially with model size, stride, and gate density.

### The constraint that forces sovereignty

Here is the part of this architecture that is not a preference. Enclavia.ai · v1.0 · August 202614 You cannot build this layer on a closed API Normalised predictive entropy needs the full next-token distribution. Cosine and Mahalanobis drift need the residual stream. Neither is exposed by the major frontier APIs, and where a truncated top-k log-probability view exists it is insufficient for a stable entropy estimate over a large vocabulary and provides nothing at all about hidden state. A vendor can withdraw or reprice what it does expose at any time. Running open-weight models on owned hardware is therefore the precondition for measuring what this paper measures — sovereignty is what makes the instrumentation physically possible, and only incidentally what satisfies the data-residency clause.

This is the argument I have found lands hardest with chief architects, because it inverts the usual framing. Sovereignty is normally sold as a constraint you accept in exchange for control, at some cost in capability. In an instrumented agentic system it is the opposite: the closed API is the constraint, because it withholds the state your controls require. A programme that commits to closed-API inference has, without necessarily realising it, also committed to never being able to prove much about its own behaviour beyond what the vendor chooses to attest.

Enclavia.ai · v1.0 · August 202615

## Mapping measurement to control evidence

Each quantity produces an artefact. Each artefact satisfies a control. This is the table that converts the preceding mathematics into an authorisation package. Measured quantity Evidence artefact Control satisfied Normalised entropy, semantic entropy, groundedness Φ Per-claim fabrication index with the retrieval trace attached, signed at release

## CFR Part 11.10(a) validation and record

integrity; EU AI Act Art. 15 accuracy and

## robustness; NIST AI RMF MEASURE 2.5

Composite drift index and intervention log Signed record of every trip, the band entered, the action taken, and the outcome

## NIST SP 800-53 r5 SI-4 monitoring; ISO/IEC

## Cl. 9.1 monitoring and measurement

CUSUM / Page-Hinkley / ADWIN alarms Change-detection history with the stated ARL0 target and the tuning parameters used

## ISO/IEC 42001 Cl. 9.1 and Cl. 10 improvement; 21

CFR Part 11.10(a) continued validated state Conformal coverage and abstention rate Per-class calibration record, realised coverage, abstention rate, staleness flags

## EU AI Act Art. 15 accuracy metrics and Art. 14

## human oversight; NIST AI RMF MEASURE 2.7

Trust scores and assignment decisions Worker competence map with the routing rationale for each dispatch NIST SP 800-53 r5 CM-3 configuration change control; ISO/IEC 42001 Cl. 8 operational control Context assembly record The exact window assembled, its provenance weights, and the configuration version

## CFR Part 11.10(e) audit trail; NIST SP 800-53 r5

AU-2 / AU-12 audit generation Threshold configuration Signed, versioned, diffable policy file with approver identity

## NIST SP 800-53 r5 CM-3; GAMP 5 configuration

management Absence of egress Static binary analysis confirming no DNS, TLS, NTP, or telemetry code path NIST SP 800-53 r5 SC-7 boundary protection; HIPAA Security Rule transmission security The last row is the one reviewers verify fastest, and it makes the general point. A control that is proved by absence — there is no outbound network code path in the binary — is stronger than a control proved by policy, because a third party can confirm it independently with a static analysis tool and does not have to take Enclavia's word for anything. That is the Astra pattern applied to compliance: make the artefact checkable, and the opacity of the process stops mattering.

Enclavia.ai · v1.0 · August 202616

## The economics of owning the stack

Per-token pricing scales with success. An agentic programme that works generates more inference, which generates a larger bill, which means the operating cost curve bends upward exactly as the business case improves. Owning the inference asset inverts that relationship: the capital cost is fixed at procurement and the marginal cost of additional work is electricity.

Exhibit 4. Break-even for a single Shepherd cell against constant-volume cloud run-rates. The reference configuration is five workers and one orchestrator, approximately 260 GB of accelerator memory, met by three 96 GB Blackwell-class cards at $11,349 each, with $900 per month for power, cooling, and support.

Reference workload Break-even After break-even Opus-class run-rate ≈ 1.6 months Near-zero marginal cost per additional task; the cloud line keeps climbing Sonnet-class run-rate ≈ 2.7 months An owned asset on the balance sheet against a perpetual operating charge Haiku-class run-rate ≈ 8.1 months Flat thereafter except electricity and support Two effects favour ownership and are deliberately excluded from the chart because neither can be quantified honestly at this stage. The first is that cloud spend rises with agent success, so the comparison understates the gap for any programme that scales. The second is the cost pressure documented in Section 2.2: Amazon has already attributed a $20B increase in its capex forecast to memory prices, and that pressure ultimately reaches the price card.

There is a third asset that does not appear on any invoice. An owned, instrumented stack accumulates the worker competence map, the calibration sets, the threshold history, and the drift analytics. That corpus is specific to the customer's own workload, improves with use, and cannot be handed over by a vendor whose API returns only text.

Enclavia.ai · v1.0 · August 202617

## Limits, and what is not yet proven

Credibility compounds, and overstating research-stage work is the fastest way to spend it. The following are stated plainly. Model Screwdriver is architecturally proven and research-stage Enclavia's gradient-free weight-steering engine adapts an open model toward a task without moving data offsite and without full fine-tuning, using a dual-headed hypernetwork that routes edits by causal tracing and injects rank-1 deltas. Across 61 logged runs on 10 unseen benchmark suites, steering from a BERT-Base scout into a BERT-Large target, the mechanism is stable: the injected delta is consistently non-zero with an average Frobenius norm of roughly 0.0003, the router converges cleanly, and isolated tasks show real gains of up to 5.3 percentage points without catastrophic forgetting. Aggregate downstream gains across the full unseen suite are not statistically significant, with paired-t p-values averaging approximately 0.36 under a strict rank-6 constraint. The geometry moves toward the target manifold; it does not yet tighten decision boundaries enough for robust zero-shot transfer. Nothing in Sections 4 through 7 depends on this maturing.

The compounding-error model assumes independence R(n) = pn treats step errors as independent, which they are not. Correlated failures make some regimes worse than the curve suggests and others better. The model is used here to establish the shape of the problem, not to predict a specific system's reliability.

Conformal coverage is marginal and assumes exchangeability The guarantee holds on average over the calibration distribution, not conditionally on every subgroup, and it degrades silently under distribution shift. This is a real limitation, not a footnote, and it is the specific reason the change-detection layer exists.

Thresholds do not transfer across domains There are no universal constants here. Entropy profiles, drift tolerances, and repetition baselines differ by task class, model, and corpus. Every deployment requires a calibration phase against in-envelope traffic, and any vendor offering pre-set thresholds that work everywhere is describing something other than what this paper describes.

Astra itself is not yet independently reviewed External mathematicians have not completed review of the manuscripts. The Lean certificates are checkable and reportedly complete, which is a strong signal, but the broader claim about scientific capability is recent and contested. This paper relies on Astra as an illustration of the verification principle, not as a settled result.

Enclavia.ai · v1.0 · August 202618

## Sourcing methodology

Every quantitative claim in Enclavia's published work carries a tier. The purpose is not modesty. In a market where a large share of circulating AI statistics are miscited or stale, being the vendor whose numbers survive scrutiny is a commercial position.

Tier Definition Used in this paper for

## Tier 1 — Independently

reported or benchmarked Primary source, company filing, peerreviewed publication, or multiple independent outlets reporting the same primary document. Astra results and token cost; Alphabet and Amazon capital expenditure and cash flow; published research on semantic entropy, CUSUM, ADWIN, MMR, conformal prediction, and the Hungarian algorithm.

Tier 2 — Derived or vendormeasured Arithmetic from stated assumptions, or measurement taken on Enclavia systems and labelled as such. Compounding-reliability figures; trust-score variance and half-life; Model Screwdriver run statistics; monitoring overhead.

Tier 3 — Forecast or internal model Analyst projection or Enclavia's own cost and capacity model. Gartner cancellation forecast; break-even months; local-inference share targets. Two corrections worth recording, because both circulate widely in this market. The frequently quoted figure of $4–5M per day for clinical trial delay cost is not current; Tufts CSDD's working estimate is in the range of $500K–800K per day, and using the inflated number invites a credibility challenge from any sponsor who knows the literature. Separately, the claim that AI capability is the binding constraint on regulated deployment is contradicted by the cancellation data, which attributes failures to cost, unclear value, and risk controls.

Enclavia.ai · v1.0 · August 202619

## Conclusion and the ask

Astra settled a question that had been open for twenty-seven years and published a certificate anyone can check. The certificate is the transferable lesson, not the theorem. Capability arrived faster than the mechanisms for trusting it, and the domains that closed that gap did so by making verification cheap, external, and adversarial rather than by asking the model to vouch for itself.

Regulated agentic AI has no Lean kernel and will not get one. What it can have is a supervision layer that measures the quantities that actually predict failure, acts on them through a graduated and change-controlled policy, and emits signed evidence as it runs. Entropy and groundedness for confident fabrication. Activationspace distance for off-task drift. Semantic autocorrelation for loops. Knapsack and marginal relevance for context. Sequential change detection for silent degradation. Exponential moving averages and constrained assignment for trust and routing. Conformal prediction for a coverage guarantee that survives contact with an auditor. None of it is novel mathematics. All of it requires access to internal model state that a closed API does not provide.

That is the whole architecture, and the reason it runs on open weights inside the customer's own boundary. The ask Enclavia is taking on a small number of design-partner programmes in regulated, IP-sensitive, or sovereigntybound environments. The engagement is deliberately narrow and measurable. Deploy one Shepherd cell against a real internal workload and measure three things together:

- Data never leaves the perimeter, confirmed by static binary analysis rather than by policy attestation.
- Drift and confident fabrication are caught before an artefact reaches a human, with the intervention log as

evidence.

- The cost crossover lands where the model says it lands, in the programme's own numbers.

If an agentic programme intends to be in the 60% that survives to 2028, the architecture decision is the one to make now, because the instrumentation cannot be retrofitted onto an interface that does not expose the state it needs.

## Dev Roy Founder & Chief Executive Officer, Enclavia.ai

dev@enclavia.ai · 703-984-9981 · enclavia.ai · Fairfax, Virginia

## Your Data · Your AI · Your Way

Enclavia.ai · v1.0 · August 202620
