Whitepaper29 pages • PDF40 min read

AI Enterprise Modernization Platform — White Paper

Platform Recovering verified business intent from legacy government systems — a deterministic-first, sovereign path to modernization with continuous authorization. Your Data. Your…

Length
29 pages
AI Enterprise Modernization Platform — White PaperNEW

AI Enterprise Modernization

Platform Recovering verified business intent from legacy government systems — a deterministic-first, sovereign path to modernization with continuous authorization. Your Data. Your AI. Your Way.

Publisher Enclavia.ai, Inc. (formerly IntraIntel.ai) · Fairfax, VA

Document AEMP-2026-03 · Technical & Executive White Paper

Version / Date Version 1.0 · August 2026

Prepared by Dev Roy, Founder & CEO / Principal Investigator; Enclavia Platform & Architecture team Audience Agency CIO / CTO / CDO / CAIO, authorizing officials, program executives, and modernization architects across DoD, the Intelligence Community, and civilian agencies Distribution Approved for public release. Prepared for discussion. Market figures per cited 2025–2026 analyst reports and U.S. GAO / OMB publications.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 1 of 29 Contents 1 Executive summary.........................................................................................................3 2 The federal legacy crisis...................................................................................................5 3 Why existing approaches fall short....................................................................................8 4 The AEMP approach: derive structure, infer meaning.........................................................10 5 Deep technical architecture.............................................................................................13 6 Sovereign deployment and continuous authorization.........................................................18 7 Federal mission scenarios..............................................................................................20 8 The business case......................................................................................................... 23 9 Evidence and methodology.............................................................................................25 10 The team..................................................................................................................... 26 11 The engagement — our ask........................................................................................... 27 A Appendix A — NIST SP 800-53 Rev. 5 control mapping..........................................................28 B Appendix B — Glossary.................................................................................................... 28 C Appendix C — Selected references.....................................................................................29 Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 2 of 29

Executive summary

The most consequential transactions the U.S. government runs — benefit determinations, tax processing, financial systems of record, logistics — still execute on code almost no one alive fully understands. AEMP recovers what that code means, proves the recovery is correct, and does it inside your own trust boundary.

In May 2025 the Government Accountability Office named eleven of the federal government's most critical legacy systems. They range from 23 to 60 years old, cost roughly $754 million a year just to keep running, and eight of the eleven are still written in COBOL or Assembly. Seven carry known cybersecurity vulnerabilities; eight cannot support a zero-trust architecture. Across the federal IT portfolio, agencies spend more than $100 billion a year, and close to 80 percent of it goes to operating and maintaining what already exists rather than building what comes next. The systems are old, the people who understood them are leaving, and the original design documents are long gone. The code is the only surviving specification — and it records how each system behaves, never why.

The prior generation of modernization tools does not close that gap. Code converters translate COBOL syntax into Java syntax and carry the opacity into a new language; IBM's own engineers call the result "JOBOL." AI coding agents generate plausible descriptions and plausible replacements, but they leave the one question that decides success unanswered and unverified: what does this system actually do for the mission?

The market has split into deterministic transpilers that preserve opacity and probabilistic agents that leave intent unproven. The middle — verified recovery of business intent — is where the risk actually lives, and it is largely unoccupied.

What AEMP does differently AEMP separates two jobs every prior tool has conflated. A deterministic Program Analysis Layer recovers all structure — control flow, data flow, dependencies, types — as facts that are correct by construction, not guessed. AI reasoning is then confined to what it does well: attaching business meaning to structures already proven to exist.

Every recovered business rule is machine-verified against the original code's behavior before it is trusted. The AI is a reasoner over verified facts, never the sole source of understanding. That is precisely what makes the output defensible to an auditor, an inspector general, or an authorizing official.

Underneath sits the same orchestration and control plane Enclavia built for sovereign, contested-environment AI. The Shepherd-AI system runs the work on open-weight models the customer owns, watches each worker's internal state for hallucination, offtask drift, and repetition loops, and corrects a failing worker before its failure reaches a person. An in-built Agentic ATO control plane binds NIST SP 800-53 controls to runtime checks and binary properties and emits signed compliance evidence as the system runs, which is what makes a continuous-authorization target realistic rather than aspirational.

None of it depends on a network connection the customer cannot control. The source code never leaves the building. Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 3 of 29 23–60 yrs age of the 11 most critical federal legacy systems (GAO, 2025) ~$754M/yr to operate those eleven systems ~80% of federal IT spend goes to O&M, not modernization What the customer receives A validated, evidence-linked map of what every legacy system does — an asset most agencies have never possessed. A modern architecture blueprint with decision records, a risk-sequenced migration roadmap, and generated code that has been behavior-tested against the original. Deployment inside your boundary, whether that is a VPC, an onpremises rack, or a fully air-gapped enclave. And a complete audit trail from every legacy artifact to every modern component, plus continuous authorization evidence you can hand to an AO on removable media without ever opening an outbound connection.

The ask We are inviting a small number of design-partner agencies to run a scoped AEMP engagement against a government-provided system and corpus and measure three things with us: that source and data never leave the perimeter, that recovered intent is machine-verified rather than asserted, and that the modernization path is cheaper and faster than a rebuild-from-scratch. Section 11 describes the engagement model.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 4 of 29

The federal legacy crisis

The code is the only surviving specification

Core government functions run on mainframes: tax administration, federal financial management, benefit eligibility and payment, military logistics, and case management at nearly every large agency. Much of it is COBOL and Assembly written across four to six decades. The engineers who designed those systems have retired, the requirements documents have been lost across reorganizations and contract recompetes, and what remains is the source itself. Source code is an exact record of behavior and a silent record of intent. It tells you, precisely, what happens when a claim is adjudicated or a payment is calculated. It does not tell you which rule encodes a statute, which encodes a 1987 policy memo, and which is a workaround someone added to survive a Y2K deadline.

That missing context — the why — is exactly what a modernization team needs and exactly what no tool has reliably recovered. The GAO has documented the scale repeatedly. Its 2016 review found systems like the Treasury's Individual Master File, then roughly 56 years old and written in Assembly, and Defense's nuclear command-and-control system still running on an IBM Series/1 with 8-inch floppy disks. Its 2019 review put the ten most critical systems at about $337 million a year to operate. Its 2025 review raised that to roughly $754 million for eleven systems and found that, of ten critical systems flagged for modernization in 2019, only three had been completed by early 2025.

The retirement clock and the closed pipeline

The workforce problem is not theoretical. Surveys and agency testimony put the typical COBOL maintainer well into a second career stage, and the training pipeline that once produced them has mostly closed. The consequence became national news in April 2020, when several state unemployment systems buckled under pandemic claim volume. New Jersey's governor made a public appeal for COBOL programmers to help repair a system he described as more than forty years old, while roughly 362,000 residents filed for benefits in a matter of weeks. That was a civilian-benefits system under load, in public view. Most federal legacy estates carry the same dependency without the same visibility.

A note on the widely cited numbers Figures like "220 billion lines of COBOL" (a 2017 Reuters estimate) and "over 800 billion lines in daily use" (a 2022 Micro Focus survey of 1,104 organizations) are survey extrapolations, not code censuses, and the often-repeated claims that COBOL handles $3 trillion in daily commerce or 95 percent of ATM transactions trace back to a single 2017 article. We cite them as directional industry estimates and anchor our own case on the primary-source GAO, OMB, and analyst figures used throughout this paper.

What failure costs

The alternative to a disciplined modernization is not the status quo; it is a series of expensive restarts. The IRS has spent on the order of $2 billion since 2009 trying to modernize the 60-year-old Individual Master File, watched a key transition milestone slip about nine years, and in March 2025 paused nearly two dozen modernization programs to regroup. The VA's replacement for a 30-year-old financial and acquisition system saw Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 5 of 29 its cost estimate climb from roughly $2.5 billion in 2019 to about $7.7 billion by late 2023, with full implementation slipping to 2030 — after two earlier replacement attempts since 1998 had already failed. These are not outliers. They are what happens when a program commits to rebuilding a system whose behavior it never fully recovered.

Figure 5 (referenced early). The modernization market bifurcates into deterministic transpilers that preserve opacity and probabilistic agents that leave intent unverified. Verified business-intent recovery, delivered inside the trust boundary, is the unoccupied upperright quadrant.

The mandate environment

Modernization is moving from discretionary spend to compliance-adjacent obligation. FITARA (2014) made agency CIOs accountable for IT investment outcomes; the MGT Act (2017) created the Technology Modernization Fund and agency working-capital funds; and Congress reauthorized the TMF through fiscal 2026. The Fund has invested roughly $1.03 billion across 68 projects since 2018, though realized savings so far are modest, which is a fair measure of how hard execution gets when the underlying systems are opaque. On the AI side, OMB Memorandum M-25-21 (April 2025) — which superseded M-24-10 and implements Executive Order 14179 — directs every agency to appoint a Chief AI Officer, to test high-impact AI before deployment, and to modernize the data, infrastructure, and security posture that AI depends on. AEMP sits at the intersection of both mandates: it is a modernization capability that is itself a governed, testable AI system.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 6 of 29

Why now

Three things changed in the last two years. Open-weight models in the Qwen, Llama, Mistral, and DeepSeek families closed most of the quality gap with proprietary frontier models on the bounded, tool-using reasoning that code understanding requires, so nearfrontier capability now runs inside an enclave rather than behind a vendor API.

Sovereignty pressure rose in parallel: in a 2026 survey of roughly 5,000 senior decisionmakers, 35 percent of chief AI officers named building and running AI in private or sovereign settings as their single biggest adoption barrier. And the reliability problem became impossible to ignore, with Gartner projecting that more than 40 percent of agentic AI projects will be canceled by the end of 2027, mostly for weak risk controls and unclear value. The demand is now for AI that is capable, sovereign, and provably under control. That is the exact envelope AEMP was built for.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 7 of 29

Why existing approaches fall short

Every serious vendor now markets "business rule extraction." Almost none of them verify the rules they extract. Understanding why is the fastest way to understand what AEMP is.

Syntax translation preserves opacity

The oldest and most common approach is language translation: a rules engine or a model rewrites COBOL into Java, line for line. The behavior is preserved, which sounds like success until an engineer has to maintain the result. Translated code inherits the structure of the original — the GO TOs, the shared working-storage, the copybook coupling — and expresses it in a language that has no idiom for any of it. IBM's own team named the failure mode: "JOBOL," Java that reads like COBOL and is no more comprehensible than the source. The agency has swapped one language it cannot maintain for another, and it still does not know which rule implements which policy.

The market bifurcation

Under the marketing, the field splits cleanly in two. On one side are deterministic transpilers — TSRI, the core conversion paths in the large cloud services, the recompileand-replatform tools — which are reliable precisely because they are mechanical, and which preserve opacity for the same reason. On the other side are probabilistic LLM agents that read a codebase and produce fluent descriptions and fluent rewrites, and which can be genuinely useful for documentation and batch jobs, but which offer no guarantee that what they assert is what the code does. Figure 5, shown earlier, places the current field on two axes: how faithfully the output captures business intent, and whether the system can run inside a sovereign boundary. The tools cluster in the lower and left regions. The upper-right — verified intent, sovereign by design — is close to empty.

Behavioral equivalence is not intent

The most sophisticated incumbents deserve credit for going further. Google's Dual Run executes real workloads on the mainframe and on the cloud target simultaneously and compares outputs; Mechanical Orchard observes live data flows and rebuilds against them. Both are real engineering, and both validate behavioral equivalence — the new system produces the same outputs as the old one. That is valuable, but it is not the same as recovering intent. A system can reproduce forty years of behavior, bugs and dead branches included, without anyone learning which behaviors are required by law, which are obsolete, and which were never supposed to happen. Equivalence answers "does it still do the same thing?" AEMP is built to answer "what is it supposed to do, and why?" — and to prove its answer.

Cloud-lock and the sovereign gap

For a federal buyer the deployment model is often disqualifying before the technical merits are even weighed. The leading agentic modernization services are cloud-exclusive by architecture; their reference designs assume managed cloud data stores and connectivity that a classified enclave, a SCIF, or a sovereignty-bound program simply cannot provide. The on-premises exceptions are narrow: one strong architecturalanalysis tool is on-prem-capable but addresses Java and .NET rather than mainframe Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 8 of 29 COBOL, and one mainframe assistant can be self-managed on a container platform. That leaves a real gap for the estates that matter most to national security: the ones that cannot phone home.

The unoccupied position

Put the three failures together and the opening is specific. Deterministic transpilers are trustworthy but opaque. AI agents are illuminating but unverified. Behavioral tools prove equivalence but not intent. And the whole capable end of the market assumes the cloud.

AEMP is designed to sit exactly where none of them do: verified recovery of business intent, produced by AI that is held accountable by deterministic ground truth, running entirely inside the customer's boundary. The rest of this paper is how that is built and how it is proven.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 9 of 29

The AEMP approach: derive structure, infer meaning

A separation of concerns

AEMP's central design decision is to refuse to let an AI model be the source of truth about a legacy system. Instead it divides the problem into two parts with a hard boundary between them. Deterministic program analysis recovers everything that can be known with certainty from the code — the control-flow graph, the data-flow and defuse chains, the call graph and dependencies, reconstructed record and type structures.

These are not inferences; they are facts, derived by the same class of techniques a compiler uses, and they are correct by construction. Only once that skeleton exists does AI reasoning enter, and its job is narrow: attach business meaning to structures that have already been proven to exist. A model may propose that a particular paragraph implements an overpayment-recovery rule, but it proposes that about a real, identified control path, not about a hallucinated one.

Figure 1. Separation of concerns. Deterministic analysis recovers structure that is correct by construction; AI reasoning attaches meaning to that structure; and every recovered rule is replayed against the original code's behavior before it is trusted.

The eight-layer, deterministic-first architecture

The platform is organized as eight layers, and the ordering is the point. The lower four are deterministic: ingest and inventory, dialect-aware parsing and copybook resolution, structural analysis, and the Enterprise Knowledge Graph that unifies the results into a single queryable model of the estate. The upper four are where AI reasons on top of that proven foundation: the Business-Intent Model, a proposed target architecture with decision records, a risk-sequenced migration plan, and finally build and verification.

Meaning is only ever attached on top of structure that has already been established. Nothing in the AI layers is trusted until it has been checked back against the deterministic layers beneath it. Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 10 of 29 Figure 2. The eight-layer, deterministic-first architecture. The foundation (layers 1–4) is mechanical and certain; the AI reasoning layers (5–8) build on it and are verified against it.

The pipeline

Operationally the layers run as a pipeline: Discover, Program Analysis, Knowledge

Graph, Business Intent, Architect, Plan, Build. Deterministic analysis feeds the

knowledge graph; reasoning models recover the Business-Intent Model of processes, rules, objects, and events; an AI architect proposes a cloud-native or hybrid target with explicit rationale; a planner sequences migration into risk-ascending waves so the safest changes ship first and confidence compounds; and specialized agents generate and verify code. Human approval gates sit at every architectural, security, and deployment decision. The system proposes and proves; a qualified person decides.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 11 of 29 Figure 3. The AEMP pipeline, from raw estate to behavior-tested code, with human approval gates at each consequential decision and machine verification on the generated output.

Machine verification

Verification is not a final QA step bolted on at the end; it is the mechanism that makes the AI layers safe to trust. When a reasoning model recovers a business rule, that rule is expressed in a form that can be executed or checked against the original program's behavior on representative inputs. If the recovered rule and the legacy code disagree, the rule is wrong, and it is sent back rather than published. The effect is that the Business-Intent Model an agency receives is not a model's opinion about the code; it is a set of claims each of which has survived a confrontation with the code itself. That is the difference between a document an inspector general will question and one they can rely on.

Human approval gates

Autonomy is bounded on purpose. AEMP is designed to do the enormous volume of analysis no human team could do by hand — reading millions of lines, tracing every data path, cross-referencing every dependency — and then to stop and present its conclusions at each decision that carries real consequence. An architecture choice, a security control, a cutover: each is a gate a person must pass. This is what keeps the platform on the right side of OMB's human-oversight requirements for high-impact AI, and it is what lets an authorizing official sign with confidence rather than faith.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 12 of 29

Deep technical architecture

This section is written for architects and technical evaluators. It describes how the platform is put together, how it keeps its own reasoning honest, and exactly what is proven versus what is still research.

Reference architecture

Everything runs inside the customer's trust boundary. Inputs enter on the left. A threetier orchestration hierarchy — the Shepherd-AI system — allocates hardware, dispatches work, and monitors the workers continuously. A shared knowledge, memory, and logging plane holds the Enterprise Knowledge Graph, the Business-Intent Model, vector and structured stores, conversation memory, a worker-competence map, and a time-indexed system log written to disk so an operator always knows what the system is doing. A security plane encrypts everything at rest, seals keys to the hardware, signs binaries and weights, and writes a tamper-evident record of every dispatch. Alongside it, the in-built Agentic ATO control plane produces continuous authorization evidence for out-of-band export. No component holds an outbound network path.

Figure 4. AEMP reference architecture. The Shepherd-AI orchestration tiers, the knowledge and memory plane, the security plane, and the Agentic ATO control plane all operate inside the customer trust boundary, with zero network egress.

The deterministic Program Analysis Layer and the Knowledge Graph

The deterministic core parses each artifact with dialect awareness — IBM Enterprise COBOL, GnuCOBOL, PL/I, JCL, and Assembly variants differ enough to matter — and resolves copybooks, includes, and file/record layouts into concrete data structures. From the parsed forms it builds the analyses a compiler back end would recognize: control-flow Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 13 of 29 graphs per program and across the call graph, data-flow and def-use chains that follow a field from the record that defines it to every point that reads or mutates it, and type and record reconstruction that recovers the shape of data the original programmers held only in their heads. All of it lands in the Enterprise Knowledge Graph, a single model that spans the whole estate and can be queried across program boundaries: which jobs write this field, which rules depend on that flag, what breaks if this file changes. Because these facts come from static analysis rather than inference, they are the fixed ground truth that everything above is measured against.

The Business-Intent Model

On top of the graph, reasoning models recover a Business-Intent Model with four kinds of element: business processes (the end-to-end flows a program participates in), business rules (the conditional logic that encodes policy), business objects (the domain entities the data represents), and domain events (the state changes that matter to the mission).

Each element is anchored to specific nodes in the knowledge graph, so every claim about meaning is traceable to the exact code and data it came from. This is what turns a pile of source into an asset an agency can reason about: a map that says, in business terms, what the system does, with a link from every statement back to the line that justifies it.

Shepherd-AI orchestration

The reasoning work is run by Shepherd-AI, a three-tier orchestrator. The Tier 1 Meta- Meta Agent owns resources: it allocates models to hardware according to the compute and memory each needs, runs the reinforcement-learning pipeline that tunes selection over time, manages session cleanup, and owns the retrain-and-steer bucket. The Tier 2 Meta Agent owns the work: it decomposes a request into tasks, selects the most trusted model for each, asks a human when it needs clarification, verifies worker output, and reprompts or reroutes on failure. Tier 3 is the pool of open-weight worker models — program analysis, business-intent reasoning, architecture, planning, build and test — each specialized and scored over time. The workers are swappable and multiple, which means the platform is never hostage to a single model's weaknesses; if a better open model appears, it enters the pool and earns its place by performance.

The drift monitor and the self-heal loop

Autonomous multi-step agents fail in ways a single chatbot does not. A small per-step error compounds across a long plan, and a worker that quietly goes off-task can corrupt an entire workflow before anyone notices. Behind a closed cloud API none of this is visible until it reaches the user. AEMP's answer is instrumentation: the drift monitor runs alongside generation and watches three internal signals, each aimed at one failure mode. Token-distribution variance catches hallucination, because a collapsing output distribution is the signature of a model committing confidently to fabricated content.

Hidden-state entropy measured against a baseline captured at the start of generation catches off-task drift. Semantic comparison across a sliding window of recent output catches repetition loops before they freeze the workflow.

The response is graduated rather than binary. On a first trip the partial output and a measured drift value go back to the Meta Agent, which re-prompts the worker or routes the task to a different model. A worker that keeps tripping is escalated to the Meta-Meta Agent's retrain-and-steer bucket. Trust is earned and remembered: after each task a worker's score updates on an exponential moving average, new = α·reward + Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 14 of 29 (1−α)·current, with a default α of 0.15 and a reward of 1 on success and 0 otherwise.

Over time the Meta Agent learns which models to rely on for which tasks — a competence map no cloud vendor can hand you, because building it requires seeing inside the models as they work. Figure 5. The drift monitor and graduated self-heal. Three internal signals map to three failure modes; the response escalates from re-prompt to reroute to bench-and-retrain; and a trust score updates after every task so selection improves with use.

The mathematics of agentic control

The drift signals above are the operational summary of a deeper measurement framework, published separately as "Proof Over Promise: The Mathematics of Agentic Control." Rather than assert robustness, AEMP measures it, using a set of quantities that are each defined, computed, and logged as the system runs: Measured quantity What it detects Entropy / groundedness fabrication index Confident fabrication — output committing to content not supported by the retrieved context Mahalanobis activation drift Off-task movement of the hidden state away from a task-anchored baseline distribution Semantic repetition score Stalls and loops, via similarity across a sliding window of recent generation Composite drift index A single fused signal combining the above for a graduated intervention threshold Context knapsack / MMR selection Whether the right evidence was placed in-context under a token budget, balancing relevance and diversity CUSUM / Page-Hinkley / ADWIN change-point A statistically meaningful shift in behavior over a run, not just a noisy spike Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 15 of 29 Measured quantity What it detects Trust EMA with Hungarian assignment Per-model competence over time, and the optimal assignment of models to tasks Conformal abstention When the system should decline to answer rather than guess, at a calibrated error rate The important property is that these are numbers, not adjectives. A reviewer can ask what the fabrication index was on a given run, what the change-point detector flagged, and where the system chose to abstain. Reliability becomes an auditable measurement rather than a marketing claim — which is the standard a federal authorizing environment should demand of any autonomous system.

Model Screwdriver — reported honestly

Owning the stack is half the thesis; adapting the models to an agency's own code and domain without moving data off-site is the other half. Model Screwdriver is Enclavia's research engine for a third path between prompt-only use and full fine-tuning. It pulls a low-rank task vector from a small scout model and maps it into a larger target through a dual-headed hypernetwork: a router head uses causal tracing to choose which layers to edit, and a generator head produces rank-1 weight deltas injected into those layers.

Because it touches a narrow set of components and never overwrites base weights, it adapts behavior while preserving the model's pretrained priors, which avoids the catastrophic forgetting that makes naive fine-tuning risky.

What is proven, and what is not Across 61 logged runs on 10 unseen benchmark suites, steering from a BERT-Base scout into a BERT-Large target, the mechanism is validated and stable: the injected weight delta is consistently non-zero (average Frobenius norm ≈ 3×10⁻⁴), the router converges cleanly, and isolated tasks show real gains — up to +5.3% on Tweet- Emotion, +2.5% on Banking intent, +2.0% on spam — without catastrophic forgetting.

But under a strict rank-6 constraint, aggregate improvements across the full unseen suite are fractional, with paired-t p-values averaging about 0.36. The geometry moves toward the target manifold; it does not yet tighten decision boundaries enough for robust zero-shot lift. So Model Screwdriver is architecturally proven and research-stage. Critically, the deployable core of AEMP — sovereign orchestration with verified recovery and self-healing reliability — does not depend on it maturing.

It is upside on top of a system that already works.

Why sovereignty is the precondition for measurement

There is a reason the drift monitor and the measurement framework are only possible on owned, open-weight models. Every signal AEMP relies on is computed from the model's internals: the token-distribution, the hidden-state activations, the residual stream. A closed frontier API exposes none of that. It returns text. You cannot measure the entropy of a distribution you are not allowed to see, or track activation drift in activations that never leave the vendor's data center. Sovereignty, in this architecture, is more than a compliance posture. It is the technical precondition for being able to prove the system is Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 16 of 29 under control at all. On-premises ownership is what makes measurable reliability possible, and measurable reliability is what makes federal authorization defensible.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 17 of 29

Sovereign deployment and continuous authorization

Deployment envelope

AEMP deploys as a hardened container image inside whatever boundary the mission requires: a VPC in a government cloud region, an on-premises rack, or a fully air-gapped enclave with no external connectivity of any kind. The image is the same across targets; only the model profile and hardware allocation change. A reference cell — roughly five workers plus an orchestrator with context headroom — fits in a single rack of commercially available accelerators. Source code and data stay put: analysis runs where the code already lives, and nothing about the estate is transmitted anywhere. For agencies whose modernization corpus is itself sensitive or classified, this is the difference between a tool they can use and one that policy forbids on the first slide.

Zero egress by construction

"No data leaves" is often a policy promise. In AEMP it is a property of the binary. The system carries no outbound network code path — no DNS resolver, no TLS client, no time-server client, no telemetry — so there is no channel through which data could exit even if something tried to send it. This is verifiable by static binary analysis rather than by trust: a reviewer can confirm the absence of those capabilities directly. Logs are written to disk in the operator's control and are destroyable at session end. Everything at rest is encrypted with AES-256-GCM under keys sealed to the hardware's TPM, binaries and model weights are signed and integrity-checked at load, access is governed by RBAC, and every dispatch leaves a tamper-evident record.

The in-built Agentic ATO control plane

The reason a continuous-authorization timeline is realistic here, rather than the usual multi-year slog, is that the system generates its own authorization evidence as it runs. Four mechanisms make it work. Each relevant NIST SP 800-53 control is bound to a runtime check or a static binary property and emits a signed, present-state attestation, so posture is a live fact rather than a document assembled after the fact. Controls that are easier to prove by absence — no DNS, no TLS, no NTP, no telemetry — are confirmed by third-party static analysis and stay stable across builds. Signed posture is exported out-of-band to the authorizing official via removable media or a sanctioned cross-domain path, keeping the compliance channel separate from the operational one. And the container is engineered for ingestion into a hardened registry and a Platform One continuous-ATO, so authorization can be inherited rather than re-earned at every site.

We target a six-month path to continuous authorization on this basis.

NIST SP 800-53 Rev. 5 control mapping

The mapping below is a summary; Appendix A carries the fuller control set. The point is not that AEMP checks boxes but that many controls are satisfied by architecture rather than by compensating process — an outbound flow that does not exist cannot be misused, and an absence is cheaper to assess than a mitigation.

Control Family How AEMP satisfies it AC-2 / AC-4 Account mgmt; information-flow enforcement Local accounts and RBAC; outbound information flows are architecturally absent Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 18 of 29 Control Family How AEMP satisfies it AU-2 / AU-12 Audit events; record generation Live logger writes a signed, replay-capable record per dispatch CM-3 / CM-5 Configuration change control Signed releases; SHA-256 weight verification at load; change gated CP-13 Alternative security mechanisms Operates with no network, no external time service, and no external identity provider IA-2 / IA-4 Identification & authentication TPM-2.0 device binding; CAC/PIV where present; offline identity SC-7 / SC-8 / SC- Boundary; transmission; at-rest protection Zero egress; AES-256-GCM for storage and any local sync; TPM- sealed keys SI-4 / SI-7 Monitoring; software & information integrity Drift monitor; integrity monitor; tamper-evident audit log

The cATO on-ramp: Platform One and Iron Bank

Continuous authorization is a defined state, not a slogan. The Defense Department's guidance rests on three pillars: continuous monitoring of controls inside the boundary, an active cyber-defense capability, and an approved DevSecOps reference design including a secure software supply chain. AEMP is built to satisfy all three, and to plug into the existing DoD on-ramp. Platform One holds a continuous ATO that programs building on it can leverage, and Iron Bank supplies more than a thousand hardened, continuously scanned container images as a body of evidence. One nuance matters and we state it plainly: Iron Bank membership does not itself confer an ATO — the reciprocity that makes "authorize once, use many" real lives in the Platform One continuous ATO and its DevSecOps reference design, with Iron Bank providing the hardened supply chain that makes downstream authorization fast. AEMP's container is engineered to fit that path rather than to route around it.

CMMC and 800-171 posture

For work touching controlled unclassified information, the compliance calendar is now concrete. The CMMC program rule took effect in December 2024 and the acquisition clause in November 2025, with Level 2 third-party assessments phasing in through 2026.

Level 2 is the 110 controls of NIST SP 800-171, and Level 3 adds a further set from 800- 172. AEMP's architecture aligns to 800-171 and 800-53 by construction — the same zeroegress, encrypted-at-rest, TPM-bound, audited design that supports the Agentic ATO also supports a CMMC posture appropriate to the award level — so a contractor deploying AEMP is not bolting compliance on afterward.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 19 of 29

Federal mission scenarios

The architecture is mission-agnostic. These five scenarios make it concrete across the civilian, defense, and intelligence estates. None require a network connection, and none send a byte outside the boundary.

Civilian benefits agency — COBOL eligibility and payment

Constraint A benefits agency runs eligibility determination and payment on a multi-decade COBOL system whose rules encode layers of statute, regulation, and settlement agreements. A single wrong rule means an improper payment or a wrongful denial, and the staff who understood the original logic have retired.

What the agency gets AEMP produces a rule-by-rule Business-Intent Model that says, in plain terms, which conditions drive each determination, with a link from every rule back to the exact code and the data it reads. Each recovered rule is machine-verified against the legacy behavior before it is trusted, so the agency can finally separate the rules required by law from the accidental behavior it has been faithfully reproducing for years — and modernize with an audit trail an inspector general can follow.

AEMP components that carry the load Deterministic data-flow and control-flow recovery over the eligibility programs; business-intent reasoning workers grounded in the knowledge graph; machine verification against original behavior; human approval gates on every rule that changes a determination.

Treasury-class financial system of record

Constraint A financial master file, decades old and part COBOL, part Assembly, is the authoritative record for transactions the government cannot afford to get wrong. Prior modernization attempts stalled because no one could prove the replacement would behave identically where it must and correctly where the original was flawed.

What the agency gets AEMP recovers the full structure of the master file and its jobs deterministically, builds a verified model of the calculations and postings, and sequences migration into riskascending waves so the least risky components move first and confidence compounds before anything critical is touched. Generated components are behavior-tested against the original, and the complete legacy-to-modern audit trail gives auditors traceability they have never had.

AEMP components that carry the load The deterministic Program Analysis Layer and Enterprise Knowledge Graph for exact structure; verified Business-Intent Model for the financial logic; the migration planner for wave sequencing; build-and-test agents under drift monitoring.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 20 of 29

DoD business system in a disconnected enclave

Constraint A defense business system must be modernized inside an enclave with no path to a commercial cloud, and the modernization corpus itself is sensitive. Every cloudhosted modernization service is disqualified on the first requirement.

What the agency gets The same AEMP image runs on a single rack inside the enclave with zero network egress, so the entire analysis-to-build workflow happens where the code already lives. The Agentic ATO control plane accrues signed authorization evidence throughout and exports it out-of-band to the authorizing official, turning what is normally a multi-year accreditation into a continuous, evidence-backed state.

AEMP components that carry the load Sovereign single-rack deployment; zero egress by construction, verifiable by static binary analysis; the in-built Agentic ATO plane; TPM-sealed keys and signed weights on the security plane.

Intelligence Community — air-gapped code understanding

Constraint An IC program needs to understand and modernize legacy analytic and mission software that can never touch an outside network, and where even the fact of what the code does is sensitive. Reasoning about the code must happen entirely inside the fence.

What the agency gets AEMP performs full deterministic analysis and verified intent recovery air-gapped, on open-weight models the program controls, and keeps all logs destroyable at session end. Because reliability is measured from the models' own internals rather than inferred, the program gets an auditable record of where the system was confident, where it drifted, and where it chose to abstain — inside the boundary, with nothing exported that should not be.

AEMP components that carry the load Air-gapped operation on owned models; the measurement framework (fabrication index, activation drift, conformal abstention) computed on-prem; RBAC scoped to the program; session-destroyable logs.

Health agency — clinical and financial legacy

Constraint A health agency carries intertwined clinical and financial legacy systems under strict privacy obligations, where sensitive records cannot leave the agency network and the interaction between clinical rules and payment logic is poorly understood.

What the agency gets AEMP maps the clinical and financial logic and the data flows between them into a single verified model without any record leaving the agency, so the agency can modernize the Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 21 of 29 payment path without breaking the clinical one, and can show privacy and audit teams exactly how data moves. Enclavia's compliance-native heritage — the platform was built to HIPAA, SOC 2, and FDA-governed standards — carries directly into this setting.

AEMP components that carry the load On-prem deployment satisfying privacy requirements; cross-domain data-flow recovery in the knowledge graph; document-intelligence and analysis workers; the audit trail from every legacy artifact to every modern component.

Across all five, the pattern is identical. The mission supplies the system and the constraint; AEMP supplies deterministic ground truth, verified intent, self-healing reliability, and an audit trail; and the Agentic ATO plane supplies the authorization evidence. Nothing in the loop depends on a connection an adversary or a policy can take away.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 22 of 29

The business case

Why 40 percent of agentic projects get cancelled

Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027, and its stated reasons are cost, unclear value, and inadequate risk controls. This is not a model-quality problem; it is a control problem. Agencies are not short of ambition

  • adoption intent is near-universal in every recent survey — they are short of ways to

trust an autonomous system in production. AEMP is designed around that exact gap: the deterministic ground truth, the machine verification, the drift monitoring, and the continuous authorization evidence exist precisely so a program can be in the roughly 60 percent that survives. The application-modernization services market is real and growing — on the order of $22.7 billion in 2025, rising toward $51 billion by 2031 at about 15 percent a year, with mainframe modernization a further $8–13 billion — but the money follows programs that can show control, not just capability.

The economics of ownership

Per-token and per-seat pricing is a tax on success: it scales with use, never depreciates, and leaves no asset behind. Owning the stack inverts that. AEMP runs on open-weight models on hardware the agency owns, so the marginal cost of additional analysis is close to the cost of electricity, and at the end of the program the agency holds an asset — the knowledge graph, the verified intent model, the competence map, the audit trail — rather than a receipt. Organizations self-hosting open models on sustained, high-volume workloads report savings in the 60–80 percent range against comparable cloud API spend; we present that as an industry-reported figure rather than a guarantee, because the crossover depends on utilization. The structural point holds regardless: a modernization program is a high-volume, long-running workload, which is exactly the profile where ownership wins.

Risk-ascending delivery

AEMP is delivered the way a careful program manager would want it: not as a big-bang rebuild but as a sequence of risk-ascending waves. The platform's analysis makes it possible to sequence work so the safest, most isolated components modernize first and each wave builds verified confidence before the next. A pilot proves the model on a bounded slice of the estate, the verification results are inspected by the agency's own people, and only then does scope expand. This is how a program avoids the failure mode that produced the multi-billion-dollar overruns in Section 2: committing everything before understanding anything.

Producibility, scaling, and distribution

Distribution rides the existing federal rails. A hardened container through a recognized registry and a Platform One continuous ATO gives DoD-wide reciprocity, so per-site reaccreditation stops being the scaling bottleneck. The platform scales from a single rack to a program-wide fleet without new central infrastructure, and it depends on no single vendor's model, because the worker pool is open-weight and swappable. The commercial model is pay-to-own: a platform license plus annual support, model integrations, and tuning, landing through scoped federal pilots and expanding to enterprise and sovereign accounts.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 23 of 29

Risks and mitigations

Risk Description Mitigation

Technical Model Screwdriver's aggregate downstream gains are not yet statistically significant The deployable core does not depend on it; the roadmap lifts the rank constraint and scales to decoder-only models Reliability An autonomous worker collapses in production Drift monitor plus graduated self-heal; the failure of any single worker is isolated from the workflow Accreditation A continuous-ATO timeline depends on the AO and cross-domain access Begin the RMF process in parallel; pre-stage signed, out-of-band evidence from day one Supply chain Model provenance and update integrity Pinned, signed open weights; reproducible builds; SBOM in the hardened registry Adoption Legacy estates are heterogeneous and messy Bounded pilot on a representative slice; agency inspects verification results before scope expands Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 24 of 29

Evidence and methodology

A white paper is only as good as the discipline behind its claims. We label every quantitative claim in this document by the strength of its source, so a reviewer can weigh each accordingly. Tier Meaning Examples in this paper Tier 1 Independently reported by a primary source GAO system ages and operating costs; OMB M-25-21 requirements; DoD cATO guidance; NIST control families; analyst market sizes Tier 2 Derived or vendor-measured Model Screwdriver's 61-run results; competitor deployment models and positioning; the self-hosting savings range Tier 3 Forecast or internal model Reference-cell sizing; the six-month continuous-ATO target; break-even projections

Proven versus research-stage

We draw a hard line between what ships and what is still research. The deployable core

  • deterministic program analysis, the Enterprise Knowledge Graph, verified businessintent recovery, Shepherd-AI orchestration, the drift monitor and self-heal loop, the

security plane, and the Agentic ATO evidence model — is engineering, not aspiration. Model Screwdriver is reported as architecturally proven and research-stage, with its non-significant aggregate results stated in Section 5.7 rather than buried. AEMP does not depend on Model Screwdriver maturing; the on-prem adaptation it promises is upside on a system that already works without it.

What a reviewer can verify independently

Three of our central claims are checkable without taking our word for it. Zero egress is a property of the binary and can be confirmed by static analysis, since a system with no DNS, TLS, NTP, or telemetry code path has no channel to exfiltrate data. The NIST control mapping uses the same control families a hardened-registry assessment uses, so the reciprocity path is concrete rather than notional. And the reliability claims are measurements — the fabrication index, activation drift, change-point flags, and abstention decisions are logged values a reviewer can inspect on a real run. We would rather be evaluated on evidence than on adjectives.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 25 of 29

The team

Enclavia.ai (formerly IntraIntel.ai) is based in Fairfax, Virginia. The leadership and advisory bench pair federal mission delivery with FDA-governed commercial AI, and together represent well over a century of combined experience in platform and systems architecture. The principal investigator for a federal engagement is Dev Roy, accountable for technical execution and the design of the Agentic ATO evidence model.

Name Role Relevant expertise

Dev Roy Founder & CEO —

Principal Investigator

20+ years across enterprise, platform, and AI architecture; federal delivery across 9+ agencies; program lead Shanon Roy COO & Co-Founder Finance and operations; scaling the commercial engine Brian Hoffman CTO & Co-Founder Secure enterprise architecture and compliant AI systems Amee Khetan Chief Growth Officer Attorney; 20+ years across legal, compliance, and software growth

Brandon Dean AI/ML Scientist Machine-learning research; mathematics

Subhashish Chatterjee Enterprise AI Architect AI/ML enterprise architecture

Raj DasGupta Advisor CTO, RIVA Solutions; cloud and cybersecurity Mukesh Pandey Advisor 20+ years in AI strategy (ex-Amazon, ex-Google)

Dr. Shishir Khetan / Alex

Lee Advisors Senior medical director; 30+ years FDA compliance Execution experience behind the platform includes hybrid multi-cloud and edge architecture across federal departments, RMF and ATO authoring and continuousauthorization practice, governed AI delivery under HIPAA, SOC 2, and FDA / 21 CFR Part 11, and production DevSecOps with hardened-container distribution.

Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 26 of 29

The engagement — our ask

We are inviting a small number of design-partner programs — in regulated, IP-sensitive, or sovereignty-bound mission areas — to run an AEMP engagement against a government-provided system and corpus and measure three things with us. First, that source code and data never leave the perimeter, verifiable by static binary analysis.

Second, that recovered business intent is machine-verified against the original behavior rather than asserted by a model. Third, that the modernization path is meaningfully cheaper and faster than a rebuild-from-scratch, measured in the agency's own numbers.

A bounded pilot on a representative slice of an estate is enough to demonstrate all three, and it is the right first step before any commitment to scale. If a modernization program needs to be in the 60 percent that survives, the architecture decision is the one to make now. We request a scoped pilot against a government-provided scenario and corpus, and we will measure the outcome with you.

Contact: Dev Roy, Founder & CEO · Enclavia.ai · Fairfax, VA · 703-984-9981 · dev.roy@enclavia.ai · enclavia.ai Your Data. Your AI. Your Way. Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 27 of 29

A Appendix A — NIST SP 800-53 Rev. 5 control mapping

Selected controls most load-bearing for a sovereign, on-premises AI system. Rev. 5 defines twenty control families and roughly 1,196 controls including enhancements; the families below are the ones AEMP's architecture most directly addresses. Where a control is satisfied by the absence of a capability, that absence is confirmable by static analysis rather than by process.

Family Representative controls AEMP implementation AC — Access Control AC-2, AC-3, AC-4, AC-6 Local accounts, RBAC, least privilege on model and data access; outbound information flows architecturally absent AU — Audit & Accountability AU-2, AU-6, AU-9, AU-12 Live time-indexed logger; signed, replay-capable record per dispatch; tamper-evident storage CM — Configuration Mgmt CM-2, CM-3, CM-5, CM-8 Signed releases; SHA-256 weight verification at load; SBOM; change gated by human approval CP — Contingency Planning CP-10, CP-13 Operates with no external network, time service, or identity provider; journaled state; clean restart IA — Identification & Auth IA-2, IA-4, IA-5 TPM-2.0 device binding; CAC/PIV where present; offline non-person-entity identity SC — System & Comms Protection SC-7, SC-8, SC-12, SC-28 Zero egress; AES-256-GCM at rest and for any local sync; TPM-sealed keys; boundary by construction SI — System & Info Integrity SI-4, SI-7, SI-10 Drift monitor and measurement framework; integrity monitor; input validation; abstention on low confidence PT / SR — Privacy & Supply Chain PT-2, SR-3, SR-4 Data never leaves the boundary; pinned, signed open weights; reproducible builds; provenance in the registry

Term Definition

Agentic ATO An in-built control plane that binds NIST controls to runtime checks and binary properties and emits signed authorization evidence continuously Business-Intent Model The recovered, verified model of a legacy system's processes, rules, objects, and events, anchored to the knowledge graph cATO Continuous Authorization to Operate; a maintained state of authorization built on continuous monitoring, active cyber defense, and an approved DevSecOps design Drift monitor The mechanism that watches a worker's internal state for hallucination, off-task drift, and repetition loops and triggers graduated self-heal Enterprise Knowledge Graph The unified, queryable model of the whole estate produced by deterministic analysis; the ground truth AI reasoning is measured against Model Screwdriver Enclavia's research-stage engine for gradient-free, on-prem weight steering that adapts a model without moving data off-site Shepherd-AI The three-tier (Meta-Meta / Meta / Worker) orchestrator that runs the reasoning work on open-weight models under continuous monitoring Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 28 of 29

Term Definition

Zero egress The property, verifiable by static binary analysis, that the system contains no outbound network code path of any kind C Appendix C — Selected references

U.S. Government Accountability Office, "Information Technology: Agencies Need to Plan

for Modernizing Critical Decades-Old Legacy Systems" (GAO-25-107795, 2025);

"Information Technology: Agencies Need to Develop Modernization Plans for Critical

Legacy Systems" (GAO-19-471, 2019); "Information Technology: Federal Agencies Need

to Address Aging Legacy Systems" (GAO-16-468, 2016); Technology Modernization Fund

review (GAO-26-107737, 2026). Office of Management and Budget, Memorandum M-25-21, "Accelerating Federal Use of

AI through Innovation, Governance, and Public Trust" (April 2025); companion M-25-22

on AI acquisition; Executive Order 14179.

U.S. Department of Defense CIO, "Continuous Authorization To Operate (cATO)"

memorandum (February 2022) and "Continuous Authorization Implementation Guide," v1.0 (March 2024); NIST SP 800-37 Rev. 2 (Risk Management Framework); NIST SP 800-53 Rev. 5 (Security and Privacy Controls); NIST SP 800-171 Rev. 2 and the CMMC program rule (32 CFR Part 170) and acquisition clause.

Gartner, "Predicts 2025: Agentic AI" press release (June 2025); MarketsandMarkets, Application Modernization Services and Mainframe Modernization market reports (2025); NTT DATA Global GenAI / Sovereign AI report (2026); Microsoft Research, BitNet b1.58 (2025).

Enclavia.ai, "Proof Over Promise: The Mathematics of Agentic Control," v1.0 (2025); "AI in a Box: Sovereign Agentic AI with In-Built Agentic ATO," DIU submission (2025). Vendor positioning drawn from public materials of IBM (watsonx Code Assistant for Z),

Amazon Web Services (AWS Transform), Google Cloud (Mainframe modernization / Dual

Run), vFunction, Mechanical Orchard, and others as of mid-2026. Enclavia.ai, Inc. · Fairfax, VA · Your Data. Your AI. Your Way. Approved for public release · Prepared for discussion Your Data. Your AI. Your Way. Page 29 of 29

By downloading, you agree to our Terms of Service and Privacy Policy. This resource is for personal and organizational use.