Last updated: 2026-10-05
Changing the Model: Why AI Systems Need an Exit Architecture
An adapter is necessary for replacing a model, but it is not sufficient. What a substitution has to demonstrate, and what a system should be designed to make easy
In a well-designed system, replacing a model should look like an ordinary change. The application should not need rebuilding, and the new component should plug into the same place as the old one. That expectation is reasonable. It is also easy to overestimate, and the reporting on the United States Department of Defence's removal of Anthropic's tools shows why.
The BBC reported in its account of the Pentagon's decision that the department "has ceased the use of Anthropic products", months after Defence Secretary Pete Hegseth designated the company a supply chain risk. The article says Claude was embedded in Maven Smart System, a larger data platform operated by Palantir, and that it was used for intelligence gathering and analysis, including work on satellite imagery and drone footage. It also says that the removal had been planned for late August, and that "it is not clear why there has been a delay". A former defence official quoted in the same article said that once tools become integrated, "it can be painful to remove them"[1].
The article does not say whether the delay came from the structure of the software, from security approval, from the classified environment, from the training of users, or from some mixture of these. This page does not try to infer the architecture of Maven Smart System from public reporting. It takes the case as a prompt. What should a system that depends on an AI supplier make easy, and what should deliberately remain difficult?
The Central Distinction FoundationalKnowledge that endures for decades — core principles
Replacing an AI provider should be architecturally straightforward and operationally cautious. Application code should not be tightly coupled to one supplier's interface. But a new model should not be treated as equivalent to the old one merely because it accepts a prompt and returns text.
The distinction the rest of the page depends on is this. Technical substitution should be designed in. Operational equivalence has to be demonstrated. A provider change should not require reconstructing the application, and it should not be reduced to changing a model name without testing what the change does.
Starting with the Intuitive Answer Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The reader's first reaction is reasonable: the provider should be a reference to a delegating class, or a value in a configuration file. At the level of application code, that is largely right. A system should depend on an interface that the organisation owns, and provider-specific code should sit behind it.
Operational workflow
|
v
Organisation-owned AI service interface
|
v
Policy, routing and audit layer
|
+---- Provider A adapter
+---- Provider B adapter
+---- Local model adapter
+---- Deterministic fallback
Application components call capabilities the organisation defines, rather than a supplier's SDK:
class IntelligenceAnalysisService:
def analyse_evidence(self, request):
raise NotImplementedError
def extract_entities(self, request):
raise NotImplementedError
def produce_brief(self, request):
raise NotImplementedError
Each supplier then gets an adapter that implements the same interface, and the choice of supplier is made in configuration. The rest of the application does not need to know which supplier is in use.
Why a Configuration Change Is Not Proof of Equivalence Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A configuration change can replace a component syntactically. It does not show that the replacement behaves appropriately. Models can differ in the input forms they accept, their context capacity, the reliability of structured output, their handling of images, their use of tools, their latency and throughput, the deployment environments they can run in, the information-handling agreements that apply to them, their refusal behaviour, their safety controls, their susceptibility to particular adversarial prompts, their support for audit, and their characteristic error patterns.
In a context like the one reported, where a model supports analysis of visual data and the preparation of one-page briefs passed up the chain of command, a substitution needs behavioural and operational validation, not only successful API calls. The configuration value should determine which approved component is loaded. It should not have the authority to declare that component operationally equivalent.
The Layers of Lock-In Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Vendor lock-in" is often used as a single vague label. A more useful account separates the forms of coupling a system can acquire, because each has a different symptom and a different remedy.
| Form of coupling | What it looks like | Design response |
|---|---|---|
| Source code | Provider SDK calls appear throughout the application | Organisation-owned interfaces, dependency injection, adapters, contract tests |
| Protocol | Downstream parsers expect one provider's response structure, tool-call format, or error types | Canonical internal request and response formats |
| Prompt | Prompts are tuned to one model's terminology and caution, and produce different structures after substitution | Provider-specific prompt packages behind a common task contract |
| Capability | The workflow needs image analysis or constrained tool use that an alternative does not support adequately | Declared capabilities, with tasks routed only to compatible components |
| Behaviour | Analysts have learned how one model signals uncertainty, what it omits, and how it refuses | Comparative evaluation, shadow operation, training, documented behavioural profiles |
| Data | Conversation state, indexes, caches, or fine-tuning records exist only in provider-specific forms | Organisation-owned source data, portable representations, migration procedures, explicit retention rules |
| Security and accreditation | Approval applies to one model, supplier, environment, purpose, and data classification | Pre-evaluated alternatives, with accreditation treated as part of the architecture |
| Organisation | Procedures, responsibilities, and professional judgement have adapted to one model | Transition plans, exercises, role analysis, maintained human capability |
Some of these forms are avoidable design defects, such as coupling in source code. Others follow from depending on a component that does context-sensitive, non-deterministic work. Separating the two tells a team what architecture can fix and what it can only manage.
Capability-Based Substitution Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A common interface should not pretend that every model is interchangeable. A signature like the following hides nearly everything that matters:
def generate_text(prompt: str) -> str:
...
A better request states the capability required and the conditions under which it must run:
request = AnalysisRequest(
task="evidence_synthesis",
input_types={"text", "image"},
output_schema="evidence_brief_v3",
provenance_required=True,
maximum_latency_seconds=20,
data_classification="restricted",
human_review_required=True,
)
The routing layer then selects only implementations that meet those requirements:
provider = registry.select(
required_capabilities=request.required_capabilities(),
policy_context=request.policy_context(),
)
If nothing satisfies the requirements, the safe response is to stop, to fall back to a narrower deterministic function, or to pass the work to a person. The system should not silently choose an inadequate model. The principle is to design for substitutable capabilities, not for interchangeable chat endpoints.
Stable and Variable Parts Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The organisation should own the parts of the system that must survive a change of supplier: the task definitions, the business rules, the internal schemas, access controls, provenance records, audit events, evaluation datasets, human-review points, acceptance thresholds, and incident procedures.
The provider adapter should hold the parts that vary: the translation to and from the provider's API, the provider's authentication, prompt formatting, model-specific tool definitions, retry behaviour, provider error handling, response normalisation, and model-specific safety handling. This boundary makes replacement tractable without pretending that the models are the same.
The Model Exit Architecture Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A system that relies on an AI supplier for important work needs an exit architecture, in the same way that a data system needs backup and recovery. The components of such an architecture are these.
- An organisation-owned gateway. Applications call a controlled internal service, not the supplier directly.
- At least one credible alternative. An adapter that has never been exercised is a hope, not a fallback.
- Canonical schemas. Internally meaningful structures should not be defined by a provider's API.
- Portable source data. Primary information, annotations, and audit records should stay accessible whichever supplier is in use.
- Evaluation suites. Representative, difficult, and safety-sensitive cases should be rerunnable against any candidate.
- Recorded provenance. Each output should identify the model, its version, the configuration, the prompt package, the tools, and the evidence used.
- Known degradation modes. The system should state what remains possible when advanced capabilities are unavailable.
- Human fallback procedures. Staff should keep enough knowledge and authority to continue critical work.
- Shadow evaluation. Alternatives can be tested on real traffic without controlling operational decisions.
- Regular exit exercises. Teams should test periodically whether a supplier could actually be removed within the time they have assumed.
Why Replacement Should Not Be Instant Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
A critical model should not be replaced casually. A rapid change of configuration can itself be dangerous if the new component reads instructions differently, omits evidence, produces incompatible classifications, refuses different categories of request, expresses uncertainty in another way, responds poorly to unusual inputs, introduces different vulnerabilities, or changes how effective human review is.
A high-consequence system should therefore make supplier removal bounded and predictable, though not necessarily instantaneous. Good design does not mean that a model can be swapped on a Friday afternoon without scrutiny. It means that the scope of the change is understood, the application does not need reconstruction, comparative tests already exist, and what remains is validation rather than excavation of the architecture.
Testing Behaviour, Not Just Connectivity Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Conventional contract tests check the interface. Each adapter should accept the canonical request, return the required schema, report errors consistently, supply provenance metadata, respect timeouts and cancellation, and never bypass logging or access controls.
Behavioural contract tests check what the model does. They should cover routine cases, difficult boundary cases, deliberately misleading inputs, incomplete evidence, conflicting evidence, adversarial instructions, unsupported requests, tasks that require the model to state uncertainty, and cases where the correct action is to escalate.
Operational contract tests check throughput, latency, recovery from failure, where data is located, logging, access control, version management, incident response, and rollback. A provider can pass the interface contract and still fail the behavioural or operational ones. Each needs its own evidence.
Differential Evaluation and Many Eyes Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
During a migration, the same benchmark cases can be sent to the incumbent model, the proposed replacement, a deterministic baseline where one exists, and human reviewers. The aim is not only to score outputs against a single expected answer. It is to compare the claims each model raises uniquely, the omissions they share, where they disagree, how confident each is without grounds, their refusal patterns, their use of evidence, and their effect on the conclusions people reach.
Several models give the system a kind of cognitive diversity, as the many-eyes discussion in the curriculum material suggests. But agreement between models is not proof, because they may share assumptions and errors. Multi-model operation is therefore useful for evaluation as well as for resilience, provided disagreements are investigated rather than averaged away.
An Activity-Theory Reading Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The reported difficulty is more than a dependency between components. Several activity systems meet in it, and each has its own object.
| Activity system | Object |
|---|---|
| Operational analysis | Produce timely and useful intelligence |
| Software development | Deliver and maintain dependable capabilities |
| Security governance | Control exposure and attack risk |
| Procurement | Obtain capability on acceptable terms |
| Legal and policy oversight | Ensure authorised use |
| Supplier organisation | Provide services under its own policies and obligations |
| Political leadership | Direct institutional policy and permitted suppliers |
The BBC reports that the Pentagon pressed Anthropic to remove its safety guardrails and give the military "unfettered access", and that Anthropic refused, citing concerns about mass surveillance and autonomous weapons[1]. On this reading, the provider's policies were part of the operational relationship, and were not only a feature of the computation it supplied.
Activity Theory helps separate a source-code problem from a procurement problem, a governance problem, a change in rules, a conflict between institutional objects, and a disruption to an established division of labour. Changing an adapter may resolve the technical coupling and leave the other contradictions untouched[activity theory as a design tool].
Practical Exercise: Could We Change the Model? Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Choose an AI-enabled system you know, and work through these questions.
Locate the coupling. Where are provider-specific APIs called? Where are model names stored? Which prompts assume a particular provider? Which response formats leak into application code? Where is conversation state held? Which tools or indexes are provider-specific? Which approval documents name the supplier directly?
State the required capability. What task is the model performing? Which input and output forms are essential? What accuracy, latency, and provenance are required? What happens when the model is uncertain? Which decisions need human approval? Which safety and data-handling requirements must hold?
Attempt substitution. Can an adapter be written without changing application logic? Does the alternative satisfy the same schema? Does it behave adequately on representative cases? What training or workflow changes would be needed? What evidence would justify deployment? How would rollback work?
Examine the organisational effects. Whose work changes? Who must learn the new model's behaviour? Who carries the migration risk? Who can approve the substitution? Which capabilities would be lost? Does the change alter responsibility or accountability?
Design Principles FoundationalKnowledge that endures for decades — core principles
- Applications should depend on capabilities the organisation owns, not on supplier SDKs.
- Provider-specific behaviour should be confined to adapters and prompt packages.
- Internal data and audit schemas should remain independent of any provider.
- Model capabilities and restrictions should be stated explicitly.
- Substitution should fail safely when its requirements cannot be met.
- Behavioural equivalence must be tested, not assumed.
- Alternative providers should be exercised before they are needed.
- Human expertise is part of the fallback architecture.
- High-consequence replacement should be controlled, not frictionless.
- Supplier exit should be treated as a normal lifecycle event, not an exceptional crisis.
Closing FoundationalKnowledge that endures for decades — core principles
An AI model should be removable without requiring the surrounding system to be rebuilt. That is an ordinary expectation of modular software design. But replaceability does not make models interchangeable commodities. Their capabilities, safety behaviour, error patterns, deployment conditions, and relationships with human practice can differ substantially. A responsible architecture therefore combines a simple technical substitution point with a demanding process of evaluation. The provider reference may be a line in a configuration file, but the justification for changing it should not be.
Related Topics
- Designing Auditable, Robust Agentic Systems — verification from outside the component being verified, and checkpoints placed where they matter.
- When Agents Fail: Brittleness, Misaligned Incentives, and Deception — why a system that works inside its tested range can fail outside it.
- From Semantic Search Towards Cognitive Agency: The Accidental Architecture of Memsearch — a local-first tool that still sends some queries to a hosted model for verification.
- Renting or Owning: Cloud Dependence, Retention, and Capability — the same exit question for cloud services in general.
References
- K. Hays and D. Bush, "Pentagon stops using Anthropic AI tools after blacklisting company, BBC told," BBC News. https://www.bbc.co.uk/news/articles/c5j9x9pr0240o