What Is Actually Portable in the AI Stack: An Exit-Cost Read for Buyers
The AI lock-in conversation is mostly conducted at the wrong altitude. Buyers ask “can we switch model providers?” and the answer, delivered at conference panels, is “yes, the APIs are mostly compatible.” That answer is technically true and practically misleading. It conflates the portability of a chat completion call with the portability of a working system that took eighteen months to build.
This piece separates what genuinely moves between vendors from what does not, and prices the exit accordingly.
The portability table
| Artefact | Portable? | Practical exit cost |
|---|---|---|
| Prompts (text) | Yes — text is text | Low |
| Prompt templates wired to tools | Partial — schema and tool definitions differ | Medium |
| Few-shot examples and system context | Yes, but re-tuning is needed | Medium |
| Evals (test sets) | Mostly — data moves; scoring needs porting | Medium |
| Embeddings | No — re-embed everything on swap | High |
| Vector indexes | Only the document store; index config is vendor-specific | Medium |
| Fine-tunes | No — re-train on the destination | High |
| Agent traces and run history | Partial — schemas differ | Medium to high |
| Guardrail configurations | No — each platform’s policy DSL is its own | High |
The chat-call compatibility layer sits at the top of this table and covers roughly the first two rows. Everything below that is where the exit cost actually lives.
Prompts: the artefact that lies about itself
A prompt is a text file. It copies cleanly to a new provider. It also silently assumes many things about the model family it was tuned against: how aggressively the model follows format instructions, how it handles long context, whether it treats JSON output as a preference or a contract. The prompt that produced stable structured output on one provider’s flagship can produce prose on another’s, without any error being raised.
Treat prompts as portable but not transferable. You will re-write thirty to fifty per cent of them on any provider swap, and you cannot predict which ones until you test. That unpredictability is part of the cost.
Evals: the partial salvation
A well-built eval suite is the only artefact that actually de-risks a model swap. If your test data, scoring rubric, and failure thresholds exist outside the vendor’s platform, you can run the candidate provider against the suite and read the regression before you commit. Buyers who skipped eval tooling because “the vendor handles it” are the ones who discover too late that their exit cost is higher than their remaining contract value.
Insist on eval artefacts that you own and that can be scored independently of any single vendor’s tooling: raw inputs, expected outputs, rubric criteria, and per-item metadata. Whatever platform-specific scoring syntax exists today is a convenience you should be able to drop.
Embeddings and fine-tunes: the hard walls
Embedding space does not transfer. If your retrieval layer embeds documents with one vendor’s embedding model, a swap requires re-embedding the corpus. The corpus copy is easy; the re-indexing window is not, and at any scale above a toy dataset it is measured in days plus real compute spend.
Fine-tunes do not transfer at all. A LoRA adapter trained against one provider’s base model is tied to that base. The destination provider will not load it, and re-training requires the original dataset, the training recipe, or both. If either is missing — and in vendor-hosted fine-tuning, often neither is handed to you — the investment has been consumed by the vendor’s platform.
Run history and audit trails
The trace data your agents have been producing is valuable for debugging, compliance review, and training future systems. Traces are exportable in principle. In practice, each vendor’s trace schema carries its own fields for model identity, tool-call framing, and cost accounting. Plan on a transformation pipeline on exit, or accept that the traces become read-only archives that no new vendor can ingest.
The exit-cost formula
Add up the pieces before you sign, not before you leave:
- Prompt re-tuning — engineering weeks, unpredictable
- Eval suite porting — mechanical but tedious
- Embedding regeneration — compute + re-index window
- Fine-tune retraining — biggest single line item if applicable
- Guardrail and policy reconfiguration — policy-DSL translation work
- Trace archival and transformation — one-time but unavoidable
For a mid-sized agentic deployment with fine-tunes and a million-row retrieval layer, the honest exit cost is a material fraction of a year’s contract value. Quote it internally that way. The vendors know this number and price their discounts to keep you from running it.
Practical mitigations
Three patterns measurably lower the cost later.
- Keep the eval suite vendor-neutral from day one. The single highest-impact decision.
- Separate prompts from runtime configuration. Prompts in source control, not only in the vendor’s UI.
- Hold the raw dataset used for any fine-tune. Without it, the fine-tune is a rental.
The providers are not malicious. They are simply charging you for sticking around. Your job is to make sure you can, on paper at least, leave.