Skip to content

Switching AI Models Is Easy. Switching Memory Isn’t.

Routing picks the next model; it doesn't carry memory. Why lock-in moves to whoever holds your context, and what to do about it by role.

Switching AI Models Is Easy. Switching Memory Isn’t.

Prices and rankings now change in a day. Routing and gateways decide which model runs the next task. Nothing decides what the new model remembers, and that is where lock-in now lives.

Key takeaways

  • Routing picks the model for the next task. It doesn’t carry memory, unwritten reasoning or tuned prompts from the old model to the new one.
  • Cheaper models don’t end lock-in. They move it from the model vendor to whoever holds your context.
  • Gateways such as OpenRouter fix the plumbing and are stateless by design, which is the right call. Memory, evaluation and an audit trail of which model answered are yours to own.

On September 22, Anthropic released Claude Opus 5.5 at 20% below the previous Opus price. About 90 minutes later, OpenAI released GPT-6 Sol and Luna at half the per-token price of the GPT-5.6 models they replace. Enterprise platforms are pitching auto-routing, where each task goes to the model that fits its quality, speed and cost. Gateways put hundreds of models behind one API key. Switching models has never been easier.

I agree with the direction. Model choice is now part of the architecture, not a once-a-year procurement decision. But the conversation leaves out half the problem.

Routing decides which model handles the next task. Nothing decides what the new model remembers.

And the price war hides a second point. When models become interchangeable, the thing that is hard to move is no longer the model. It is the context: the history, the documents, the decisions and the tests that make AI useful on your work. Lock-in doesn’t disappear when models get cheaper. It moves to whoever holds the context. The question every enterprise should be asking is who that is.

Four people, one problem

The problem looks different depending on where you sit. These are composites of patterns we see across teams, not any single client.

The developer

A provider has an outage and a fallback kicks in. Same API, different model. The JSON still parses, but the agent’s judgment has shifted, and nothing in the logs says a switch happened.

What broke: a silent model change
The product owner

A “latest” model alias upgrades itself overnight. The support bot now writes longer answers and escalates less. Better or worse? Without a test set of real tickets, nobody can say.

What broke: no way to measure
The executive

Finance asks to move to the cheaper vendor. Swapping the chat model is a config change. The document search was built on the old vendor’s embedding model, so years of contracts need re-indexing.

What broke: a hidden migration cost
The manager

A six-week market study lives in one analyst’s chat history. The analyst moves on, the license lapses, and the reasoning behind the recommendation goes with them.

What broke: knowledge walks out

The same thing happens at home: switch chatbots mid-project and the new one greets you like a stranger.

What actually survives a switch

“Switching models” covers three different moves: a new version from the same vendor, a new vendor called directly, or a new vendor reached through a gateway. Each carries some things over and leaves others behind. This matrix is the checklist we start every engagement with.

Matrix of eight things that do or do not survive a model switch, by switch type (new version same vendor, new vendor direct, new vendor via a gateway), and what fixes each

Figure 1. Eight things that move, or don’t, when you change models. A gateway improves the plumbing rows (tools, billing, audit). The memory rows stay red or amber until you own them.

Conversation and project memory. History, project memory and files live inside each vendor’s app. Consumer assistants have started to offer memory import; Claude, for example, accepts memories exported from other assistants as pasted text. That helps, and it proves the point: what moves is a one-page summary, which is exactly what we tell individuals to write before they switch. Nothing like it exists for an enterprise’s thousands of conversations and project workspaces.

Reasoning that was never written down. Modern models reason before they answer, and some vendors return that reasoning encrypted, readable only by the model that produced it. A gateway can carry the format across. It can’t carry the meaning. Any conclusion that stayed in the model’s head is gone.

Prompt behavior. Prompts are tuned, often by trial and error, to one model’s habits. The next model follows instructions more or less literally, formats differently and refuses different things. Nothing errors. Quality just moves.

Embeddings and search indexes. If your AI searches your own documents, they were turned into vectors by one embedding model, and vectors from different embedding models aren’t compatible. Change the embedder and you rebuild the index.

Tool calls. Most vendors now accept similar tool schemas, so a direct switch means rework rather than a rewrite. Gateways normalize most of the rest.

Audit trail. Switches don’t always come from you. Anthropic’s Opus 5.5 declines some requests in sensitive categories such as cybersecurity. By default the API returns a refusal. With the opt-in server-side fallback, Anthropic retries on the model it recommends for that category, and the response’s model field reports which model actually answered. That is well designed, and it only helps if you log it. An agent that ignores the field can run a mix of models from one step to the next without anyone knowing.

Prompt cache. Providers cache long, repeated context to cut cost and latency. The cache belongs to one model, so the first days after a switch cost more than the price sheet suggests.

What gateways fix, and what they don’t

To test the thesis we looked closely at OpenRouter, one of the most widely used model gateways. It is a good product, and it is clear about where its job ends.

Need Status What OpenRouter provides (per its docs, Sept 2026)
Many vendors, one API Solved Hundreds of models behind one OpenAI-compatible endpoint, one key, one bill.
Resilience Solved Provider failover plus a models fallback list; failed requests aren’t billed.
Config without redeploys Solved Presets store model choice, fallbacks, system prompt and parameters.
Governance Strong SSO/SCIM, SOC 2, zero-data-retention routing, US and EU in-region processing, per-key budgets.
Tools & structured output Mostly Normalized tool calling, an Agent SDK and an MCP server.
Reasoning continuity Format only Reasoning blocks are normalized and re-emitted in each provider’s format. Encrypted content still means something only to the model that wrote it.
Cache Partly Sticky routing keeps a conversation on the same provider, per model. A model switch starts cold.
Memory By design, no The Responses API is stateless and rejects stored conversations. Chat history in its web app lives in your browser only.
Choosing the model Market signal The Auto Router classifies each prompt into one of about 30 task types, then ranks models by what the OpenRouter community “actually spends on over a trailing 7-day window” for that type. A useful signal, but it is the market’s preference, not your quality bar.

That is the right design for a gateway. OpenRouter holds none of your context, so you can leave as easily as you arrived. It also draws a clean line: everything below it is plumbing you can rent, and everything above it is memory you have to own.

Layered stack: apps on top, a context layer you own (project briefs, conversations, embeddings, golden sets, model-of-record), a stateless model gateway, and model providers at the bottom

Figure 2. The line between plumbing and memory. Gateways and models are rentable and swappable. The context layer is where lock-in now lives, so it should belong to you.

Platforms that route across models for you, from enterprise search to agent platforms, often offer to hold that context too, because context is what makes routing smart. That can be a fair trade. Make it knowingly, and ask one question before you sign: if we leave, what do we take with us?

The same test applies to consultancies, including us. A context layer is only yours if it runs in your own tenancy, in open formats: your database, plain JSON briefs, standard vector stores, OpenTelemetry traces. That is how we build it, so the team can keep running it without us.

What to do, by role

Developers

“Can I swap the model without breaking the agent?”

  • Keep conversation state in your own store and send a trimmed context bundle each call, not the full history.
  • Pin exact model versions for long-running work and log every fallback, with the model that answered.
  • Pass reasoning blocks back on tool calls within one model, and drop them when the model changes.

Product owners

“Is the new model better for my users?”

  • Own a golden set of real tasks. No model swap ships without passing it.
  • If a model grades the golden set, pin that judge model too.
  • Canary every switch to 5–10% of traffic first.

Business executives

“What does switching really cost, and who holds our context?”

  • Budget for re-testing and re-indexing, not just token price.
  • Require that the context layer lives in your tenancy, in open formats, whatever you buy.
  • Treat stored AI history as a business record, with retention and legal hold.

Managers

“Does our work survive a departure or a license change?”

  • Every AI-assisted project keeps a brief outside the chat: question, sources, findings, decisions, open items.
  • Make “write the checkpoint” part of the definition of done.

For everyone else: before you switch tools, ask the current one to summarize everything you decided on one page, save it somewhere you control, and paste it into the new one.

Technical companion: The Model Portability Playbook covers the reference architecture, a seven-step switch runbook, dual-index embedding migration, a cloud-by-cloud mapping and a maturity model.

Who owns the memory of the work?

The price of intelligence will keep falling, and the best model for a given task will keep changing. That is good news, as long as changing models doesn’t mean losing the work. Routing and gateways give you the freedom to switch. A context layer you own, an eval gate and a record of which model answered make switching safe.

Model portability without context portability is half an architecture.

Walk through Figure 1 on one of your workloads

Book 30 minutes with a Periscope engineer. We’ll score one AI workload against the matrix and show you which rows are red.

Book 30 minutes

Sources

  1. SiliconANGLE, Anthropic releases Claude Opus 5.5 and OpenAI counters with two cheaper GPT-6 models, September 22, 2026.
  2. The Next Web, OpenAI cuts GPT-6 prices in half with Sol and Luna, September 2026.
  3. Anthropic, Refusals and fallback (default refusal; opt-in server-side fallback; model field reports the answering model).
  4. Anthropic, Import and export your memory from Claude.
  5. OpenRouter, Auto Router; Responses API; Reasoning tokens; Prompt caching; Presets; FAQ; Enterprise.

💬Discussion & Notes

Comments powered by Garrul.