Skip to content

The Model Portability Playbook

Reference architecture, switch runbook, dual-index embedding migration, cloud mapping and a six-week plan for keeping AI work when the model changes.

The Model Portability Playbook

How to change AI models without losing the work: a reference architecture, a switch runbook, embedding migration, where each piece runs on each cloud, and how we roll it out with a team.

This is the technical companion to Switching AI models is easy. Switching memory isn’t. That essay makes the case: routing and gateways decide which model runs the next task, but nothing decides what the new model remembers, so lock-in moves to whoever holds the context. This page is the how.

A reference architecture for switching without forgetting

The pattern works on any major cloud, on Cloudflare’s edge or on your own hardware. Every model stays replaceable because nothing the business depends on lives inside a model or a vendor’s app.

Reference architecture: channels feed an agent runtime; a context layer you own holds briefs, conversations, versioned embeddings and golden sets; a stateless gateway routes to several model providers; an eval gate approves model changes; everything is traced

Figure 1. Context-portable AI architecture. The agent runtime assembles a trimmed context bundle from stores you own, in open formats, and sends it through a stateless gateway. Model changes must pass the eval gate. Every output is traced with the model that actually answered.

1. The context layer

Four stores, all in your own tenancy and in open formats, so neither a platform nor a consultancy becomes the new lock-in.

  • Project brief store. Each long-running piece of work keeps a structured brief: question, sources, findings, decisions, open items. Plain JSON or Markdown in your own database.
  • Conversations and checkpoints. History in your own store, plus a checkpoint summary at each phase boundary, written in plain language any model can read.
  • Knowledge index. Documents, chunks and vectors, with the embedding model pinned and versioned separately from the model that writes answers.
  • Golden sets and eval results. A few hundred real tasks from your own work, with expected outcomes, scored against every model version you run.

2. Govern what you keep

Owning your context means owning its risk. The moment AI history lands in your database it becomes a business record, subject to retention, eDiscovery and privacy law. Design for that on day one:

  • Redact before you store. Tokenize or strip PII and PHI on the way in, not after. Keep the mapping in a separate vault with its own access policy.
  • Retention by class. Different schedules for conversations, briefs and eval data; legal hold that overrides deletion; deletion on request that reaches checkpoints and vector stores too, not only the raw log.
  • Least privilege on context. Row-level access so an agent retrieves only what the calling user may see. Agents get their own scoped, expiring identities, never a shared service key.
  • Segregate regulated data. PHI and similar data stay in a governed boundary, served by models in that boundary, often open-weight models on your own hardware.

Identity has always been the control plane for data. With AI it is also the control plane for memory.

3. Budget the context

“Keep the full history” doesn’t mean “send the full history”. Resending everything on every call is slow, expensive, and dilutes the model’s attention. Send a context bundle instead: the current brief, the latest checkpoint summary, the last few turns, and passages retrieved for this step. Summaries are written at phase boundaries and replace the turns they cover. The same bundle works for any model, which is what makes a switch cheap.

4. The gateway

One API in front of every provider. Route by task, pin exact versions for long-running work, and treat every fallback as a model change that gets logged. Keep model choice, prompts and parameters in versioned config so a switch is a reviewed change, not a code deploy. Pass reasoning blocks back on tool calls within one model, and drop them when the model changes; the new model can’t use them.

5. The eval gate, with a pinned judge

No model change reaches production until it clears the golden set on quality, cost, latency and refusal rate. One detail most teams miss: if a model grades the golden set, that judge is also a model, and it can change too. Pin the judge’s version, change it only on its own schedule, and keep a human-graded sample to check the judge against.

6. Observability and audit

Trace every call with OpenTelemetry and record, per output, the model that actually answered: provider, model, version, preset and prompt version. Read it from the response, not from your request. With Anthropic’s opt-in server-side fallback, for example, a declined request is retried on a different model and the response’s model field says so. Your log should say so too.

The model switch runbook

With the architecture in place, a switch becomes a routine change with a rollback plan. These are the seven steps we rehearse, in order.

Seven-step model switch runbook in three phases: prepare (inventory and pin, checkpoint), prove (shadow run, golden-set eval), switch (canary, cut over, retire and re-index)

Figure 2. The seven-step switch. Switch between phases of work, never mid-thought. When a long project changes model, have the new model restate the brief in its own words before it continues.

When the embedder changes: run two indexes

Re-indexing is the most expensive part of a vendor move and the easiest to get wrong. Build the new index next to the old one, compare them on golden queries, then flip the read pointer.

Embedding migration with two indexes: v1 keeps serving while v2 backfills; golden queries compare recall before flipping the read pointer

Figure 3. Dual-index embedding migration. Index v1 keeps serving while v2 backfills and receives dual writes. Reads flip only when v2 matches or beats v1 on recall for your golden queries.

Where each piece runs

You don’t need a new platform. Every major cloud ships the building blocks, and the open-source versions are mature. Pick the column that matches where your data already lives.

Component AWS Microsoft Azure Google Cloud Cloudflare Open source / on-prem
Model access Amazon Bedrock Microsoft Foundry Models Gemini Enterprise Agent Platform (the evolution of Vertex AI), Model Garden Workers AI vLLM or NVIDIA NIM serving open-weight models on DGX
Gateway & routing Bedrock Intelligent Prompt Routing; LiteLLM API Management AI gateway Agent Gateway AI Gateway (fallbacks, dynamic routing) LiteLLM, OpenRouter
Conversations & memory Bedrock AgentCore Memory; DynamoDB Foundry Agent Service memory; Cosmos DB Memory Bank; AlloyDB Agents SDK on Durable Objects; D1 Postgres
Knowledge index Bedrock Knowledge Bases with S3 Vectors or OpenSearch Azure AI Search AlloyDB AI; BigQuery vector search Vectorize, R2 pgvector, Qdrant
Eval gate Bedrock Evaluations Foundry evaluations Agent Evaluation AI Gateway evaluations promptfoo, Langfuse, Braintrust
Tracing & audit CloudWatch, AgentCore Observability Azure Monitor, Application Insights Agent Observability, Cloud Trace AI Gateway logs, Logpush OpenTelemetry, Langfuse
Identity & policy IAM, AgentCore Identity Entra ID Agent Identity, IAM Access, Zero Trust Keycloak, SPIFFE/SPIRE

Two placement rules from our own work. Keep the context layer next to the data it describes, because residency and access rules follow the data, not the model. And run sensitive workloads on open-weight models on your own hardware, where a model switch is a download, not a contract.

How ready is your team?

Most teams we meet sit at level one or two. They have a gateway or a single vendor and assume that makes them flexible.

Four-level model portability maturity ladder: level 1 locked in, level 2 plumbed, level 3 context-owned, level 4 switch-ready

Figure 4. Model portability maturity. Level 2 feels flexible because the API is swappable. Only levels 3 and 4 keep the work itself when the model changes.

Six weeks to your first switch-ready workload

We start with one production workload, chosen with you, and take it to level 3 or 4 with a real model switch and a tested rollback. The architecture, runbooks and eval gate then carry to the next workloads, which your team can run with or without us.

Six-week Periscope engagement plan: assess in weeks 1–2, architect in weeks 2–3, prove in weeks 3–5, enable in weeks 5–6, with the first live model switch at the end of week 5

Figure 5. Engagement plan for the first workload. Phases overlap so the eval gate is ready before the switch drill.

Weeks 1–2: Assess

  • Inventory AI workloads: model, version, embedder, prompts, where history lives.
  • Score each against the survival matrix and pick the first workload.
  • Map residency, retention, identity and audit requirements.

Output: Portability scorecard and risk register.

Weeks 2–3: Architect

  • Design the context layer on your current cloud, in open formats.
  • Configure the gateway: routing, pinning, fallbacks, presets.
  • Define the model-of-record schema, retention classes and tracing.

Output: Target architecture and prioritized backlog.

Weeks 3–5: Prove

  • Build the golden set for the first workload with your subject experts.
  • Stand up the eval gate with a pinned judge, and dashboards.
  • Run one live switch using the seven-step runbook.

Output: One workload switched in production, rollback tested.

Weeks 5–6: Enable

  • Role-based workshops (below).
  • Runbooks, brief templates and checkpoint prompts.
  • Roadmap and owners for the next workloads; quarterly switch drills.

Output: A team that can repeat it without us.

Workshops we run with your team

  • Executive briefing (90 minutes). The economics of switching, where lock-in now lives, and the questions to ask every AI vendor.
  • Product and delivery leads (half day). Writing golden sets, reading eval results, canary decisions, and briefs as a project habit.
  • Engineering bootcamp (two days). Hands-on: gateway config, context stores and budgets, dual-index migration, tracing and the switch runbook.
  • Everyone, lunch and learn (45 minutes). Using AI tools safely, keeping your own notes, and switching tools without starting over.

How we measure it

  • Switch lead time. Days from “we want model X” to X live in production, with the eval gate passed.
  • Pinned workloads. Share of production workloads on exact model versions, not floating aliases.
  • Golden-set coverage. Share of workloads with a golden set, and a pinned judge, run on every model change.
  • Owned context. Share of conversations and projects whose state lives in your stores, not only in a vendor app.
  • Model-of-record coverage. Share of outputs traceable to the provider, model, version and prompt that actually answered.
  • Re-index time. Hours to rebuild and validate the knowledge index on a new embedder.

Pick your first workload

Book 30 minutes with a Periscope engineer. We’ll score one AI workload against the survival matrix and sketch its path to switch-ready. Book 30 minutes.

Sources

  1. Anthropic, Refusals and fallback.
  2. OpenRouter, Reasoning tokens; Presets; Model fallbacks.
  3. Google Cloud, Introducing Gemini Enterprise Agent Platform, April 22, 2026; Memory Bank.
  4. Microsoft, Microsoft Foundry Models overview; Memory in Foundry Agent Service.
  5. AWS, Intelligent prompt routing in Amazon Bedrock; AgentCore Memory; S3 Vectors with Bedrock Knowledge Bases.
  6. Cloudflare, AI Gateway and fallbacks.

💬Discussion & Notes

Comments powered by Garrul.