The one-screen answer: As reported, the 2026-07-28 MCP revision removes the protocol-level session — no Mcp-Session-Id header, no initialize handshake — so any request can hit any instance. That means you run your MCP server behind a plain round-robin load balancer with no sticky sessions, no session affinity, and no shared in-memory session store. You can delete the ip_hash and sticky-cookie directives a stateful MCP server used to need, add a dumb /health probe, and let a horizontal autoscaler treat every instance as interchangeable. The one catch: long-running work moved to the Tasks extension, and because any poll can land on any instance, your task state must move to a shared durable store (Redis/Postgres/object store) — the transport is stateless, but your task state is not. This is the deploy-shape how-to; for the SEP-by-SEP details see the migration checklist and what breaks and why.

Before / after: the topology actually changes#

The old shape forced every client back to the same process, because the session lived in that process's memory. The new shape does not.

BEFORE (2025-11-25, stateful)
client ──> LB (ip_hash / sticky cookie) ──> instance A  [session map in RAM]
                                       \──> instance B  [session map in RAM]
                       shared session store (Redis, keyed by Mcp-Session-Id)

AFTER (2026-07-28, stateless)
client ──> LB (round-robin) ──┬──> instance A  (no session state)
                              ├──> instance B  (no session state)
                              └──> instance C  (no session state)
                       shared TASK store (Redis/Postgres, keyed by task id)
                       ^ only long-running work touches this

The session store and the affinity rule both disappear. A new box appears — but only tasks touch it, and only when you run long work.

What it means: the migration is mostly deletion. You remove routing complexity; you add one store, and only if you use Tasks.

Reverse proxy: round-robin, and delete the sticky bits#

nginx round-robins by default. The entire config is an upstream block and a proxy_pass. Note what is gone — the commented lines are the directives a stateful MCP server previously required.

upstream mcp_backend {
    # ip_hash;                       # DELETE: no session to pin
    # sticky cookie mcp_sid;         # DELETE: no Mcp-Session-Id to route on
    server mcp-1.internal:8080;
    server mcp-2.internal:8080;
    server mcp-3.internal:8080;
    keepalive 64;
}

server {
    listen 443 ssl;
    location /mcp {
        proxy_pass http://mcp_backend;   # plain round-robin
        proxy_http_version 1.1;
        proxy_set_header Connection "";
    }
    location /health {
        proxy_pass http://mcp_backend;
        access_log off;
    }
}

If your nginx config still has an ip_hash line after July 28, you are paying for statefulness you no longer have.

What it means: the LB config gets shorter, cheaper, and boring — which is exactly what you want infrastructure to be.

Health checks and autoscaling: instances are now interchangeable#

Because no request is pinned, a liveness probe only has to answer "is this process up?" — not "does this process own your session?" A trivial handler is enough.

app.get("/health", (_req, res) => res.status(200).send("ok"));

That plugs straight into a Kubernetes HPA (or any scale-on-CPU controller):

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: { name: mcp-server }
spec:
  scaleTargetRef: { apiVersion: apps/v1, kind: Deployment, name: mcp-server }
  minReplicas: 0            # stateless servers can scale to zero
  maxReplicas: 20
  metrics:
    - type: Resource
      resource: { name: cpu, target: { type: Utilization, averageUtilization: 70 } }

Scaling out adds capacity with no rebalancing; scaling to zero strands nothing. If you are still choosing where these instances run, the agent-sandbox comparison covers Cloud Run vs E2B vs Modal vs Fly.

What it means: you get real horizontal scale and scale-to-zero for free, because the protocol stopped forcing affinity.

Externalize Tasks state — the one place statelessness bites#

Long-running work moved to the Tasks extension (SEP-2663), as reported: a tools/call returns a task handle, the client polls tasks/get, supplies input via tasks/update, and cancels via tasks/cancel. States run working → input_required → completed / failed / cancelled. Note that tasks/list was removed — it could not be scoped safely without a session.

Here is the whole rule in code: write on create, read on poll, always from the shared store — never process memory.

// tools/call (async) — persist the record so ANY instance can answer later
async function onTasksCall(req) {
  const taskId = req.params.taskId ?? crypto.randomUUID();  // idempotent
  await redis.set(`task:${taskId}`,
    JSON.stringify({ state: "working", input: req.params, result: null }),
    { NX: true });                     // NX = safe to retry the same call
  enqueueWorker(taskId);               // worker also reads/writes the store
  return { taskId, state: "working" };
}

// tasks/get — read from the SHARED store, because the poll can land anywhere
async function onTasksGet({ taskId }) {
  const rec = await redis.get(`task:${taskId}`);
  if (!rec) return { error: "unknown task" };
  return JSON.parse(rec);              // any instance answers identically
}

The call that created the task on instance A will, under round-robin, be polled on instance B or C. That only works if B and C read the same Redis (or Postgres) record. Keep the async tools/call idempotent so a client retry after a dropped connection does not spawn a duplicate job.

What it means: budget for exactly one new dependency — a durable task store — and nothing else. The transport went stateless so you would not need a session store; Tasks is the one workload that still needs shared state.

Migration safety: you do not have to flip everything Tuesday#

The stateless core is a release candidate as of 2026-07-27 and finalizes Tuesday, July 28, 2026. Nothing breaks immediately: as reported, a 2026-07-28 client falls back to the old initialize handshake when it meets an older server, and deprecated features keep working for at least a ~12-month window. Migration is on your schedule.

Because instances are interchangeable, you roll them one at a time behind the LB — no cutover event. Pull a node, deploy the stateless build, health-check it, return it to the pool, repeat. The week's roundup tracks which clients and SDKs have already shipped support.

What it means: the deadline is a floor for capability, not a wall. Deploy when your fleet is ready, not when the clock strikes.

Gotchas#

For the full code-level diff of headers, handshake, and _meta, the migration checklist is the companion to this deploy guide.