High availability

Several copies of a server behind a load balancer keep serving when one of them stops. A client on MCP 2026-07-28 doesn't care which copy answers: every request carries what it needs, so the copies only have to share their secrets. A client on an older protocol version has a session, and the copy that didn't create it needs to find it: that's what a shared Redis is for. In the default mode, any copy loads the session from Redis. In distributed mode, the copy that made a session keeps serving it, the others relay its requests to that copy through Redis, and when that copy stops, another takes its sessions over. Running Several Instances teaches the default mode step by step.

@FrontMcp({ info, apps, redis: { url: process.env.REDIS_URL } })  // the same on every instance
MCP_SESSION_SECRET=… FRONTMCP_BIND_ADDRESS=all node dist/node/help-desk.bundle.js   # the same secret everywhere

Reference

What instances share

WhatEvery instance needsWithout it
MCP_SESSION_SECRETThe same value, for clients before MCP 2026-07-28In production, those clients' initialize answers 500. An id made under another secret gets 404 invalid session id. See secrets.
JWT_SECRET, VAULT_SECRETThe same values, with local or remote auth, or tools that ask the user more than onceTokens and answers from one instance fail on another. A production server with redis and neither secret warns at startup: requestState is signed with a per-process key. See running several instances.
redisThe same RedisSessions of older clients, background tasks and pending questions stay in the instance that created them.
auth.tokenStorage, throttle.storageRedis too, if you use local/remote auth or rate limitsSign-ins that change instance fail, and each instance counts its own limits.
The same buildTools, resources and promptsA client sees a different server depending on where it lands. Compare toolsHash in /readyz: see checking replicas.

Clients on MCP 2026-07-28 open no sessions, so for them redis only matters when a tool starts a background task or asks a question.

How a session moves between instances

With redis set, a session is written to Redis when a client opens it, with the id of the instance that holds it (its nodeId). When a request of that session reaches an instance that doesn't have it in memory, that instance reads it from Redis, logs Recreating transport from stored session, serves the request, and writes it back. So a load balancer can send an older client's requests anywhere, without affinity, and a restarted or replaced instance picks up where the old one left off. It does that only for an id it can decrypt, so every instance needs the same MCP_SESSION_SECRET (caveats).

A session's key lives an hour, or redis.defaultTtlMs, and the instance that holds the session pushes it out while it serves the session's requests (how long sessions last). An instance that can't reach Redis when it starts keeps sessions in memory, retries the connection, and uses Redis once it answers, with no restart. Its /readyz answers 503 until then, so a load balancer that checks readiness sends it no traffic (when Redis can't be reached).

Distributed mode

FRONTMCP_DEPLOYMENT_MODE=distributed, which frontmcp build --target distributed sets in its entry (node dist/distributed/index.js), puts the server in distributed mode:

  • It listens on every interface, unless http.security.bindAddress or FRONTMCP_BIND_ADDRESS says otherwise. The audit says BIND_ALL_INTERFACES_DISTRIBUTED.
  • /healthz reports "deployment": "distributed".
  • Every response carries the id of the instance that served it, X-FrontMCP-Machine-Id: <id>, for load balancers that route by it: MCP answers of every protocol version, /healthz and 404s. An older client's initialize also sets a cookie, __frontmcp_node=<id>; Path=/; Max-Age=86400; HttpOnly; SameSite=Strict. (In 1.8.7 only the responses to older clients had the header.)
  • With redis, it starts the HA manager ([HA] Manager started for distributed deployment): a heartbeat key for each instance (mcp:ha:heartbeat:<id>), refreshed every 10 seconds and expiring after 30; a scanner for the sessions of stopped instances, every 10 seconds; a bus that records which instance holds each session (mcp:bus:session:<session id>); and the relay, [HA] Request relay ready — sessions owned by other nodes are served through them. It also stores SSE events in Redis, under mcp:events:.

How a session's requests are served

The request reachesWhat happens
The instance that holds the sessionIt serves it.
Another instance, while the holder runsIt relays the request to the holder through Redis, and streams the answer back. The holder runs it, and the answer's X-FrontMCP-Machine-Id names the holder.
Another instance, after the holder shut down (SIGTERM or SIGINT)The instance takes the session over from Redis at once, serves it, and from then on holds it.
Another instance, after the holder died without shutting down (a crash, SIGKILL, a lost machine)Until the holder's heartbeat expires: 503, with Retry-After set to the heartbeat's lifetime (30) and {"code":-32000,"message":"This MCP session is served by another instance of the server, which did not answer the relayed request. Retry after 30s; if that instance has stopped, the session is taken over by another instance once its heartbeat expires."}. After that, it takes the session over as above.

A server that shuts down deletes its heartbeat key (stopping). When instance desk-a was stopped with SIGTERM, it exited 0, only desk-b's heartbeat was left in Redis, and the session's next request, sent to desk-b a tenth of a second later, was answered 200, with [HA] Session owner is gone — the session will be taken over, Recreating transport from stored session and [HA] Took over session from a stopped node { …, previousNodeId: 'desk-a' } in its log. Before 1.9.2 the heartbeat stayed until it expired, so after a SIGTERM the session answered 503 for 30 seconds. An instance that crashes still leaves its heartbeat behind: after desk-a died, desk-b answered 503 with Retry-After: 30.

A load balancer doesn't need to route by the header or the cookie: round-robin works, and routing by them saves the relay through Redis. Every instance needs the same MCP_SESSION_SECRET, as in the default mode. Relayed requests pass through Redis with their headers, so keep Redis on a private network. Distributed mode without redis works on one instance: it only adds the header and the cookie.

deployments[].ha in frontmcp.config tunes the HA manager: heartbeatIntervalMs, heartbeatTtlMs, takeoverGracePeriodMs and redisKeyPrefix. The distributed build writes them into serverless-setup.js as FRONTMCP_HA_* variables, which a variable already set wins over. With ha: { heartbeatIntervalMs: 2000, heartbeatTtlMs: 6000 }, the 503 said Retry after 6s and the takeover came six seconds after the holder stopped. deployments[].server.cookies sets the cookie's name, Domain and SameSite: cookies: { affinity: "desk_node", sameSite: "Lax" } gave desk_node=<id>; …; SameSite=Lax.

Machine ids

Each instance has an id. It's the nodeId of the sessions it creates, and the value of the header and cookie in distributed mode.

SourceWhen
setMachineIdOverride(id)Whenever your code has called it, over MACHINE_ID too. See below.
MACHINE_IDWhenever it's set.
HOSTNAME, else the operating system's host nameIn distributed mode. Kubernetes sets HOSTNAME to the pod's name.
A random id per processIn serverless mode.
A random id, new on every startOtherwise. FrontMCP's source mentions .frontmcp/machine-id, in the working folder outside production, and a file that MACHINE_ID_PATH names; starting the node build in its project folder wrote no such file.

Outside production, FrontMCP encrypts session ids with a key made from the machine id when MCP_SESSION_SECRET isn't set, so without MACHINE_ID the key changes on every start: after a restart, an old session id gets 404 invalid session id, Redis or not. Production requires the secret for sessions, so there the machine id only names the instance: sharing MACHINE_ID between instances doesn't share anything else.

Your code reads the id with getMachineId() from @frontmcp/utils, the function FrontMCP itself calls, and replaces it for the whole process with setMachineIdOverride(id); setMachineIdOverride(undefined) goes back to the sources above. Call it before the server starts. The header follows a later call, but the heartbeat doesn't: in distributed mode with redis, an instance started as help-desk-7b8f9, then given desk-a, kept mcp:ha:heartbeat:help-desk-7b8f9 while its responses said X-FrontMCP-Machine-Id: desk-a.

main.ts
import "reflect-metadata";
import { getMachineId, setMachineIdOverride } from "@frontmcp/utils";

setMachineIdOverride(`desk-${process.env.REGION}-${process.pid}`); // before @FrontMcp starts the server

// later, anywhere in the process
console.log(getMachineId()); // desk-eu-west-1-4182

@frontmcp/utils is installed with @frontmcp/sdk; add it to your own dependencies, at the same version, to import it. create({ machineId }) sets the same override for one in-process server, but only in an ES module project: in a CommonJS one, FrontMCP 1.9.4 sets it on a copy of @frontmcp/utils that the server never reads, and the id stays as it was. Call setMachineIdOverride() there instead. This was checked with FrontMCP 1.9.4 on Node: create({ machineId }) in both kinds of project, and the override against distributed mode's header and Redis heartbeat.

Caveats

  • In the default mode, sessions are shared, streams aren't: a client that holds an event stream open, like the older SSE transport, holds it with one instance. Give such clients affinity at the load balancer.
  • An instance serves a session id only when it decrypts under its own MCP_SESSION_SECRET. An id made under another secret gets 404 invalid session id, and the log says mcp-session-id is not a session this server verified: give every instance the same secret. After you rotate it, a client with an old id gets that 404 once and starts a new session. (In 1.8.5, a public server served such an id from Redis.)
  • In the default mode, a DELETE of a session ends it on the instance that receives it and deletes it from Redis; a later request there gets 404 Session not found. Another instance that has the session in memory, because it created the session or served it before, checks with Redis (EXISTS) on each of the session's requests that the session is still stored, and drops it when it isn't: its next request got 404 {"code":-32001,"message":"session expired"}, and its log said The stored session is gone — dropping its transport here. A session whose key expired gets the same 404, on the instance that holds it too (how long sessions last). (Changed in 1.9.3: in 1.9.2 such an instance kept serving a deleted or expired session, and answered its next request 200.) In distributed mode the DELETE is relayed to the holder, which answers 204 with its own X-FrontMCP-Machine-Id and deletes the stored session, and the session is gone everywhere.
  • That check doesn't make a session depend on Redis. When Redis refuses connections, or doesn't answer within transport.persistence.sessionCheckTimeoutMs (500 ms by default), the instance serves the session from memory, logs [TransportService] Could not confirm the session is still stored — serving it here, and checks again on the session's next request. With the Redis container paused, which keeps the connection open and answers nothing, each call on a session the instance held answered 200 after half a second, with The session store did not answer within 500 ms in the log; with sessionCheckTimeoutMs: 2000, after two seconds. A request for a session the instance doesn't hold has no copy to fall back on, and waits for Redis: it got no answer while Redis was paused, and was answered once it came back. (Changed in 1.9.4: in 1.9.3 a call on a session the instance held waited for a paused Redis too, and got no answer in three minutes.) See when Redis can't be reached.
  • This was checked with FrontMCP 1.9.4: two instances of a Node build on one machine sharing Valkey 8 in Docker, with a DELETE on one, and Redis stopped and paused, and two of the distributed build sharing it: relayed calls, a relayed DELETE and a holder stopped with SIGTERM. The crash was seen with 1.9.2. Two Node instances behind nginx in Docker Compose, and a shorter heartbeat from deployments[].ha, were last run with 1.9.1. No Kubernetes was used, and no load balancer routed by the cookie or the header.
  • A question to the user waits on the instance that asked. For clients before 2026-07-28, a tool's this.elicit() waits on the instance it runs on, and the answer reaches it from another instance through Redis publish and subscribe, encrypted when MCP_ELICITATION_SECRET (or MCP_SESSION_SECRET) is set, so every instance needs the same secret: see Several instances and encryption.

Usage

Running two instances that share sessions

The server has redis set from the environment, as in Redis. Start two instances with the same secret, on two ports:

export MCP_SESSION_SECRET="$(openssl rand -hex 32)"
MACHINE_ID=desk-a PORT=4001 REDIS_URL=redis://127.0.0.1:6379 node dist/node/help-desk.bundle.js &
MACHINE_ID=desk-b PORT=4002 REDIS_URL=redis://127.0.0.1:6379 node dist/node/help-desk.bundle.js &

A client on MCP 2025-06-18 sends initialize to port 4001, and gets an Mcp-Session-Id. It sends its next tools/call, with that id, to port 4002, which has never seen it: the call succeeds, and the second instance logs:

[TransportService] Recreating transport from stored session
[StreamableHttpAdapter] Marked transport as pre-initialized for session recreation

Without redis, the second instance answers 404 session not initialized, and the client has to initialize again. In a container, add FRONTMCP_BIND_ADDRESS=all so the load balancer can reach each instance.

Checking that replicas serve the same tools

/readyz has a hash of the tool names, the same on every instance built from the same code:

curl -s http://127.0.0.1:4001/readyz | jq -r .catalog.toolsHash
curl -s http://127.0.0.1:4002/readyz | jq -r .catalog.toolsHash
<the same 64 hex characters>
<the same 64 hex characters>

During a rolling deploy, two hashes mean two versions are serving. It covers tool names only, not their schemas.

Routing by instance, in distributed mode without Redis

A load balancer that pins a client to an instance can use the cookie distributed mode sets on initialize:

FRONTMCP_DEPLOYMENT_MODE=distributed HOSTNAME=help-desk-7b8f9-abc12 node dist/node/help-desk.bundle.js
X-FrontMCP-Machine-Id: help-desk-7b8f9-abc12
Set-Cookie: __frontmcp_node=help-desk-7b8f9-abc12; Path=/; Max-Age=86400; HttpOnly; SameSite=Strict

Without redis, sessions stay in memory, so this only helps clients that return to the same instance, and one that stops takes its sessions with it. With redis, any instance serves a session, by relaying it to the instance that holds it (above), and the cookie saves that relay.


Troubleshooting

A request answers 500, and the log says RemoteTransporter.handleRequest() is not implemented

Distributed mode with redis before 1.9, and the request's Mcp-Session-Id belongs to a session another instance created. Upgrade: since 1.9 the request is relayed to that instance. See distributed mode.

503, This MCP session is served by another instance of the server, which did not answer the relayed request

Distributed mode: the instance that holds the session didn't answer the relay, most often because it died without shutting down, and its heartbeat hasn't expired yet. Retry after Retry-After seconds; once the heartbeat expires, the instance that gets the request takes the session over. The log says Relay to session owner failed. If the holder is running, check that it reaches the same Redis. See how a session's requests are served.

TypeError: socket.destroySoon is not a function, and the instance exits

The instance served a request relayed from another instance, in distributed mode with redis, on FrontMCP 1.9.1 or 1.9.2. Upgrade to 1.9.3, where it keeps running. On those versions, route each session's requests to the instance that holds it, by the __frontmcp_node cookie or the X-FrontMCP-Machine-Id header, so that nothing is relayed. See how a session's requests are served.

[HA] Failed to start HA manager — running without HA, this.redis.hset is not a function or [HA] Orphan scan failed

Distributed mode with redis before 1.8.7. Upgrade: it starts and opens sessions from 1.8.7 on.

Another instance answers 404 session not initialized

It can read the session id but can't find the session: redis isn't set on every instance, points at another Redis or database, uses another keyPrefix, or couldn't be reached when that instance started. Check each instance's startup log for Session store connection validated successfully.

404 invalid session id from an instance that has Redis

The id wasn't made under this instance's MCP_SESSION_SECRET: the instances have different secrets, or the secret was rotated after the client started its session. The client starts a new one. Give every instance the same secret.

signature verification failed on some requests

The instances have different JWT_SECRETs. Give them the same one. See running several instances.