# High availability

> Run a FrontMCP server as several instances behind a load balancer: what they must share, how an older client's session moves between them through Redis, machine ids, and what distributed mode does: relaying a session's requests to its instance, and taking over the sessions of a stopped one.

Source: https://frontmcp.dev/reference/deployment/high-availability

Several copies of a server behind a load balancer keep serving when one of them stops. A client on MCP 2026-07-28 doesn't care which copy answers: every request carries what it needs, so the copies only have to share their secrets. A client on an older protocol version has a session, and the copy that didn't create it needs to find it: that's what a shared Redis is for. In the default mode, any copy loads the session from Redis. In distributed mode, the copy that made a session keeps serving it, the others relay its requests to that copy through Redis, and when that copy stops, another takes its sessions over. [Running Several Instances](https://frontmcp.dev/learn/running-several-instances) teaches the default mode step by step.

```ts
@FrontMcp({ info, apps, redis: { url: process.env.REDIS_URL } })  // the same on every instance
```

```bash
MCP_SESSION_SECRET=… FRONTMCP_BIND_ADDRESS=all node dist/node/help-desk.bundle.js   # the same secret everywhere
```

---

## Reference

### What instances share

| What | Every instance needs | Without it |
| --- | --- | --- |
| `MCP_SESSION_SECRET` | The same value, for clients before MCP 2026-07-28 | In production, those clients' `initialize` answers `500`. An id made under another secret gets `404` `invalid session id`. See [secrets](https://frontmcp.dev/reference/auth/production#secrets). |
| `JWT_SECRET`, `VAULT_SECRET` | The same values, with `local` or `remote` auth, or tools that ask the user more than once | Tokens and answers from one instance fail on another. A production server with `redis` and neither secret warns at startup: `requestState is signed with a per-process key`. See [running several instances](https://frontmcp.dev/reference/auth/production#running-several-instances). |
| `redis` | The same Redis | Sessions of older clients, background tasks and pending questions stay in the instance that created them. |
| `auth.tokenStorage`, `throttle.storage` | Redis too, if you use `local`/`remote` auth or rate limits | Sign-ins that change instance fail, and each instance counts its own limits. |
| The same build | Tools, resources and prompts | A client sees a different server depending on where it lands. Compare `toolsHash` in `/readyz`: see [checking replicas](#checking-that-replicas-serve-the-same-tools). |

Clients on MCP 2026-07-28 open no sessions, so for them `redis` only matters when a tool starts a background task or asks a question.

### How a session moves between instances

With `redis` set, a session is written to Redis when a client opens it, with the id of the instance that holds it (its `nodeId`). When a request of that session reaches an instance that doesn't have it in memory, that instance reads it from Redis, logs `Recreating transport from stored session`, serves the request, and writes it back. So a load balancer can send an older client's requests anywhere, without affinity, and a restarted or replaced instance picks up where the old one left off. It does that only for an id it can decrypt, so every instance needs the same `MCP_SESSION_SECRET` ([caveats](#caveats)).

A session's key lives an hour, or `redis.defaultTtlMs`, and the instance that holds the session pushes it out while it serves the session's requests ([how long sessions last](https://frontmcp.dev/reference/deployment/redis#how-long-sessions-last)). An instance that can't reach Redis when it starts keeps sessions in memory, retries the connection, and uses Redis once it answers, with no restart. Its `/readyz` answers `503` until then, so a load balancer that checks readiness sends it no traffic ([when Redis can't be reached](https://frontmcp.dev/reference/deployment/redis#when-redis-cant-be-reached)).

> **Note**
In 1.8.6, requests that the session's own instance served didn't refresh the key, so a session expired an hour after it was last loaded and an instance that stopped after that left its clients to start over; and an instance that couldn't reach Redis at startup never used it until it restarted. Both are fixed.

### Distributed mode

`FRONTMCP_DEPLOYMENT_MODE=distributed`, which `frontmcp build --target distributed` sets in its entry (`node dist/distributed/index.js`), puts the server in distributed mode:

- It listens on every interface, unless `http.security.bindAddress` or `FRONTMCP_BIND_ADDRESS` says otherwise. The audit says `BIND_ALL_INTERFACES_DISTRIBUTED`.
- `/healthz` reports `"deployment": "distributed"`.
- Every response carries the id of the instance that served it, `X-FrontMCP-Machine-Id: <id>`, for load balancers that route by it: MCP answers of every protocol version, `/healthz` and `404`s. An older client's `initialize` also sets a cookie, `__frontmcp_node=<id>; Path=/; Max-Age=86400; HttpOnly; SameSite=Strict`. (In 1.8.7 only the responses to older clients had the header.)
- With `redis`, it starts the HA manager (`[HA] Manager started for distributed deployment`): a heartbeat key for each instance (`mcp:ha:heartbeat:<id>`), refreshed every 10 seconds and expiring after 30; a scanner for the sessions of stopped instances, every 10 seconds; a bus that records which instance holds each session (`mcp:bus:session:<session id>`); and the relay, `[HA] Request relay ready — sessions owned by other nodes are served through them`. It also stores SSE events in Redis, under `mcp:events:`.

#### How a session's requests are served

| The request reaches | What happens |
| --- | --- |
| The instance that holds the session | It serves it. |
| Another instance, while the holder runs | It relays the request to the holder through Redis, and streams the answer back. The holder runs it, and the answer's `X-FrontMCP-Machine-Id` names the holder. |
| Another instance, after the holder shut down (`SIGTERM` or `SIGINT`) | The instance takes the session over from Redis at once, serves it, and from then on holds it. |
| Another instance, after the holder died without shutting down (a crash, `SIGKILL`, a lost machine) | Until the holder's heartbeat expires: `503`, with `Retry-After` set to the heartbeat's lifetime (`30`) and `{"code":-32000,"message":"This MCP session is served by another instance of the server, which did not answer the relayed request. Retry after 30s; if that instance has stopped, the session is taken over by another instance once its heartbeat expires."}`. After that, it takes the session over as above. |

A server that shuts down deletes its heartbeat key ([stopping](https://frontmcp.dev/reference/deployment/node#stopping)). When instance `desk-a` was stopped with `SIGTERM`, it exited `0`, only `desk-b`'s heartbeat was left in Redis, and the session's next request, sent to `desk-b` a tenth of a second later, was answered `200`, with `[HA] Session owner is gone — the session will be taken over`, `Recreating transport from stored session` and `[HA] Took over session from a stopped node { …, previousNodeId: 'desk-a' }` in its log. Before 1.9.2 the heartbeat stayed until it expired, so after a `SIGTERM` the session answered `503` for 30 seconds. An instance that crashes still leaves its heartbeat behind: after `desk-a` died, `desk-b` answered `503` with `Retry-After: 30`.

> **Note**
In 1.9.1 and 1.9.2 the holder answered a relayed request, then, about half a second later, exited with code `1` and `TypeError: socket.destroySoon is not a function` (from `@hono/node-server` 1.19.17, which the MCP SDK uses), and its other sessions waited for its heartbeat to expire. Now it keeps running: with the same `@hono/node-server`, it answered three relayed calls and a relayed `DELETE`, and its `/healthz` still answered `200` after each.

A load balancer doesn't need to route by the header or the cookie: round-robin works, and routing by them saves the relay through Redis. Every instance needs the same `MCP_SESSION_SECRET`, as in the default mode. Relayed requests pass through Redis with their headers, so keep Redis on a private network. Distributed mode without `redis` works on one instance: it only adds the header and the cookie.

`deployments[].ha` in `frontmcp.config` tunes the HA manager: `heartbeatIntervalMs`, `heartbeatTtlMs`, `takeoverGracePeriodMs` and `redisKeyPrefix`. The distributed build writes them into `serverless-setup.js` as `FRONTMCP_HA_*` variables, which a variable already set wins over. With `ha: { heartbeatIntervalMs: 2000, heartbeatTtlMs: 6000 }`, the `503` said `Retry after 6s` and the takeover came six seconds after the holder stopped. `deployments[].server.cookies` sets the cookie's name, `Domain` and `SameSite`: `cookies: { affinity: "desk_node", sameSite: "Lax" }` gave `desk_node=<id>; …; SameSite=Lax`.

> **Note**
In 1.8.7, a request whose session another instance had made answered `500 Internal Server Error`, with `RemoteTransporter.handleRequest() is not implemented` in the log, whether that instance was running or stopped, and the scanner's `claimed session` changed nothing. An older client had to reach the instance that made its session, and the session died with it. `deployments[].ha` was read by nothing.

> **Note**
In 1.8.6, the HA parts were handed the `redis` options where they expected a Redis client. The server logged `[HA] Failed to start HA manager — running without HA` with `TypeError: this.redis.set is not a function`, every ten seconds `[HA] Orphan scan failed: TypeError: this.redis.keys is not a function`, and every `initialize` answered `500` with `this.redis.hset is not a function`, so no client before MCP 2026-07-28 could connect. All of that is gone.

### Machine ids

Each instance has an id. It's the `nodeId` of the sessions it creates, and the value of the header and cookie in distributed mode.

| Source | When |
| --- | --- |
| `setMachineIdOverride(id)` | Whenever your code has called it, over `MACHINE_ID` too. See below. |
| `MACHINE_ID` | Whenever it's set. |
| `HOSTNAME`, else the operating system's host name | In distributed mode. Kubernetes sets `HOSTNAME` to the pod's name. |
| A random id per process | In serverless mode. |
| A random id, new on every start | Otherwise. FrontMCP's source mentions `.frontmcp/machine-id`, in the working folder outside production, and a file that `MACHINE_ID_PATH` names; starting the `node` build in its project folder wrote no such file. |

Outside production, FrontMCP encrypts session ids with a key made from the machine id when `MCP_SESSION_SECRET` isn't set, so without `MACHINE_ID` the key changes on every start: after a restart, an old session id gets `404` `invalid session id`, Redis or not. Production requires the secret for sessions, so there the machine id only names the instance: sharing `MACHINE_ID` between instances doesn't share anything else.

Your code reads the id with `getMachineId()` from `@frontmcp/utils`, the function FrontMCP itself calls, and replaces it for the whole process with `setMachineIdOverride(id)`; `setMachineIdOverride(undefined)` goes back to the sources above. Call it before the server starts. The header follows a later call, but the heartbeat doesn't: in distributed mode with `redis`, an instance started as `help-desk-7b8f9`, then given `desk-a`, kept `mcp:ha:heartbeat:help-desk-7b8f9` while its responses said `X-FrontMCP-Machine-Id: desk-a`.

```ts main.ts
import "reflect-metadata";
import { getMachineId, setMachineIdOverride } from "@frontmcp/utils";

setMachineIdOverride(`desk-${process.env.REGION}-${process.pid}`); // before @FrontMcp starts the server

// later, anywhere in the process
console.log(getMachineId()); // desk-eu-west-1-4182
```

`@frontmcp/utils` is installed with `@frontmcp/sdk`; add it to your own dependencies, at the same version, to import it. [`create({ machineId })`](https://frontmcp.dev/reference/sdk/create) sets the same override for one in-process server, but only in an ES module project: in a CommonJS one, FrontMCP 1.9.4 sets it on a copy of `@frontmcp/utils` that the server never reads, and the id stays as it was. Call `setMachineIdOverride()` there instead. This was checked with FrontMCP 1.9.4 on Node: `create({ machineId })` in both kinds of project, and the override against distributed mode's header and Redis heartbeat.

#### Caveats

- In the default mode, sessions are shared, streams aren't: a client that holds an event stream open, like the older SSE transport, holds it with one instance. Give such clients affinity at the load balancer.
- An instance serves a session id only when it decrypts under its own `MCP_SESSION_SECRET`. An id made under another secret gets `404` `invalid session id`, and the log says `mcp-session-id is not a session this server verified`: give every instance the same secret. After you rotate it, a client with an old id gets that `404` once and starts a new session. (In 1.8.5, a public server served such an id from Redis.)
- In the default mode, a `DELETE` of a session ends it on the instance that receives it and deletes it from Redis; a later request there gets `404` `Session not found`. Another instance that has the session in memory, because it created the session or served it before, checks with Redis (`EXISTS`) on each of the session's requests that the session is still stored, and drops it when it isn't: its next request got `404` `{"code":-32001,"message":"session expired"}`, and its log said `The stored session is gone — dropping its transport here`. A session whose key expired gets the same `404`, on the instance that holds it too ([how long sessions last](https://frontmcp.dev/reference/deployment/redis#how-long-sessions-last)). (Changed in 1.9.3: in 1.9.2 such an instance kept serving a deleted or expired session, and answered its next request `200`.) In distributed mode the `DELETE` is relayed to the holder, which answers `204` with its own `X-FrontMCP-Machine-Id` and deletes the stored session, and the session is gone everywhere.
- That check doesn't make a session depend on Redis. When Redis refuses connections, or doesn't answer within `transport.persistence.sessionCheckTimeoutMs` (500 ms by default), the instance serves the session from memory, logs `[TransportService] Could not confirm the session is still stored — serving it here`, and checks again on the session's next request. With the Redis container paused, which keeps the connection open and answers nothing, each call on a session the instance held answered `200` after half a second, with `The session store did not answer within 500 ms` in the log; with `sessionCheckTimeoutMs: 2000`, after two seconds. A request for a session the instance doesn't hold has no copy to fall back on, and waits for Redis: it got no answer while Redis was paused, and was answered once it came back. (Changed in 1.9.4: in 1.9.3 a call on a session the instance held waited for a paused Redis too, and got no answer in three minutes.) See [when Redis can't be reached](https://frontmcp.dev/reference/deployment/redis#when-redis-cant-be-reached).
- This was checked with FrontMCP 1.9.4: two instances of a Node build on one machine sharing Valkey 8 in Docker, with a `DELETE` on one, and Redis stopped and paused, and two of the distributed build sharing it: relayed calls, a relayed `DELETE` and a holder stopped with `SIGTERM`. The crash was seen with 1.9.2. Two Node instances behind nginx in Docker Compose, and a shorter heartbeat from `deployments[].ha`, were last run with 1.9.1. No Kubernetes was used, and no load balancer routed by the cookie or the header.
- **A question to the user waits on the instance that asked.** For clients before 2026-07-28, a tool's [`this.elicit()`](https://frontmcp.dev/reference/sdk/elicit) waits on the instance it runs on, and the answer reaches it from another instance through Redis publish and subscribe, encrypted when `MCP_ELICITATION_SECRET` (or `MCP_SESSION_SECRET`) is set, so every instance needs the same secret: see [Several instances and encryption](https://frontmcp.dev/reference/sdk/elicit#several-instances-and-encryption).

---

## Usage

### Running two instances that share sessions

The server has `redis` set from the environment, as in [Redis](https://frontmcp.dev/reference/deployment/redis#redis). Start two instances with the same secret, on two ports:

```bash
export MCP_SESSION_SECRET="$(openssl rand -hex 32)"
MACHINE_ID=desk-a PORT=4001 REDIS_URL=redis://127.0.0.1:6379 node dist/node/help-desk.bundle.js &
MACHINE_ID=desk-b PORT=4002 REDIS_URL=redis://127.0.0.1:6379 node dist/node/help-desk.bundle.js &
```

A client on MCP 2025-06-18 sends `initialize` to port 4001, and gets an `Mcp-Session-Id`. It sends its next `tools/call`, with that id, to port 4002, which has never seen it: the call succeeds, and the second instance logs:

```text
[TransportService] Recreating transport from stored session
[StreamableHttpAdapter] Marked transport as pre-initialized for session recreation
```

Without `redis`, the second instance answers `404` `session not initialized`, and the client has to `initialize` again. In a container, add `FRONTMCP_BIND_ADDRESS=all` so the load balancer can reach each instance.

### Checking that replicas serve the same tools

`/readyz` has a hash of the tool names, the same on every instance built from the same code:

```bash
curl -s http://127.0.0.1:4001/readyz | jq -r .catalog.toolsHash
curl -s http://127.0.0.1:4002/readyz | jq -r .catalog.toolsHash
```

```text
<the same 64 hex characters>
<the same 64 hex characters>
```

During a rolling deploy, two hashes mean two versions are serving. It covers tool names only, not their schemas.

### Routing by instance, in distributed mode without Redis

A load balancer that pins a client to an instance can use the cookie distributed mode sets on `initialize`:

```bash
FRONTMCP_DEPLOYMENT_MODE=distributed HOSTNAME=help-desk-7b8f9-abc12 node dist/node/help-desk.bundle.js
```

```http
X-FrontMCP-Machine-Id: help-desk-7b8f9-abc12
Set-Cookie: __frontmcp_node=help-desk-7b8f9-abc12; Path=/; Max-Age=86400; HttpOnly; SameSite=Strict
```

Without `redis`, sessions stay in memory, so this only helps clients that return to the same instance, and one that stops takes its sessions with it. With `redis`, any instance serves a session, by relaying it to the instance that holds it ([above](#how-a-sessions-requests-are-served)), and the cookie saves that relay.

---

## Troubleshooting

### A request answers `500`, and the log says `RemoteTransporter.handleRequest() is not implemented`

Distributed mode with `redis` before 1.9, and the request's `Mcp-Session-Id` belongs to a session another instance created. Upgrade: since 1.9 the request is relayed to that instance. See [distributed mode](#distributed-mode).

### `503`, `This MCP session is served by another instance of the server, which did not answer the relayed request`

Distributed mode: the instance that holds the session didn't answer the relay, most often because it died without shutting down, and its heartbeat hasn't expired yet. Retry after `Retry-After` seconds; once the heartbeat expires, the instance that gets the request takes the session over. The log says `Relay to session owner failed`. If the holder is running, check that it reaches the same Redis. See [how a session's requests are served](#how-a-sessions-requests-are-served).

### `TypeError: socket.destroySoon is not a function`, and the instance exits

The instance served a request relayed from another instance, in distributed mode with `redis`, on FrontMCP 1.9.1 or 1.9.2. Upgrade to 1.9.3, where it keeps running. On those versions, route each session's requests to the instance that holds it, by the `__frontmcp_node` cookie or the `X-FrontMCP-Machine-Id` header, so that nothing is relayed. See [how a session's requests are served](#how-a-sessions-requests-are-served).

### `[HA] Failed to start HA manager — running without HA`, `this.redis.hset is not a function` or `[HA] Orphan scan failed`

Distributed mode with `redis` before 1.8.7. Upgrade: it starts and opens sessions from 1.8.7 on.

### Another instance answers `404` `session not initialized`

It can read the session id but can't find the session: `redis` isn't set on every instance, points at another Redis or database, uses another `keyPrefix`, or couldn't be reached when that instance started. Check each instance's startup log for `Session store connection validated successfully`.

### `404` `invalid session id` from an instance that has Redis

The id wasn't made under this instance's `MCP_SESSION_SECRET`: the instances have different secrets, or the secret was rotated after the client started its session. The client starts a new one. Give every instance the same secret.

### `signature verification failed` on some requests

The instances have different `JWT_SECRET`s. Give them the same one. See [running several instances](https://frontmcp.dev/reference/auth/production#running-several-instances).
