Running Several Instances
One server is one process. When it restarts, it's gone for a moment, and when it's busy, every client waits. Running several copies of it, instances, behind a load balancer fixes both, and the load balancer sends each request to whichever instance it likes. That's only safe for what every instance can answer the same way. This lesson shows what instances must share, why clients on MCP 2026-07-28 need less of it than older clients, and how Redis carries the rest.
You will learn
- Why an instance's memory is its own, and what that breaks
- What clients on MCP 2026-07-28 need from several instances, and what older clients need
- Which secrets every instance must share
- How to share sessions and rate limits through Redis, with a load balancer in front
- The traps in FrontMCP 1.9.4:
REDIS_URLalone, and rate limits that need Redis of their own
Each instance has its own memory
Support agents add notes to tickets. add_note keeps them in a provider, which lives in the server's memory. That's fine for one instance. The tests below start two instances from the same configuration, as two processes behind a load balancer would be, and send calls to each. Open the Tests tab:
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { config } from "./config";
// One tools/call, as a 2026-07-28 client sends it
function call(name: string, args: Record<string, unknown>) {
return new Request("https://desk.example.com/", {
method: "POST",
headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/call", "mcp-name": name },
body: JSON.stringify({
jsonrpc: "2.0",
id: 1,
method: "tools/call",
params: { name, arguments: args, _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } },
}),
});
}
async function send(instance: (request: Request) => Promise<Response>, name: string, args: Record<string, unknown>) {
return (await (await instance(call(name, args))).json()).result.structuredContent;
}
test("either instance answers a call that carries what it needs", async () => {
const a = await FrontMcpInstance.createFetchHandler(config);
const b = await FrontMcpInstance.createFetchHandler(config);
expect(await send(a, "get_ticket", { id: "T-1" })).toEqual({ id: "T-1", title: "Cannot log in", status: "open" });
expect(await send(b, "get_ticket", { id: "T-1" })).toEqual({ id: "T-1", title: "Cannot log in", status: "open" });
});
test("🚩 a note added on one instance is missing on the other", async () => {
const a = await FrontMcpInstance.createFetchHandler(config);
const b = await FrontMcpInstance.createFetchHandler(config);
await send(a, "add_note", { id: "T-1", note: "Asked for a screenshot" });
expect(await send(a, "get_notes", { id: "T-1" })).toEqual({ id: "T-1", notes: ["Asked for a screenshot"] });
expect(await send(b, "get_notes", { id: "T-1" })).toEqual({ id: "T-1", notes: [] });
});Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
Each createFetchHandler() builds a server of its own, with its own providers, the way each process does. get_ticket works on both, because everything it needs is in the code and the request. The note only exists on the instance that took it, so an agent whose next request lands on the other instance sees no notes at all.
The fix is to keep the notes somewhere every instance reaches, and give each instance a store that talks to it. A provider can be handed to the server from outside, with useValue:
import { Provider } from "@frontmcp/sdk";
// ✅ The notes live wherever `db` is, and every instance is given the same one
@Provider({ name: "NoteStore" })
export class NoteStore {
constructor(private readonly db: Map<string, string[]>) {}
add(id: string, note: string) {
this.db.set(id, [...this.list(id), note]);
}
list(id: string) {
return this.db.get(id) ?? [];
}
}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
Both instances run in this page, so handing them one Map is the Playground's stand-in for a database they both connect to. Two real processes share nothing, not even a module variable: in production, NoteStore reads and writes your database, or Redis, and each instance is started with one connected to the same place. The tools don't change, which is the point of putting state in a provider.
What clients need from several instances
The same question, "does another instance know about this?", decides what FrontMCP's own state needs, and the answer depends on the client.
A client on MCP 2026-07-28 keeps no session. Every request carries what the server needs to answer it, the way get_ticket did above. A client on an older protocol version starts with initialize, gets an Mcp-Session-Id, and sends it with every request after that. The session lives in the memory of the instance that created it.
To see the difference, the help desk gets two more tools next to get_ticket: one that asks the user two questions, and one with a rate limit:
@Tool({
name: "assign_ticket",
description: "Assign a ticket to a support agent. Asks the user who, then asks to confirm.",
inputSchema: { id: z.string() },
})
export class AssignTicket extends ToolContext {
async execute({ id }: { id: string }) {
const who = await this.elicit(`Who should take ${id}?`, z.object({ agent: z.enum(["Nour", "Sam", "Priya"]) }));
const agent = who.status === "accept" ? who.content?.agent : undefined;
if (!agent) return { id, assigned: null };
const sure = await this.elicit(`Assign ${id} to ${agent}?`, z.object({ confirm: z.boolean() }));
return { id, assigned: sure.status === "accept" && sure.content?.confirm === true ? agent : null };
}
}
@Tool({
name: "export_tickets",
description: "Export every support ticket as CSV. At most 2 exports an hour.",
inputSchema: {},
rateLimit: { maxRequests: 2, windowMs: 60 * 60_000 },
})
export class ExportTickets extends ToolContext {
async execute() {
return { csv: ["id,title,status", ...tickets.map((t) => `${t.id},${t.title},${t.status}`)].join("\n") };
}
}Built with the main.ts from the next section, and started as two instances on one machine with no shared state and nothing set but the port:
PORT=4001 node dist/node/help-desk.bundle.js &
PORT=4002 node dist/node/help-desk.bundle.js &What a client sees when its requests reach both:
| The client | Sends | Gets |
|---|---|---|
| On MCP 2026-07-28 | get_ticket to 4001, then to 4002 | The ticket, from both. |
| On MCP 2025-06-18 | initialize to 4001, then get_ticket with its session id to 4002 | 404 {"code":-32000,"message":"invalid session id"} from 4002, which can't read an id that 4001 encrypted with its own secret. |
| On MCP 2026-07-28 | assign_ticket to 4001, the answer to its first question to 4002, the answer to its second to 4001 | The first question again, Who should take T-1?, and 4001 logs rejected requestState with reason: 'bad-signature' and a hint that names VAULT_SECRET. |
| Either | export_tickets five times, alternating | Four exports: each instance counts its own two. |
The second row shows what a session id is: it's encrypted, and an instance serves only an id it can decrypt. Give both instances the same MCP_SESSION_SECRET and 4002 reads the id, but it still has no such session, so it answers 404 session not initialized. Keeping the session where 4002 can find it is Redis's job, below.
The third row is the surprising one. Between questions, a 2026-07-28 client carries the earlier answers in a requestState the server signed, and each instance made up its own signing key. Nothing is stored, so Redis doesn't help; a shared secret does (elicit() says which one signs it). With the same VAULT_SECRET on both instances, the same client gets {"id":"T-1","assigned":"Priya"}. Without it, the rejected state is logged with the fix:
WARN mcp-20260728: rejected requestState {
method: 'tools/call',
reason: 'bad-signature',
hint: 'requestState is signed with a per-process key; set VAULT_SECRET (or JWT_SECRET) to the same value on every instance so a round can land on any of them'
}
So a 2026-07-28 client needs every instance to run the same build with the same secrets. An older client also needs its session to be found wherever its request lands.
The secrets every instance shares
| Variable | Every instance needs the same one when | Otherwise |
|---|---|---|
MCP_SESSION_SECRET | Clients before MCP 2026-07-28 connect. Production requires it anyway (previous lesson). | It encrypts their session ids. An instance answers 404 invalid session id to an id made under another secret (High availability). |
VAULT_SECRET, or JWT_SECRET | A tool asks the user more than one question. | The first question is asked again, as above. |
JWT_SECRET | The server uses local or remote auth. | Tokens issued by one instance fail on another with signature verification failed. |
Generate each once, keep it in your platform's secret store, and give every instance the same value. Auth in production lists the rest of what authenticated servers share.
A production server that shares state through redis, or through a transport.persistence object, and has neither VAULT_SECRET nor JWT_SECRET, warns once as it starts:
WARN requestState is signed with a per-process key (neither VAULT_SECRET nor JWT_SECRET is set). Multi-round tools (elicit/sample) restart when a round lands on another instance; set VAULT_SECRET (or JWT_SECRET) to the same value on every instance.
A server without either option gets no warning at startup, and the hint above is the only sign, when a round lands on the wrong instance. A blank value counts as unset.
Sharing sessions and limits through Redis
@FrontMcp({ redis }) moves the state FrontMCP keeps between requests into Redis: older clients' sessions, background tasks, and questions waiting for an answer. Rate-limit counts have an option of their own, throttle.storage. Read the address from the environment, so the same build runs with or without Redis:
import "reflect-metadata";
import { FrontMcp } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";
const url = process.env.REDIS_URL; // like redis://redis:6379, or rediss://default:…@cache.example.com:6380
@FrontMcp({
info: { name: "help-desk", version: "1.0.0" },
apps: [HelpDesk],
http: { entryPath: "/mcp" },
elicitation: { enabled: true },
...(url ? { redis: { url }, throttle: { enabled: true, storage: { type: "redis", redis: { url } } } } : {}),
})
export default class Server {}Both options take the URL: redis as { url }, and throttle.storage under its redis key. (redis also takes host, port, password, db and tls as fields.)
Docker Compose can run the whole setup on one machine: Redis, two instances of the image from the last lesson, and nginx in front of them, sending requests to each in turn:
services:
redis:
image: redis:7-alpine
healthcheck:
test: ["CMD", "redis-cli", "ping"]
interval: 3s
timeout: 5s
retries: 3
desk:
build:
context: ..
dockerfile: ci/Dockerfile
init: true
deploy:
replicas: 2
environment:
MCP_SESSION_SECRET: ${MCP_SESSION_SECRET:?set MCP_SESSION_SECRET}
VAULT_SECRET: ${VAULT_SECRET:?set VAULT_SECRET}
REDIS_URL: redis://redis:6379
depends_on:
redis:
condition: service_healthy
lb:
image: nginx:1.27-alpine
ports:
- "127.0.0.1:3000:80"
volumes:
- ./nginx.conf:/etc/nginx/conf.d/default.conf:ro
depends_on:
- deskupstream desk {
server desk:3000;
}
server {
listen 80;
location / {
proxy_pass http://desk;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header Connection "";
proxy_buffering off;
proxy_read_timeout 1h;
}
}export MCP_SESSION_SECRET="$(openssl rand -hex 32)" VAULT_SECRET="$(openssl rand -hex 32)"
docker compose -f ci/compose.yml up -d --buildEach instance logs that it found Redis:
[TransportService] redis session store will be initialized for transport persistence
[TransportService] Session store connection validated successfully
Now every client in the table above is served through http://127.0.0.1:3000/mcp. The older client's initialize and its next four calls all succeed, though nginx sent some of them to the other instance, which logs Recreating transport from stored session as it picks the session up from Redis. assign_ticket finishes with {"id":"T-1","assigned":"Priya"}, and the third export in the hour is refused, whichever instance gets it. Redis holds a key for each:
docker compose -f ci/compose.yml exec redis redis-cli --scanmcp:session:<session id>
mcp:guard:export_tickets:global:rl:<window>
/readyz now pings Redis too, so a load balancer that checks it stops sending traffic to an instance that can't reach it. Redis has every option, tls for managed Redis among them, and High availability the rest of what instances share.
What else lives in one instance's memory
| State | Shared through |
|---|---|
| Sessions of clients before MCP 2026-07-28 | redis |
| Background tasks, and questions waiting for an answer | redis |
| Rate-limit and concurrency counts | throttle.storage (Guard options) |
| Plugin stores, like the Cache plugin's | The plugin's type: "global-store", which uses redis |
| Job runs | jobs.store (Where runs are kept) |
Sign-ins and refresh tokens of local and remote auth | auth.tokenStorage (Local auth) |
| Your own state, like the notes | Your database, or Redis, given to every instance |
FrontMCP checks redis as the server is built, and refuses some of these at startup when they're wrong, which a test can catch before you deploy. The Playground can't reach Redis, but it can build the server:
import { test, expect } from "@frontmcp/testing";
import { App, FrontMcpInstance, redisOptionsSchema } from "@frontmcp/sdk";
import { CachePlugin } from "@frontmcp/plugin-cache";
import { GetTicket, HelpDesk } from "./help-desk.app";
const info = { name: "help-desk", version: "1.0.0" };
test("`redis` takes the URL a managed Redis hands you", () => {
expect(redisOptionsSchema.parse({ url: "rediss://default:s3cr3t@cache.example.com:6380/2" })).toMatchObject({
host: "cache.example.com",
port: 6380,
password: "s3cr3t",
db: 2,
tls: true,
});
});
test("a `redis` URL that names another user is refused", async () => {
const server = FrontMcpInstance.createFetchHandler({ info, apps: [HelpDesk], redis: { url: "rediss://support:s3cr3t@cache.example.com:6380" } });
await expect(server).rejects.toThrow("only the default user is supported");
});
test("🚩 a plugin's global store needs `redis` on the server", async () => {
@App({ id: "cached", name: "Cached", tools: [GetTicket], plugins: [CachePlugin.init({ type: "global-store" })] })
class Cached {}
await expect(FrontMcpInstance.createFetchHandler({ info, apps: [Cached] })).rejects.toThrow(
'Plugin "CachePlugin" requires global "redis" configuration.',
);
});Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
redis takes a redis:// or rediss:// URL, the form managed Redis services hand you, and splits it into host, port, password, db and tls. A URL that names a user other than default is refused: the error says redis.url names the ACL user "support"; only the default user is supported, and a deployed server exits with it at startup. The last test is about plugins: one whose store is "global-store" uses the server's redis, and a server without it doesn't start.
Deep diveWhat distributed mode addsShow detailsHide details
The node build above is all several instances need: any of them loads an older client's session from Redis. frontmcp build --target distributed (or FRONTMCP_DEPLOYMENT_MODE=distributed) does it the other way round: an instance keeps serving the sessions it made, and another instance that gets one of their requests relays it to that instance over Redis and streams the answer back. Each instance writes a heartbeat to Redis. One stopped with SIGTERM deletes its heartbeat as it shuts down, so whichever instance gets its sessions' next requests takes them over at once; one that crashed has them taken over once its heartbeat expires, 30 seconds after it stopped, and until then those requests answer 503 with Retry-After: 30. Every answer names the instance that served it in an X-FrontMCP-Machine-Id header, for load balancers that route by it. High availability has the details.
Changed in 1.9: in 1.8.7 a request that carried a session another instance had made answered 500 (RemoteTransporter.handleRequest() is not implemented), so in distributed mode an older client had to reach the instance that made its session.
Vercel and Lambda functions are several instances by themselves: each can run many copies at once, and none keeps anything between invocations. They need the same things: shared secrets, and for older clients' sessions a store every copy reaches, like a Redis in the function's network or, on Vercel, Vercel KV, which can't serve background tasks and questions waiting for an answer: FrontMCP skips tasks with a warning, and refuses elicitation unless it has a Redis of its own.
Deep diveWhat about sticky sessions?Show detailsHide details
A load balancer can pin each client to one instance, usually with a cookie, so an older client's session is always where its requests go. That avoids Redis for sessions, but an instance that stops or restarts takes its sessions with it, and the rest of the state in the table above still isn't shared. Redis doesn't make pinning useless either: a client that holds an event stream open, like one on the older SSE transport, holds it with one instance, and only that instance can write to it. Give such clients affinity at the load balancer.
Recap
- Each instance has its own memory: its providers, and FrontMCP's own state. A request that lands on another instance doesn't see it.
- Clients on MCP 2026-07-28 carry everything in each request, so they need every instance to run the same build with the same secrets. Older clients also need their session found, through
redisor a load balancer that pins them. - Give every instance the same
MCP_SESSION_SECRET, the sameVAULT_SECRETfor tools that ask more than one question, and the sameJWT_SECRETforlocalandremoteauth. redisshares sessions, tasks and pending questions,throttle.storageshares rate limits, and your own state belongs in a database or Redis that every instance is given. Both options take aredis://URL.REDIS_URLorREDIS_HOSTalone moves tasks and pending questions, but not sessions or rate limits: read it intoredisandthrottle.storage.- A Redis that's down at startup or goes away later is retried: the instance keeps working in memory, and rate-limited tools fail with
GUARD_STORAGE_UNAVAILABLEunlessthrottle.storagehasfallback: "memory".
Try some challenges
Each challenge runs hidden checks against your code. Edit the code, then press Check.
Challenge 1 of 3
Give every instance the same on-call list
set_on_call records which agent is on call, and who_is_on_call tells the model. The checks start two instances with createServer(db), handing both the same db, and createServer ignores it, so each instance keeps its own list. Make every instance keep the on-call agent in the db it's given.
import { HelpDesk } from "./on-call.tools";
export function createServer(db: Map<string, string>) {
return {
info: { name: "help-desk", version: "1.0.0" },
apps: [HelpDesk],
};
}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.