Guard options

FrontMCP's guard limits tool calls: how often a tool runs (rateLimit), how many calls of it run at once (concurrency), and how long one call may take (timeout). A tool sets its own limits in @Tool. @FrontMcp({ throttle }) sets defaults for every tool, adds limits for the server as a whole, and filters requests by IP address. A call over a limit never reaches execute(), and the caller gets an error that says which limit it hit. Limiting and Timing Out Calls teaches the same options step by step.

@Tool({ name, inputSchema, rateLimit, concurrency, timeout })

@FrontMcp({
  info, apps,
  throttle: { enabled, defaultRateLimit, defaultConcurrency, defaultTimeout, global, globalConcurrency, ipFilter, storage, keyPrefix },
})

Reference

A tool's limits

Set rateLimit, concurrency and timeout in a tool's options. They work without any throttle on the server.

export-tickets.tool.ts
import { Tool, ToolContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV. At most 2 exports an hour per caller, one at a time.",
  inputSchema: {},
  rateLimit: { maxRequests: 2, windowMs: 60 * 60_000, partitionBy: "userId" },
  concurrency: { maxConcurrent: 1, queueTimeoutMs: 5_000 },
  timeout: { executeMs: 30_000 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: await buildCsv({ signal: this.signal }) };
  }
}

See more examples below.

rateLimit

How many calls of the tool a window allows.

FieldTypeDefaultDescription
maxRequestsnumberrequiredCalls allowed per window. A positive integer.
windowMsnumber60_000The window's length, in milliseconds.
partitionByPartitionKey"global"Whose calls share a count. See partitionBy.

A call over the limit fails with RATE_LIMIT_EXCEEDED and Rate limit exceeded. Retry after N seconds. Windows are aligned to the clock (an hourly window runs from one full UTC hour to the next), and N is the time left in the current one.

The count is an estimate over two windows: calls in the current window, plus the previous window's calls weighted by how much of the previous window is still within windowMs of now. So right after a busy window, fewer calls get through than maxRequests, and once "Retry after N seconds" has passed, one call can get through and the next be refused again. (That depends on the clock, so no example below shows it.)

concurrency

How many calls of the tool may run at the same moment.

FieldTypeDefaultDescription
maxConcurrentnumberrequiredCalls allowed to run at once. A positive integer.
queueTimeoutMsnumber0How long a call waits for a free slot, in milliseconds. 0 refuses it at once.
partitionByPartitionKey"global"Whose calls share the slots. See partitionBy.

A call that finds no free slot fails with CONCURRENCY_LIMIT and Concurrency limit reached for "merge_customers" (max: 1). With a queue, it checks for a free slot every 100 ms at first, backing off to once a second, and fails with QUEUE_TIMEOUT and Queue timeout for "merge_customers" after waiting 5000ms for a concurrency slot if none frees up in time. A slot is taken before the input is validated and given back when execute() ends, whether it succeeded, failed or timed out.

timeout

FieldTypeDefaultDescription
executeMsnumberrequiredHow long execute() may run, in milliseconds. A positive integer.

When time runs out, the call fails with EXECUTION_TIMEOUT and Execution of "lookup_customer" timed out after 500ms, and this.signal is aborted with that same error. execute() itself keeps running unless it listens to the signal, and it keeps its concurrency slot until it ends. See A timeout ends the call, not the work.

partitionBy

partitionBy decides which calls share a count (for rateLimit) or a set of slots (for concurrency). FrontMCP turns it into a key, and calls with the same key share:

partitionByKey
"global""global": every call shares one count.
"userId"The caller's userId, or sessionId when there's none.
"session"sessionId.
"ip"The client IP address; without one, "user:" and the caller's userId, or else "ip:unresolved".
(ctx) => stringWhatever the function returns.

The key is built from a PartitionKeyContext:

FieldWhat it holds in FrontMCP 1.9.3
userIdThe signed-in caller's client id from authInfo, like "static:6ab9f1eb8f7d" for a static token. Missing for anonymous callers.
sessionIdFor MCP 2026-07-28 clients, which have no session: "user:" and the userId, or "anonymous". For older clients in a session, over Streamable HTTP or the old SSE transport, the session id. (Before 1.8.5, SSE clients counted as having none.)
clientIpThe address the request came from. Behind a proxy, set the FRONTMCP_TRUST_PROXY=true environment variable (and FRONTMCP_TRUSTED_PROXY_DEPTH for more than one hop) to read X-Forwarded-For instead. Under createFetchHandler(), it's the address the runtime reports, when it reports one (Cloudflare Workers, Bun, Deno). The Playground has none.

So under MCP 2026-07-28, "userId" and "session" both give each signed-in caller a count of their own, and all anonymous callers share one, the key "anonymous". Nobody gets around a limit by not signing in, but one busy anonymous client can use up the limit for every other one. See it run.

@FrontMcp({ throttle })

throttle configures the guard for the whole server.

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: "userId" },
    defaultConcurrency: { maxConcurrent: 5 },
    defaultTimeout: { executeMs: 30_000 },
    global: { maxRequests: 1_000, windowMs: 60_000, partitionBy: "ip" },
    ipFilter: { denyList: ["192.0.2.0/24"] },
  },
})
export default class Server {}
OptionTypeDefaultDescription
enabledbooleanrequiredtrue turns every option below on. false turns the guard off: no rate limits or concurrency limits anywhere, including each tool's own.
defaultRateLimitRateLimitConfignoneA rateLimit for every tool that doesn't set its own. Each tool counts separately.
defaultConcurrencyConcurrencyConfignoneA concurrency for every tool that doesn't set its own. Each tool has its own slots.
defaultTimeoutTimeoutConfignoneA timeout for every tool that doesn't set its own.
globalRateLimitConfignoneOne rate limit for every request the server gets, of any kind. See global.
globalConcurrencyConcurrencyConfignoneHow many tool calls may run at once, across every tool. See globalConcurrency.
ipFilterIpFilterConfignoneAllow or refuse requests by client IP. See ipFilter.
storageStorageConfigmemoryWhere counts and slots are kept. See storage.
keyPrefixstring"mcp:guard:"Prefix for every key the guard stores.

A tool's own rateLimit, concurrency or timeout replaces the default: it doesn't add to it.

When the guard runs

Server configurationWhat's enforced
No throttleEach tool's own rateLimit, concurrency and timeout.
throttle: { enabled: true, … }Each tool's own limits, and every throttle option.
throttle: { enabled: false, … }Only each tool's own timeout. Every other limit, and every throttle option, is ignored.

Only tool calls are limited. Resources, prompts and other requests don't count toward a tool's rateLimit or a default one; global is the one limit that counts them. An agent is a tool to its clients, invoke_<name>, and @Agent takes the same three options, which limit it like any tool (since FrontMCP 1.8.3): see Limiting who calls an agent, and how often.

The limits are checked after the caller's authorization and before the arguments are validated, so a call with invalid arguments counts toward rateLimit like any other. See it run.

global

A rate limit for the whole server, { maxRequests, windowMs?, partitionBy? }, like rateLimit. Every request counts once, whatever its method: tools/list and resources/read as much as tools/call. It's checked when the request arrives, before FrontMCP knows which tool it's for, so a request over it isn't a tool error: it gets HTTP 429, a Retry-After header with the seconds to wait, and a JSON-RPC error:

{ "jsonrpc": "2.0", "id": 7, "error": { "code": -32029, "message": "Rate limit exceeded. Retry after 36 seconds" } }

With partitionBy "global" or "ip", it's checked before authentication; with "userId", "session" or a function, right after it, so the caller is known.

globalConcurrency

A cap on tool calls running at once across every tool, { maxConcurrent, queueTimeoutMs?, partitionBy? }. It's checked before the tool's own concurrency. A call over it fails with CONCURRENCY_LIMIT and Concurrency limit reached for "global" (max: 2). A tool that calls another with this.callTool() doesn't take a second global slot for the inner call, so maxConcurrent: 1 doesn't deadlock it.

ipFilter

FieldTypeDefaultDescription
allowListstring[][]IP addresses or CIDR ranges, IPv4 or IPv6, to allow.
denyListstring[][]IP addresses or CIDR ranges to refuse.
defaultAction"allow" | "deny""allow"What happens to an address on neither list, and to a request whose address isn't known.
trustProxybooleanfalseNot read. Setting it logs a warning. Use the FRONTMCP_TRUST_PROXY environment variable.
trustedProxyDepthnumber1Not read. Setting it logs a warning. Use FRONTMCP_TRUSTED_PROXY_DEPTH.

Each request is checked before anything else: first the deny list, then the allow list, then defaultAction. An allowList on its own therefore refuses nothing: add defaultAction: "deny" to refuse everyone else. A refused request gets HTTP 403 and JSON-RPC error -32001, Forbidden: client IP rejected by ipFilter.

A request whose address isn't known gets defaultAction. With "deny", that refuses every request from a runtime that doesn't report addresses, including the Playground. Behind a proxy that isn't trusted, every request has the proxy's address. See it run.

storage and keyPrefix

By default the guard keeps its counts and slots in memory, so each server process has its own, and FrontMCP logs a warning that memory storage isn't suitable for several instances. To share limits between instances, give it a store:

throttle: {
  enabled: true,
  storage: { type: "redis", redis: { url: process.env.REDIS_URL } },
  defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: "userId" },
},

storage is @frontmcp/utils' StorageConfig: type is "memory", "redis", "upstash", "vercel-kv", "cloudflare-kv" or "auto", with that backend's options under its own key (redis: { url } or redis: { config: { host, port, password, db, tls } }, upstash, vercelKv, cloudflareKv). A storage without type is "auto": Upstash if UPSTASH_REDIS_REST_URL and UPSTASH_REDIS_REST_TOKEN are set, Vercel KV for KV_REST_API_URL and KV_REST_API_TOKEN, Redis for REDIS_URL or REDIS_HOST, and memory otherwise. If the store can't be reached, the server fails to start in production (NODE_ENV=production), with a GuardStorageUnavailableError that names throttle.storage, and falls back to memory, with a warning, otherwise. Rate limits fail closed on purpose: fallback: "memory" inside storage starts the server anyway, with counts of its own for each instance, and fallback: "error" refuses to start in every environment. See both. The same choice applies if the store stops answering while the server runs: see that too. The top-level redis option isn't used by the guard. The Playground can't reach Redis, so a shared store runs on a real server only.

keyPrefix is put before every key the guard stores, like mcp:guard:export_tickets:global:rl:… for a rate count. Give servers that share a store different prefixes if their limits should be separate. Before 1.8.6 the default prefix left an empty segment, mcp:guard::export_tickets:…, so an instance on 1.8.6 or later doesn't read the counts an older one wrote, and limits briefly split during a rolling deploy.

What a caller gets back

A tool's limits, and every default, answer with an ordinary tool error: an isError result over HTTP 200, with no Retry-After header. The code is in _meta.code, and the message is kept in production. global and ipFilter answer the whole request instead:

LimitAnswerCodeMessage
rateLimit, defaultRateLimitTool errorRATE_LIMIT_EXCEEDEDRate limit exceeded. Retry after 1200 seconds
concurrency, defaultConcurrencyTool errorCONCURRENCY_LIMITConcurrency limit reached for "merge_customers" (max: 1)
globalConcurrencyTool errorCONCURRENCY_LIMITConcurrency limit reached for "global" (max: 2)
Any of them, with queueTimeoutMsTool errorQUEUE_TIMEOUTQueue timeout for "merge_customers" after waiting 5000ms for a concurrency slot
timeout, defaultTimeoutTool errorEXECUTION_TIMEOUTExecution of "lookup_customer" timed out after 500ms
globalHTTP 429, Retry-AfterJSON-RPC -32029Rate limit exceeded. Retry after 36 seconds
ipFilterHTTP 403JSON-RPC -32001Forbidden: client IP rejected by ipFilter
A tool's limit, while throttle.storage is downTool errorGUARD_STORAGE_UNAVAILABLEIn production, Service temporarily unavailable: the rate-limit store cannot be reached. Retry shortly. See while the server runs.
global, while throttle.storage is downHTTP 503, Retry-After: 1JSON-RPC -32603, data.code GUARD_STORAGE_UNAVAILABLEService temporarily unavailable: the rate-limit store cannot be reached

The error classes page lists the classes behind these codes.

Building blocks

@frontmcp/sdk re-exports the guard's own library, @frontmcp/guard, for limits the options can't express, like a limit per ticket rather than per caller:

ExportWhat it does
createGuardManager({ config, logger? })Returns a Promise<GuardManager> for a GuardConfig (enabled is required but not read). Memory storage unless config.storage says otherwise.
GuardManagercheckRateLimit(name, config?, ctx?) → { allowed, remaining, resetMs, retryAfterMs? }, counting the call when it's allowed. checkGlobalRateLimit(ctx?). acquireSemaphore(name, config?, ctx?) and acquireGlobalSemaphore(ctx?) → a ticket with release(), or null when no slot is free. checkIpFilter(ip?), isIpAllowListed(ip?), destroy(). Without a config, each falls back to the manager's defaults.
withTimeout(fn, ms, name)Runs fn(signal) and rejects with ExecutionTimeoutError after ms, aborting signal.
IpFilternew IpFilter(ipFilterConfig).check(ip) → { allowed, reason, matchedRule? }, where reason is "denylisted", "allowlisted" or "default".
SlidingWindowRateLimiter, DistributedSemaphoreThe rate counter and the slot pool the manager uses, over a storage adapter.
resolvePartitionKey(partitionBy, ctx), buildStorageKey(name, key, suffix?)How keys are made, as in the partitionBy table.
guardConfigSchema, rateLimitConfigSchema, concurrencyConfigSchema, timeoutConfigSchema, ipFilterConfigSchema, partitionKeySchemaThe Zod schemas FrontMCP validates these options with. ThrottleConfig is the type of throttle.
GuardError and its subclassesExecutionTimeoutError, ConcurrencyLimitError, QueueTimeoutError, IpBlockedError, IpNotAllowedError, each with code and statusCode.

To refuse a call from your own limit, fail with RateLimitError (a public error, RATE_LIMIT_EXCEEDED), or pass a guard error to this.fail(). Don't throw a guard error from execute(): only ExecutionTimeoutError keeps its code that way; the others arrive as TOOL_EXECUTION_ERROR, hidden in production. See it run.

Caveats

  • enabled is required in a throttle object, even to set only defaults, and the defaults, global, globalConcurrency and ipFilter only work with enabled: true.
  • enabled: false turns off each tool's own limits too, all but timeout. Keep it for local testing.
  • Counts live in memory by default, one set per process. Several instances, or serverless functions, need storage.
  • An in-process server from create() takes throttle as @FrontMcp does.
  • Anonymous callers share one count under "userId" and "session" (see partitionBy).
  • A timeout doesn't stop execute(). Pass this.signal on to the work, and don't give a tool that changes things a timeout shorter than its work.
  • Say the limit in the tool's description, so the model can plan around it instead of learning it from an error.
  • In a browser, calls may not overlap: without the TC39 AsyncContext, FrontMCP's browser build serves one request at a time, so concurrency and globalConcurrency have nothing to limit there. The Playground provides a minimal AsyncContext, so its calls overlap as they do on Node.

Usage

Limiting one tool

A tool's limits

Example 1 of 3

Rate limit

At most two exports an hour. The third call is refused, with the time until the window ends.

Open
import { Tool, ToolContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV. At most 2 exports an hour.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
  rateLimit: { maxRequests: 2, windowMs: 60 * 60_000 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title,status\nT-1,Cannot log in,open\nT-2,Invoice total is wrong,closed" };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Setting defaults for every tool

With enabled: true, throttle gives every tool the limits it doesn't set itself. Each tool counts on its own: search_tickets and get_ticket get five calls an hour each, not five between them. export_tickets sets its own rateLimit, which replaces the default. sync_crm runs into the default timeout, and weekly_report into the default concurrency:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { ExportTickets, GetTicket, SearchTickets, SyncCrm, WeeklyReport } from "./tickets.tools";

@App({ id: "help-desk", name: "Help Desk", tools: [SearchTickets, GetTicket, ExportTickets, SyncCrm, WeeklyReport] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    defaultRateLimit: { maxRequests: 5, windowMs: 60 * 60_000 },
    defaultConcurrency: { maxConcurrent: 1 },
    defaultTimeout: { executeMs: 300 },
  },
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Limiting the whole server

Server-wide limits

Example 1 of 2

Requests

global counts every request, whatever it is. Here the server takes ten an hour, a tiny limit to make the point: the Playground used six when it started (five discovery and list requests, and the call), so press Call until it's refused. In the first test, nine lists and a call get through, and the eleventh request is refused before FrontMCP looks at what it is. The second test sends raw HTTP, to see the status and the Retry-After header, which mcp doesn't show.

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { SearchTickets } from "./search-tickets.tool";

@App({ id: "help-desk", name: "Help Desk", tools: [SearchTickets] })
export class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: { enabled: true, global: { maxRequests: 10, windowMs: 60 * 60_000 } },
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Choosing whose calls count

With partitionBy: "userId", each signed-in caller gets a count of their own, and every anonymous caller shares one. The Playground's client is anonymous, so the first test shows the shared count. The second builds the same server with static API keys and calls it as two different callers. The partitionBy function in get_ticket records the context it's given, to show what a function can key on:

Open
import { Tool, ToolContext, z, type PartitionKeyContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every ticket as CSV. Once an hour per signed-in caller; once an hour for all anonymous callers together.",
  inputSchema: {},
  rateLimit: { maxRequests: 1, windowMs: 60 * 60_000, partitionBy: "userId" },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title\nT-1,Cannot log in" };
  }
}

export const seen: PartitionKeyContext[] = [];

@Tool({
  name: "get_ticket",
  description: "Get one ticket",
  inputSchema: { id: z.string() },
  rateLimit: {
    maxRequests: 100,
    partitionBy: (ctx) => {
      seen.push(ctx);
      return ctx.userId ?? ctx.clientIp ?? "anonymous";
    },
  },
})
export class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in" };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Filtering by IP address

ipFilter checks each request's address before anything else. The Playground's requests have no address, so a deny list lets them through (they get defaultAction, "allow"), and the tests use IpFilter, the class behind the option, to show how addresses are matched. The last test builds the server with defaultAction: "deny", which refuses every request whose address isn't known:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { SearchTickets } from "./search-tickets.tool";

@App({ id: "help-desk", name: "Help Desk", tools: [SearchTickets] })
export class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    ipFilter: { denyList: ["192.0.2.0/24", "2001:db8::/32"] },
  },
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

On a real server behind a load balancer, set FRONTMCP_TRUST_PROXY=true so FrontMCP reads the address from X-Forwarded-For; otherwise every request has the load balancer's address, or none.

Limiting by an argument

partitionBy sees the caller, not the arguments. For "each ticket can be escalated once an hour", keep a GuardManager of your own and key it by ticket. Fail with RateLimitError, so the model reads the same message and code as for any rate limit:

Open
import { createGuardManager, RateLimitError, Tool, ToolContext, z } from "@frontmcp/sdk";

const limits = createGuardManager({ config: { enabled: true } });

@Tool({
  name: "escalate_ticket",
  description: "Page the on-call engineer about a ticket. Each ticket can be escalated once an hour.",
  inputSchema: { id: z.string().describe("Ticket id, like T-1") },
})
export class EscalateTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    const check = await (await limits).checkRateLimit(`escalate:${id}`, { maxRequests: 1, windowMs: 60 * 60_000 });
    if (!check.allowed) this.fail(new RateLimitError(Math.ceil(check.retryAfterMs! / 1000)));
    return { id, escalated: true };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The manager above keeps its counts in memory, per process. Give it the same storage as throttle to share them.

Turning the guard off

throttle: { enabled: false } switches off every rate limit and concurrency limit on the server, the tools' own included, and every throttle default. A tool's own timeout still applies:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { ExportTickets, LookupCustomer, SyncCrm } from "./tickets.tools";

@App({ id: "help-desk", name: "Help Desk", tools: [ExportTickets, SyncCrm, LookupCustomer] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: { enabled: false, defaultTimeout: { executeMs: 100 } }, // 🚩 for local testing only
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.


Troubleshooting

Rate limit exceeded. Retry after N seconds, and after N seconds it's refused again

The count is an estimate over two windows: when the window ends, the calls from the one before still count, less and less as time goes on. The first call after the wait gets through; the next may not, until enough of the old window has passed. If callers retry exactly on time, use a longer window with a higher maxRequests, or tell the model in the description to space its calls out.

Invalid input: expected boolean, received undefined at throttle.enabled

The server's class declaration throws this when throttle has no enabled. It's required, even when you only set defaults: add enabled: true. TypeScript reports it too.

Too small: expected number to be >0 at rateLimit.maxRequests

maxRequests, maxConcurrent, windowMs and executeMs must be positive integers; the tool fails when it's declared. To turn a tool off, remove it from its app, or use availableWhen. To lift its limit, remove rateLimit.

FrontMCP checks these options with the schemas it exports, so you can check a config the same way:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "search_tickets", description: "Search tickets by title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    return { query, tickets: [] };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The throttle defaults have no effect

defaultRateLimit, defaultConcurrency, defaultTimeout, global, globalConcurrency and ipFilter need enabled: true. With enabled: false, they're all ignored, and so are the tools' own limits except timeout. Also check that the tool doesn't set its own limit: that replaces the default.

Limits aren't shared between server instances

Each process keeps its own counts unless throttle.storage points them at a shared store. Check its shape: it's { type: "redis", redis: { url } } (or redis: { config: { host, port } }), not { provider: "redis", host, port }. A storage without type auto-detects from environment variables (REDIS_URL, REDIS_HOST, UPSTASH_REDIS_REST_URL, KV_REST_API_URL) and quietly uses memory if none is set, so the second form looks like it works. The top-level redis option isn't used by the guard. See storage.

The server doesn't start: throttle.storage (redis) is unavailable

throttle.storage points at a store the server can't reach. Rate limits are a security control, so they fail closed: in production, unless fallback says otherwise, the server refuses to start instead of running with limits that nobody shares. The error is a GuardStorageUnavailableError, with code GUARD_STORAGE_UNAVAILABLE, and its message ends with the way out: Set throttle.storage.fallback: 'memory' to start with per-instance counters instead. Fix the address or the credentials, or choose the fallback on purpose, knowing that each instance then counts alone. Before 1.8.6, the error was the storage client's own Failed to connect to Redis, with nothing about throttle.storage.

The tests use a Redis address that nothing listens on:

Open
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance, GuardStorageUnavailableError } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";

const info = { name: "help-desk", version: "1.0.0" };
const unreachableRedis = { type: "redis", redis: { config: { host: "127.0.0.1", port: 1 } } } as const;

test('`fallback: "error"` stops the server, and the error names `throttle.storage`', async () => {
  const starting = FrontMcpInstance.createDirect({
    info,
    apps: [HelpDesk],
    throttle: { enabled: true, storage: { ...unreachableRedis, fallback: "error" } },
  });
  await expect(starting).rejects.toBeInstanceOf(GuardStorageUnavailableError);
  await expect(starting).rejects.toMatchObject({ code: "GUARD_STORAGE_UNAVAILABLE", statusCode: 503, storageType: "redis" });
  await expect(starting).rejects.toThrow(/^throttle\.storage \(redis\) is unavailable: /);
  await expect(starting).rejects.toThrow("Rate limits fail closed");
});

test('`fallback: "memory"` starts the server with counts of its own', async () => {
  const server = await FrontMcpInstance.createDirect({
    info,
    apps: [HelpDesk],
    throttle: { enabled: true, storage: { ...unreachableRedis, fallback: "memory" } },
  });
  const result = await server.callTool("get_ticket", { id: "T-1" });
  await server.dispose();
  expect(result.isError).toBeFalsy();
});

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Calls fail with GUARD_STORAGE_UNAVAILABLE while the server runs

The store answered at startup and then stopped, like a Redis that went down. A call that needs a limit check or a concurrency slot can't be counted, so with fallback: "error", the default in production, it is refused: a GuardStorageUnavailableError, code GUARD_STORAGE_UNAVAILABLE, status 503. In development, its message says what failed: throttle.storage (redis) is unavailable: <the store's error>. Rate limits fail closed, so this call was refused. Set throttle.storage.fallback: 'memory' to keep serving with per-instance counters while it is down. In production a client reads Service temporarily unavailable: the rate-limit store cannot be reached. Retry shortly., which doesn't give away where the store is. With throttle.global, the whole request is refused instead, with HTTP 503 and Retry-After: 1. With fallback: "memory", the default outside production, the guard switches to counters of its own for each instance, logs it once, tries the store again every 30 seconds, and logs when it answers. Calls made while the store is down are counted per instance, so the limit can be reached on every instance in turn.

The tests pass their own Redis client, one that answers until a test takes it down:

Open
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance, GuardStorageUnavailableError, createErrorHandler, toMcpError } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";
import { FakeRedis } from "./redis.example";

const start = (redis: FakeRedis, fallback: "error" | "memory") =>
  FrontMcpInstance.createDirect({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDesk],
    throttle: { enabled: true, storage: { type: "redis", redis: { client: redis as never }, fallback } },
  });

test('`fallback: "error"` refuses calls while the store is down, and serves again when it answers', async () => {
  const redis = new FakeRedis();
  const server = await start(redis, "error");
  try {
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });

    redis.down = true;
    const refused = await server.callTool("get_ticket", { id: "T-1" }).catch((error) => error);
    expect(refused).toBeInstanceOf(GuardStorageUnavailableError);
    expect(refused).toMatchObject({ code: "GUARD_STORAGE_UNAVAILABLE", statusCode: 503, storageType: "redis" });
    expect(refused.message).toContain("Rate limits fail closed, so this call was refused.");
    expect(toMcpError(refused).isPublic).toBe(true);
    const production = createErrorHandler({ isDevelopment: false }).handle(refused);
    expect(production.content).toEqual([
      { type: "text", text: "Service temporarily unavailable: the rate-limit store cannot be reached. Retry shortly." },
    ]);

    redis.down = false;
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });
  } finally {
    await server.dispose();
  }
});

test("with `global`, the whole request gets 503 and Retry-After", async () => {
  const redis = new FakeRedis();
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDesk],
    throttle: { enabled: true, global: { maxRequests: 100, windowMs: 60_000 }, storage: { type: "redis", redis: { client: redis as never }, fallback: "error" } },
  });
  redis.down = true;
  const response = await handler(
    new Request("https://desk.example.com/", {
      method: "POST",
      headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/list" },
      body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/list", params: { _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } } }),
    }),
  );
  expect(response.status).toBe(503);
  expect(response.headers.get("retry-after")).toBe("1");
  expect(await response.json()).toMatchObject({
    error: { code: -32603, message: "Service temporarily unavailable: the rate-limit store cannot be reached", data: { code: "GUARD_STORAGE_UNAVAILABLE" } },
  });
});

test('`fallback: "memory"` keeps serving, with counts of its own', async () => {
  const redis = new FakeRedis();
  const server = await start(redis, "memory");
  try {
    redis.down = true;
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });
    await expect(server.callTool("get_ticket", { id: "T-1" })).rejects.toMatchObject({ code: "RATE_LIMIT_EXCEEDED" });
  } finally {
    await server.dispose();
  }
});

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

New in 1.8.7: this runtime case. Starting without the store was already handled in 1.8.6. Changed in 1.9: before, production clients read the development message, with the store's error and address, and with throttle.global the request failed with a bare 500.

Every request gets -32001 Forbidden: client IP rejected by ipFilter

ipFilter has defaultAction: "deny", and the requests' address isn't on the allow list: the server is behind a proxy it doesn't trust, so every request has the proxy's address, or it runs under createFetchHandler() in a runtime that doesn't report addresses, so it has none. Set FRONTMCP_TRUST_PROXY=true (and FRONTMCP_TRUSTED_PROXY_DEPTH for more than one proxy) so it reads X-Forwarded-For. ipFilter.trustProxy in the config does nothing.

An allowList doesn't keep other addresses out

An address on neither list gets defaultAction, which is "allow". Add defaultAction: "deny". See the test above.

Tool "…" execution failed: Concurrency limit reached for …, with TOOL_EXECUTION_ERROR

execute() threw a guard error of its own, like a ConcurrencyLimitError for a slot from a GuardManager. Thrown, only ExecutionTimeoutError keeps its code; the others are wrapped as TOOL_EXECUTION_ERROR and hidden in production. Pass the error to this.fail(), which keeps its code and message, or fail with RateLimitError, as in Limiting by an argument:

Open
import { ConcurrencyLimitError, createGuardManager, Tool, ToolContext } from "@frontmcp/sdk";

const limits = createGuardManager({ config: { enabled: true } });
const oneAtATime = { maxConcurrent: 1 };

@Tool({ name: "start_export", description: "Export tickets. One export at a time.", inputSchema: {} })
export class StartExport extends ToolContext {
  async execute() {
    const slot = await (await limits).acquireSemaphore("export", oneAtATime);
    if (!slot) throw new ConcurrencyLimitError("export", 1); // 🚩 thrown: TOOL_EXECUTION_ERROR
    await new Promise((resolve) => setTimeout(resolve, 200));
    await slot.release();
    return { exported: true };
  }
}

@Tool({ name: "start_import", description: "Import tickets. One import at a time.", inputSchema: {} })
export class StartImport extends ToolContext {
  async execute() {
    const slot = await (await limits).acquireSemaphore("import", oneAtATime);
    if (!slot) this.fail(new ConcurrencyLimitError("import", 1)); // ✅ keeps CONCURRENCY_LIMIT
    await new Promise((resolve) => setTimeout(resolve, 200));
    await slot.release();
    return { imported: true };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

One anonymous client used up everyone's limit

Under "userId" and "session", all anonymous callers share the key "anonymous". Have callers sign in, so each gets a count of their own, or, on a real server, count anonymous callers by address: partitionBy: (ctx) => ctx.userId ?? ctx.clientIp ?? "anonymous".