# Guard options

> Every FrontMCP option that limits tool calls, a tool's rateLimit, concurrency and timeout and @FrontMcp's throttle, whose calls each limit counts, and the error each one produces.

Source: https://frontmcp.dev/reference/sdk/guard

FrontMCP's guard limits tool calls: how often a tool runs (`rateLimit`), how many calls of it run at once (`concurrency`), and how long one call may take (`timeout`). A tool sets its own limits in `@Tool`. `@FrontMcp({ throttle })` sets defaults for every tool, adds limits for the server as a whole, and filters requests by IP address. A call over a limit never reaches `execute()`, and the caller gets an error that says which limit it hit. [Limiting and Timing Out Calls](https://frontmcp.dev/learn/limiting-calls) teaches the same options step by step.

```ts
@Tool({ name, inputSchema, rateLimit, concurrency, timeout })

@FrontMcp({
  info, apps,
  throttle: { enabled, defaultRateLimit, defaultConcurrency, defaultTimeout, global, globalConcurrency, ipFilter, storage, keyPrefix },
})
```

---

## Reference

### A tool's limits

Set `rateLimit`, `concurrency` and `timeout` in a tool's options. They work without any `throttle` on the server.

```ts export-tickets.tool.ts
import { Tool, ToolContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV. At most 2 exports an hour per caller, one at a time.",
  inputSchema: {},
  rateLimit: { maxRequests: 2, windowMs: 60 * 60_000, partitionBy: "userId" },
  concurrency: { maxConcurrent: 1, queueTimeoutMs: 5_000 },
  timeout: { executeMs: 30_000 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: await buildCsv({ signal: this.signal }) };
  }
}
```

[See more examples below.](#usage)

#### `rateLimit`

How many calls of the tool a window allows.

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `maxRequests` | `number` | required | Calls allowed per window. A positive integer. |
| `windowMs` | `number` | `60_000` | The window's length, in milliseconds. |
| `partitionBy` | `PartitionKey` | `"global"` | Whose calls share a count. See [`partitionBy`](#partitionby). |

A call over the limit fails with `RATE_LIMIT_EXCEEDED` and `Rate limit exceeded. Retry after N seconds`. Windows are aligned to the clock (an hourly window runs from one full UTC hour to the next), and N is the time left in the current one.

The count is an estimate over two windows: calls in the current window, plus the previous window's calls weighted by how much of the previous window is still within `windowMs` of now. So right after a busy window, fewer calls get through than `maxRequests`, and once "Retry after N seconds" has passed, one call can get through and the next be refused again. (That depends on the clock, so no example below shows it.)

#### `concurrency`

How many calls of the tool may run at the same moment.

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `maxConcurrent` | `number` | required | Calls allowed to run at once. A positive integer. |
| `queueTimeoutMs` | `number` | `0` | How long a call waits for a free slot, in milliseconds. `0` refuses it at once. |
| `partitionBy` | `PartitionKey` | `"global"` | Whose calls share the slots. See [`partitionBy`](#partitionby). |

A call that finds no free slot fails with `CONCURRENCY_LIMIT` and `Concurrency limit reached for "merge_customers" (max: 1)`. With a queue, it checks for a free slot every 100 ms at first, backing off to once a second, and fails with `QUEUE_TIMEOUT` and `Queue timeout for "merge_customers" after waiting 5000ms for a concurrency slot` if none frees up in time. A slot is taken before the input is validated and given back when `execute()` ends, whether it succeeded, failed or timed out.

#### `timeout`

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `executeMs` | `number` | required | How long `execute()` may run, in milliseconds. A positive integer. |

When time runs out, the call fails with `EXECUTION_TIMEOUT` and `Execution of "lookup_customer" timed out after 500ms`, and `this.signal` is aborted with that same error. `execute()` itself keeps running unless it listens to the signal, and it keeps its `concurrency` slot until it ends. See [A timeout ends the call, not the work](https://frontmcp.dev/learn/limiting-calls#a-timeout-ends-the-call-not-the-work).

#### `partitionBy`

`partitionBy` decides which calls share a count (for `rateLimit`) or a set of slots (for `concurrency`). FrontMCP turns it into a key, and calls with the same key share:

| `partitionBy` | Key |
| --- | --- |
| `"global"` | `"global"`: every call shares one count. |
| `"userId"` | The caller's `userId`, or `sessionId` when there's none. |
| `"session"` | `sessionId`. |
| `"ip"` | The client IP address; without one, `"user:"` and the caller's `userId`, or else `"ip:unresolved"`. |
| `(ctx) => string` | Whatever the function returns. |

The key is built from a `PartitionKeyContext`:

| Field | What it holds in FrontMCP 1.9.3 |
| --- | --- |
| `userId` | The signed-in caller's client id from [`authInfo`](https://frontmcp.dev/reference/sdk/context), like `"static:6ab9f1eb8f7d"` for a static token. Missing for anonymous callers. |
| `sessionId` | For MCP 2026-07-28 clients, which have no session: `"user:"` and the `userId`, or `"anonymous"`. For older clients in a session, over Streamable HTTP or the old SSE transport, the session id. (Before 1.8.5, SSE clients counted as having none.) |
| `clientIp` | The address the request came from. Behind a proxy, set the `FRONTMCP_TRUST_PROXY=true` environment variable (and `FRONTMCP_TRUSTED_PROXY_DEPTH` for more than one hop) to read `X-Forwarded-For` instead. Under `createFetchHandler()`, it's the address the runtime reports, when it reports one (Cloudflare Workers, Bun, Deno). The Playground has none. |

So under MCP 2026-07-28, `"userId"` and `"session"` both give each signed-in caller a count of their own, and **all anonymous callers share one**, the key `"anonymous"`. Nobody gets around a limit by not signing in, but one busy anonymous client can use up the limit for every other one. [See it run.](#choosing-whose-calls-count)

### `@FrontMcp({ throttle })`

`throttle` configures the guard for the whole server.

```ts main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: "userId" },
    defaultConcurrency: { maxConcurrent: 5 },
    defaultTimeout: { executeMs: 30_000 },
    global: { maxRequests: 1_000, windowMs: 60_000, partitionBy: "ip" },
    ipFilter: { denyList: ["192.0.2.0/24"] },
  },
})
export default class Server {}
```

| Option | Type | Default | Description |
| --- | --- | --- | --- |
| `enabled` | `boolean` | required | `true` turns every option below on. `false` turns the guard off: no rate limits or concurrency limits anywhere, including each tool's own. |
| `defaultRateLimit` | `RateLimitConfig` | none | A [`rateLimit`](#ratelimit) for every tool that doesn't set its own. Each tool counts separately. |
| `defaultConcurrency` | `ConcurrencyConfig` | none | A [`concurrency`](#concurrency) for every tool that doesn't set its own. Each tool has its own slots. |
| `defaultTimeout` | `TimeoutConfig` | none | A [`timeout`](#timeout) for every tool that doesn't set its own. |
| `global` | `RateLimitConfig` | none | One rate limit for every request the server gets, of any kind. See [`global`](#global). |
| `globalConcurrency` | `ConcurrencyConfig` | none | How many tool calls may run at once, across every tool. See [`globalConcurrency`](#globalconcurrency). |
| `ipFilter` | `IpFilterConfig` | none | Allow or refuse requests by client IP. See [`ipFilter`](#ipfilter). |
| `storage` | `StorageConfig` | memory | Where counts and slots are kept. See [`storage`](#storage-and-keyprefix). |
| `keyPrefix` | `string` | `"mcp:guard:"` | Prefix for every key the guard stores. |

A tool's own `rateLimit`, `concurrency` or `timeout` replaces the default: it doesn't add to it.

#### When the guard runs

| Server configuration | What's enforced |
| --- | --- |
| No `throttle` | Each tool's own `rateLimit`, `concurrency` and `timeout`. |
| `throttle: { enabled: true, … }` | Each tool's own limits, and every `throttle` option. |
| `throttle: { enabled: false, … }` | Only each tool's own `timeout`. Every other limit, and every `throttle` option, is ignored. |

Only tool calls are limited. Resources, prompts and other requests don't count toward a tool's `rateLimit` or a default one; `global` is the one limit that counts them. An agent is a tool to its clients, `invoke_<name>`, and `@Agent` takes the same three options, which limit it like any tool (since FrontMCP 1.8.3): see [Limiting who calls an agent, and how often](https://frontmcp.dev/reference/sdk/agent#limiting-who-calls-an-agent-and-how-often).

The limits are checked after the caller's [authorization](https://frontmcp.dev/learn/authorizing-calls) and before the arguments are validated, so a call with invalid arguments counts toward `rateLimit` like any other. [See it run.](#setting-defaults-for-every-tool)

#### `global`

A rate limit for the whole server, `{ maxRequests, windowMs?, partitionBy? }`, like [`rateLimit`](#ratelimit). Every request counts once, whatever its method: `tools/list` and `resources/read` as much as `tools/call`. It's checked when the request arrives, before FrontMCP knows which tool it's for, so a request over it isn't a tool error: it gets HTTP `429`, a `Retry-After` header with the seconds to wait, and a JSON-RPC error:

```json
{ "jsonrpc": "2.0", "id": 7, "error": { "code": -32029, "message": "Rate limit exceeded. Retry after 36 seconds" } }
```

With `partitionBy` `"global"` or `"ip"`, it's checked before authentication; with `"userId"`, `"session"` or a function, right after it, so the caller is known.

#### `globalConcurrency`

A cap on tool calls running at once across every tool, `{ maxConcurrent, queueTimeoutMs?, partitionBy? }`. It's checked before the tool's own `concurrency`. A call over it fails with `CONCURRENCY_LIMIT` and `Concurrency limit reached for "global" (max: 2)`. A tool that calls another with [`this.callTool()`](https://frontmcp.dev/reference/sdk/call-tool) doesn't take a second global slot for the inner call, so `maxConcurrent: 1` doesn't deadlock it.

#### `ipFilter`

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `allowList` | `string[]` | `[]` | IP addresses or CIDR ranges, IPv4 or IPv6, to allow. |
| `denyList` | `string[]` | `[]` | IP addresses or CIDR ranges to refuse. |
| `defaultAction` | `"allow" \| "deny"` | `"allow"` | What happens to an address on neither list, and to a request whose address isn't known. |
| `trustProxy` | `boolean` | `false` | **Not read.** Setting it logs a warning. Use the `FRONTMCP_TRUST_PROXY` environment variable. |
| `trustedProxyDepth` | `number` | `1` | **Not read.** Setting it logs a warning. Use `FRONTMCP_TRUSTED_PROXY_DEPTH`. |

Each request is checked before anything else: first the deny list, then the allow list, then `defaultAction`. An `allowList` on its own therefore refuses nothing: add `defaultAction: "deny"` to refuse everyone else. A refused request gets HTTP `403` and JSON-RPC error `-32001`, `Forbidden: client IP rejected by ipFilter`.

A request whose address isn't known gets `defaultAction`. With `"deny"`, that refuses every request from a runtime that doesn't report addresses, including the Playground. Behind a proxy that isn't trusted, every request has the proxy's address. [See it run.](#filtering-by-ip-address)

#### `storage` and `keyPrefix`

By default the guard keeps its counts and slots in memory, so each server process has its own, and FrontMCP logs a warning that memory storage isn't suitable for several instances. To share limits between instances, give it a store:

```ts
throttle: {
  enabled: true,
  storage: { type: "redis", redis: { url: process.env.REDIS_URL } },
  defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: "userId" },
},
```

`storage` is `@frontmcp/utils`' `StorageConfig`: `type` is `"memory"`, `"redis"`, `"upstash"`, `"vercel-kv"`, `"cloudflare-kv"` or `"auto"`, with that backend's options under its own key (`redis: { url }` or `redis: { config: { host, port, password, db, tls } }`, `upstash`, `vercelKv`, `cloudflareKv`). A `storage` without `type` is `"auto"`: Upstash if `UPSTASH_REDIS_REST_URL` and `UPSTASH_REDIS_REST_TOKEN` are set, Vercel KV for `KV_REST_API_URL` and `KV_REST_API_TOKEN`, Redis for `REDIS_URL` or `REDIS_HOST`, and memory otherwise. If the store can't be reached, the server fails to start in production (`NODE_ENV=production`), with a `GuardStorageUnavailableError` that names `throttle.storage`, and falls back to memory, with a warning, otherwise. Rate limits fail closed on purpose: `fallback: "memory"` inside `storage` starts the server anyway, with counts of its own for each instance, and `fallback: "error"` refuses to start in every environment. [See both.](#the-server-doesnt-start-throttlestorage-redis-is-unavailable) The same choice applies if the store stops answering while the server runs: [see that too.](#calls-fail-with-guard_storage_unavailable-while-the-server-runs) The top-level `redis` option isn't used by the guard. The Playground can't reach Redis, so a shared store runs on a real server only.

`keyPrefix` is put before every key the guard stores, like `mcp:guard:export_tickets:global:rl:…` for a rate count. Give servers that share a store different prefixes if their limits should be separate. Before 1.8.6 the default prefix left an empty segment, `mcp:guard::export_tickets:…`, so an instance on 1.8.6 or later doesn't read the counts an older one wrote, and limits briefly split during a rolling deploy.

### What a caller gets back

A tool's limits, and every default, answer with an ordinary tool error: an `isError` result over HTTP `200`, with no `Retry-After` header. The code is in `_meta.code`, and the message is kept in production. `global` and `ipFilter` answer the whole request instead:

| Limit | Answer | Code | Message |
| --- | --- | --- | --- |
| `rateLimit`, `defaultRateLimit` | Tool error | `RATE_LIMIT_EXCEEDED` | `Rate limit exceeded. Retry after 1200 seconds` |
| `concurrency`, `defaultConcurrency` | Tool error | `CONCURRENCY_LIMIT` | `Concurrency limit reached for "merge_customers" (max: 1)` |
| `globalConcurrency` | Tool error | `CONCURRENCY_LIMIT` | `Concurrency limit reached for "global" (max: 2)` |
| Any of them, with `queueTimeoutMs` | Tool error | `QUEUE_TIMEOUT` | `Queue timeout for "merge_customers" after waiting 5000ms for a concurrency slot` |
| `timeout`, `defaultTimeout` | Tool error | `EXECUTION_TIMEOUT` | `Execution of "lookup_customer" timed out after 500ms` |
| `global` | HTTP `429`, `Retry-After` | JSON-RPC `-32029` | `Rate limit exceeded. Retry after 36 seconds` |
| `ipFilter` | HTTP `403` | JSON-RPC `-32001` | `Forbidden: client IP rejected by ipFilter` |
| A tool's limit, while `throttle.storage` is down | Tool error | `GUARD_STORAGE_UNAVAILABLE` | In production, `Service temporarily unavailable: the rate-limit store cannot be reached. Retry shortly.` See [while the server runs](#calls-fail-with-guard_storage_unavailable-while-the-server-runs). |
| `global`, while `throttle.storage` is down | HTTP `503`, `Retry-After: 1` | JSON-RPC `-32603`, `data.code` `GUARD_STORAGE_UNAVAILABLE` | `Service temporarily unavailable: the rate-limit store cannot be reached` |

The [error classes](https://frontmcp.dev/reference/sdk/error-classes#rate-limits-and-the-guard) page lists the classes behind these codes.

### Building blocks

`@frontmcp/sdk` re-exports the guard's own library, `@frontmcp/guard`, for limits the options can't express, like a limit per ticket rather than per caller:

| Export | What it does |
| --- | --- |
| `createGuardManager({ config, logger? })` | Returns a `Promise<GuardManager>` for a `GuardConfig` (`enabled` is required but not read). Memory storage unless `config.storage` says otherwise. |
| `GuardManager` | `checkRateLimit(name, config?, ctx?)` → `{ allowed, remaining, resetMs, retryAfterMs? }`, counting the call when it's allowed. `checkGlobalRateLimit(ctx?)`. `acquireSemaphore(name, config?, ctx?)` and `acquireGlobalSemaphore(ctx?)` → a ticket with `release()`, or `null` when no slot is free. `checkIpFilter(ip?)`, `isIpAllowListed(ip?)`, `destroy()`. Without a `config`, each falls back to the manager's defaults. |
| `withTimeout(fn, ms, name)` | Runs `fn(signal)` and rejects with `ExecutionTimeoutError` after `ms`, aborting `signal`. |
| `IpFilter` | `new IpFilter(ipFilterConfig).check(ip)` → `{ allowed, reason, matchedRule? }`, where `reason` is `"denylisted"`, `"allowlisted"` or `"default"`. |
| `SlidingWindowRateLimiter`, `DistributedSemaphore` | The rate counter and the slot pool the manager uses, over a storage adapter. |
| `resolvePartitionKey(partitionBy, ctx)`, `buildStorageKey(name, key, suffix?)` | How keys are made, as in the [`partitionBy`](#partitionby) table. |
| `guardConfigSchema`, `rateLimitConfigSchema`, `concurrencyConfigSchema`, `timeoutConfigSchema`, `ipFilterConfigSchema`, `partitionKeySchema` | The Zod schemas FrontMCP validates these options with. `ThrottleConfig` is the type of `throttle`. |
| `GuardError` and its subclasses | `ExecutionTimeoutError`, `ConcurrencyLimitError`, `QueueTimeoutError`, `IpBlockedError`, `IpNotAllowedError`, each with `code` and `statusCode`. |

To refuse a call from your own limit, fail with `RateLimitError` (a public error, `RATE_LIMIT_EXCEEDED`), or pass a guard error to [`this.fail()`](https://frontmcp.dev/reference/sdk/fail). Don't `throw` a guard error from `execute()`: only `ExecutionTimeoutError` keeps its code that way; the others arrive as `TOOL_EXECUTION_ERROR`, hidden in production. [See it run.](#limiting-by-an-argument)

#### Caveats

- **`enabled` is required** in a `throttle` object, even to set only defaults, and the defaults, `global`, `globalConcurrency` and `ipFilter` only work with `enabled: true`.
- **`enabled: false` turns off each tool's own limits too**, all but `timeout`. Keep it for local testing.
- **Counts live in memory by default**, one set per process. Several instances, or serverless functions, need [`storage`](#storage-and-keyprefix).
- An in-process server from [`create()`](https://frontmcp.dev/reference/sdk/create#keeping-the-servers-options) takes `throttle` as `@FrontMcp` does.
- **Anonymous callers share one count** under `"userId"` and `"session"` (see [`partitionBy`](#partitionby)).
- **A timeout doesn't stop `execute()`**. Pass `this.signal` on to the work, and don't give a tool that changes things a timeout shorter than its work.
- **Say the limit in the tool's description**, so the model can plan around it instead of learning it from an error.
- **In a browser, calls may not overlap**: without the TC39 `AsyncContext`, FrontMCP's browser build serves [one request at a time](https://frontmcp.dev/reference/sdk/create-fetch-handler#in-a-browser), so `concurrency` and `globalConcurrency` have nothing to limit there. The Playground provides a minimal `AsyncContext`, so its calls overlap as they do on Node.

---

## Usage

### Limiting one tool

<Examples title="A tool's limits">

#### Example: Rate limit
At most two exports an hour. The third call is refused, with the time until the window ends.

```ts export-tickets.tool.ts active
import { Tool, ToolContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV. At most 2 exports an hour.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
  rateLimit: { maxRequests: 2, windowMs: 60 * 60_000 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title,status\nT-1,Cannot log in,open\nT-2,Invoice total is wrong,closed" };
  }
}
```

```ts export.test.ts
import { test, expect } from "@frontmcp/testing";

test("the third export in an hour is refused", async ({ mcp }) => {
  expect(await mcp.tools.call("export_tickets", {})).toBeSuccessful();
  expect(await mcp.tools.call("export_tickets", {})).toBeSuccessful();

  const third = await mcp.tools.call("export_tickets", {});
  expect(third).toBeError("RATE_LIMIT_EXCEEDED");
  expect(third.text()).toMatch(/^Rate limit exceeded\. Retry after \d+ seconds$/);
});
```

#### Example: Concurrency
One merge at a time, and a merge that finds another running waits up to a second for it. Three merges sent together: the first runs, one of the other two gets the slot when it frees up, and the last runs out of time.

```ts merge-customers.tool.ts active
import { Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({
  name: "merge_customers",
  description: "Merge one customer record into another. One merge runs at a time; another waits up to a second for its turn.",
  inputSchema: { from: z.string(), into: z.string() },
  annotations: { destructiveHint: true },
  concurrency: { maxConcurrent: 1, queueTimeoutMs: 1_000 },
})
export class MergeCustomers extends ToolContext {
  async execute({ from, into }: { from: string; into: string }) {
    await new Promise((resolve) => setTimeout(resolve, 600)); // moving tickets takes a while
    return { merged: from, into };
  }
}
```

```ts merge.test.ts
import { test, expect } from "@frontmcp/testing";

test("two of three merges run, one after the other; the third runs out of time", async ({ mcp }) => {
  const results = await Promise.all(
    ["C-7", "C-8", "C-9"].map((from) => mcp.tools.call("merge_customers", { from, into: "C-2" })),
  );
  const codes = results.map((r) => r.raw._meta?.code ?? "ok");
  expect(codes.filter((c) => c === "ok")).toHaveLength(2);
  expect(codes.filter((c) => c === "QUEUE_TIMEOUT")).toHaveLength(1);

  const timedOut = results.find((r) => r.raw._meta?.code === "QUEUE_TIMEOUT")!;
  expect(timedOut.text()).toBe('Queue timeout for "merge_customers" after waiting 1000ms for a concurrency slot');
});
```

#### Example: Timeout
The CRM sometimes takes two seconds. The tool gives up after half a second, and passes `this.signal` on, so the request stops too.

```ts lookup-customer.tool.ts active
import { Tool, ToolContext, z } from "@frontmcp/sdk";
import { crmLookup } from "./crm";

@Tool({
  name: "lookup_customer",
  description: "Look up a customer in the CRM. Gives up after half a second if the CRM is slow.",
  inputSchema: { id: z.string() },
  annotations: { readOnlyHint: true, openWorldHint: true },
  timeout: { executeMs: 500 },
})
export class LookupCustomer extends ToolContext {
  async execute({ id }: { id: string }) {
    return crmLookup(id, this.signal);
  }
}
```

```ts crm.ts
export const cancelled: string[] = [];

// Stands in for a CRM client that takes an AbortSignal, like fetch() does.
export function crmLookup(id: string, signal?: AbortSignal) {
  return new Promise<{ id: string; name: string }>((resolve, reject) => {
    const timer = setTimeout(() => resolve({ id, name: "Globex" }), id === "C-9" ? 2_000 : 50);
    signal?.addEventListener("abort", () => {
      clearTimeout(timer);
      cancelled.push(`${id}: ${signal.reason.message}`);
      reject(signal.reason);
    });
  });
}
```

```ts lookup.test.ts
import { test, expect } from "@frontmcp/testing";
import { cancelled } from "./crm";

test("a slow lookup times out, and its request is cancelled", async ({ mcp }) => {
  const result = await mcp.tools.call("lookup_customer", { id: "C-9" });
  expect(result).toBeError("EXECUTION_TIMEOUT");
  expect(result.text()).toBe('Execution of "lookup_customer" timed out after 500ms');
  expect(cancelled).toEqual(['C-9: Execution of "lookup_customer" timed out after 500ms']);
});

test("a fast one answers", async ({ mcp }) => {
  expect(await mcp.tools.call("lookup_customer", { id: "C-1" })).toBeSuccessful();
});
```

### Setting defaults for every tool

With `enabled: true`, `throttle` gives every tool the limits it doesn't set itself. Each tool counts on its own: `search_tickets` and `get_ticket` get five calls an hour each, not five between them. `export_tickets` sets its own `rateLimit`, which replaces the default. `sync_crm` runs into the default timeout, and `weekly_report` into the default concurrency:

```ts main.ts active
import { App, FrontMcp } from "@frontmcp/sdk";
import { ExportTickets, GetTicket, SearchTickets, SyncCrm, WeeklyReport } from "./tickets.tools";

@App({ id: "help-desk", name: "Help Desk", tools: [SearchTickets, GetTicket, ExportTickets, SyncCrm, WeeklyReport] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    defaultRateLimit: { maxRequests: 5, windowMs: 60 * 60_000 },
    defaultConcurrency: { maxConcurrent: 1 },
    defaultTimeout: { executeMs: 300 },
  },
})
export default class Server {}
```

```ts tickets.tools.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";

const wait = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms));

@Tool({ name: "search_tickets", description: "Search tickets by title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    return { query, tickets: [{ id: "T-1", title: "Cannot log in" }] };
  }
}

@Tool({ name: "get_ticket", description: "Get one ticket", inputSchema: { id: z.string() } })
export class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in" };
  }
}

@Tool({
  name: "export_tickets",
  description: "Export every ticket as CSV. Once an hour.",
  inputSchema: {},
  rateLimit: { maxRequests: 1, windowMs: 60 * 60_000 }, // replaces defaultRateLimit
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title\nT-1,Cannot log in" };
  }
}

@Tool({ name: "sync_crm", description: "Copy tickets to the CRM", inputSchema: {} })
export class SyncCrm extends ToolContext {
  async execute() {
    await wait(1_000);
    return { synced: true };
  }
}

@Tool({ name: "weekly_report", description: "Ticket counts for the last week", inputSchema: {} })
export class WeeklyReport extends ToolContext {
  async execute() {
    await wait(200);
    return { opened: 41, closed: 38 };
  }
}
```

```ts defaults.test.ts
import { test, expect } from "@frontmcp/testing";

test("each tool has its own count of five", async ({ mcp }) => {
  for (let i = 0; i < 5; i++) expect(await mcp.tools.call("search_tickets", { query: "log" })).toBeSuccessful();
  expect(await mcp.tools.call("search_tickets", { query: "log" })).toBeError("RATE_LIMIT_EXCEEDED");
  expect(await mcp.tools.call("get_ticket", { id: "T-1" })).toBeSuccessful();
});

test("a tool's own rateLimit replaces the default", async ({ mcp }) => {
  expect(await mcp.tools.call("export_tickets", {})).toBeSuccessful();
  expect(await mcp.tools.call("export_tickets", {})).toBeError("RATE_LIMIT_EXCEEDED");
});

test("defaultTimeout applies to tools without a timeout", async ({ mcp }) => {
  const result = await mcp.tools.call("sync_crm", {});
  expect(result).toBeError("EXECUTION_TIMEOUT");
  expect(result.text()).toBe('Execution of "sync_crm" timed out after 300ms');
});

test("calls with invalid arguments count too", async ({ mcp }) => {
  // get_ticket has been called once, above.
  for (let i = 0; i < 4; i++) expect(await mcp.tools.call("get_ticket", { id: 5 })).toBeError("INVALID_INPUT");
  expect(await mcp.tools.call("get_ticket", { id: "T-1" })).toBeError("RATE_LIMIT_EXCEEDED");
});

test("other requests don't count toward a tool's limit", async ({ mcp }) => {
  for (let i = 0; i < 10; i++) await mcp.tools.list();
  expect(await mcp.tools.call("weekly_report", {})).toBeSuccessful();
});

test("defaultConcurrency lets one call of each tool run at a time", async ({ mcp }) => {
  const [first, second] = await Promise.all([mcp.tools.call("weekly_report", {}), mcp.tools.call("weekly_report", {})]);
  expect(first).toBeSuccessful();
  expect(second).toBeError("CONCURRENCY_LIMIT");
  expect(second.text()).toBe('Concurrency limit reached for "weekly_report" (max: 1)');
});
```

### Limiting the whole server

<Examples title="Server-wide limits">

#### Example: Requests
`global` counts every request, whatever it is. Here the server takes ten an hour, a tiny limit to make the point: the Playground used six when it started (five discovery and list requests, and the call), so press **Call** until it's refused. In the first test, nine lists and a call get through, and the eleventh request is refused before FrontMCP looks at what it is. The second test sends raw HTTP, to see the status and the `Retry-After` header, which `mcp` doesn't show.

```ts main.ts active
import { App, FrontMcp } from "@frontmcp/sdk";
import { SearchTickets } from "./search-tickets.tool";

@App({ id: "help-desk", name: "Help Desk", tools: [SearchTickets] })
export class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: { enabled: true, global: { maxRequests: 10, windowMs: 60 * 60_000 } },
})
export default class Server {}
```

```ts search-tickets.tool.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "search_tickets", description: "Search tickets by title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    return { query, tickets: [] };
  }
}
```

```ts send.ts
// Sends one MCP 2026-07-28 request to a fetch handler, as a client would.
export async function send(
  handler: (request: Request) => Promise<Response>,
  method: string,
  params: Record<string, unknown> = {},
  headers: Record<string, string> = {},
) {
  const response = await handler(
    new Request("https://desk.example.com/", {
      method: "POST",
      headers: {
        "content-type": "application/json",
        accept: "application/json, text/event-stream",
        "mcp-protocol-version": "2026-07-28",
        "mcp-method": method,
        ...(typeof params.name === "string" ? { "mcp-name": params.name } : {}),
        ...headers,
      },
      body: JSON.stringify({
        jsonrpc: "2.0",
        id: 1,
        method,
        params: { ...params, _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } },
      }),
    }),
  );
  return { status: response.status, retryAfter: response.headers.get("retry-after"), body: await response.json() };
}
```

```ts global.test.ts
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDeskApp } from "./main";
import { send } from "./send";

test("the eleventh request of any kind is refused", async ({ mcp }) => {
  for (let i = 0; i < 5; i++) await mcp.tools.list();
  for (let i = 0; i < 4; i++) await mcp.resources.list();
  expect(await mcp.tools.call("search_tickets", { query: "log" })).toBeSuccessful();

  const eleventh = await mcp.tools.call("search_tickets", { query: "log" });
  expect(eleventh.error).toMatchObject({ code: -32029, message: expect.stringMatching(/^Rate limit exceeded\. Retry after \d+ seconds$/) });
});

test("over HTTP, it's a 429 with Retry-After", async () => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDeskApp],
    throttle: { enabled: true, global: { maxRequests: 1, windowMs: 60_000 } },
  });
  expect((await send(handler, "tools/list")).status).toBe(200);

  const refused = await send(handler, "tools/list");
  expect(refused.status).toBe(429);
  expect(refused.retryAfter).toMatch(/^\d+$/);
  expect(refused.body.error).toEqual({ code: -32029, message: `Rate limit exceeded. Retry after ${refused.retryAfter} seconds` });
});
```

#### Example: Concurrency
`globalConcurrency` caps tool calls running at once, whichever tools they are. Two reports run; a third call, to a different tool, is refused. `report_digest` calls both reports with `this.callTool()`, and those inner calls don't take slots of their own.

```ts main.ts active
import { App, FrontMcp } from "@frontmcp/sdk";
import { MonthlyReport, ReportDigest, SearchTickets, WeeklyReport } from "./reports.tools";

@App({ id: "help-desk", name: "Help Desk", tools: [WeeklyReport, MonthlyReport, ReportDigest, SearchTickets] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: { enabled: true, globalConcurrency: { maxConcurrent: 2 } },
})
export default class Server {}
```

```ts reports.tools.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";

const slowQuery = () => new Promise((resolve) => setTimeout(resolve, 300));

@Tool({ name: "weekly_report", description: "Ticket counts for the last week", inputSchema: {} })
export class WeeklyReport extends ToolContext {
  async execute() {
    await slowQuery();
    return { opened: 41, closed: 38 };
  }
}

@Tool({ name: "monthly_report", description: "Ticket counts for the last month", inputSchema: {} })
export class MonthlyReport extends ToolContext {
  async execute() {
    await slowQuery();
    return { opened: 170, closed: 162 };
  }
}

@Tool({ name: "report_digest", description: "Weekly and monthly ticket counts together", inputSchema: {} })
export class ReportDigest extends ToolContext {
  async execute() {
    const weekly = await this.callTool("weekly_report", {});
    const monthly = await this.callTool("monthly_report", {});
    return { weekly: weekly.structuredContent, monthly: monthly.structuredContent };
  }
}

@Tool({ name: "search_tickets", description: "Search tickets by title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    return { query, tickets: [] };
  }
}
```

```ts concurrency.test.ts
import { test, expect } from "@frontmcp/testing";

test("a third call at once is refused, whichever tool it's for", async ({ mcp }) => {
  const [weekly, monthly, search] = await Promise.all([
    mcp.tools.call("weekly_report", {}),
    mcp.tools.call("monthly_report", {}),
    mcp.tools.call("search_tickets", { query: "log" }),
  ]);
  expect(weekly).toBeSuccessful();
  expect(monthly).toBeSuccessful();
  expect(search).toBeError("CONCURRENCY_LIMIT");
  expect(search.text()).toBe('Concurrency limit reached for "global" (max: 2)');
});

test("a tool's this.callTool() calls don't take global slots", async ({ mcp }) => {
  const [digest, weekly] = await Promise.all([mcp.tools.call("report_digest", {}), mcp.tools.call("weekly_report", {})]);
  expect(weekly).toBeSuccessful();
  expect(digest.json()).toEqual({ weekly: { opened: 41, closed: 38 }, monthly: { opened: 170, closed: 162 } });
});
```

### Choosing whose calls count

With `partitionBy: "userId"`, each signed-in caller gets a count of their own, and every anonymous caller shares one. The Playground's client is anonymous, so the first test shows the shared count. The second builds the same server with [static API keys](https://frontmcp.dev/learn/authenticating-clients) and calls it as two different callers. The `partitionBy` function in `get_ticket` records the context it's given, to show what a function can key on:

```ts tickets.tools.ts active
import { Tool, ToolContext, z, type PartitionKeyContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every ticket as CSV. Once an hour per signed-in caller; once an hour for all anonymous callers together.",
  inputSchema: {},
  rateLimit: { maxRequests: 1, windowMs: 60 * 60_000, partitionBy: "userId" },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title\nT-1,Cannot log in" };
  }
}

export const seen: PartitionKeyContext[] = [];

@Tool({
  name: "get_ticket",
  description: "Get one ticket",
  inputSchema: { id: z.string() },
  rateLimit: {
    maxRequests: 100,
    partitionBy: (ctx) => {
      seen.push(ctx);
      return ctx.userId ?? ctx.clientIp ?? "anonymous";
    },
  },
})
export class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in" };
  }
}
```

```ts main.ts
import { App, FrontMcp } from "@frontmcp/sdk";
import { ExportTickets, GetTicket } from "./tickets.tools";

@App({ id: "help-desk", name: "Help Desk", tools: [ExportTickets, GetTicket] })
export class HelpDeskApp {}

@FrontMcp({ info: { name: "help-desk", version: "1.0.0" }, apps: [HelpDeskApp] })
export default class Server {}
```

```ts send.ts
// Sends one MCP 2026-07-28 request to a fetch handler, as a client would.
export async function send(
  handler: (request: Request) => Promise<Response>,
  method: string,
  params: Record<string, unknown> = {},
  headers: Record<string, string> = {},
) {
  const response = await handler(
    new Request("https://desk.example.com/", {
      method: "POST",
      headers: {
        "content-type": "application/json",
        accept: "application/json, text/event-stream",
        "mcp-protocol-version": "2026-07-28",
        "mcp-method": method,
        ...(typeof params.name === "string" ? { "mcp-name": params.name } : {}),
        ...headers,
      },
      body: JSON.stringify({
        jsonrpc: "2.0",
        id: 1,
        method,
        params: { ...params, _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } },
      }),
    }),
  );
  return { status: response.status, retryAfter: response.headers.get("retry-after"), body: await response.json() };
}
```

```ts partition.test.ts
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDeskApp } from "./main";
import { seen } from "./tickets.tools";
import { send } from "./send";

test("anonymous callers share one count", async ({ mcp }) => {
  expect(await mcp.tools.call("export_tickets", {})).toBeSuccessful();
  expect(await mcp.tools.call("export_tickets", {})).toBeError("RATE_LIMIT_EXCEEDED");
});

test("an anonymous caller's context has only a sessionId", async ({ mcp }) => {
  await mcp.tools.call("get_ticket", { id: "T-1" });
  expect(seen[seen.length - 1]).toEqual({ sessionId: "anonymous" });
});

test("each signed-in caller has a count of their own", async () => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDeskApp],
    auth: { mode: "static", tokens: ["key-ana", "key-ben"] },
  });
  const exportAs = async (key: string) =>
    (await send(handler, "tools/call", { name: "export_tickets", arguments: {} }, { authorization: `Bearer ${key}` })).body.result;

  expect((await exportAs("key-ana")).isError).toBeUndefined();
  expect((await exportAs("key-ana"))._meta.code).toBe("RATE_LIMIT_EXCEEDED");
  expect((await exportAs("key-ben")).isError).toBeUndefined();

  await send(handler, "tools/call", { name: "get_ticket", arguments: { id: "T-1" } }, { authorization: "Bearer key-ana" });
  const ctx = seen[seen.length - 1];
  expect(ctx.userId).toMatch(/^static:/);
  expect(ctx.sessionId).toBe(`user:${ctx.userId}`);
});
```

### Filtering by IP address

`ipFilter` checks each request's address before anything else. The Playground's requests have no address, so a deny list lets them through (they get `defaultAction`, `"allow"`), and the tests use `IpFilter`, the class behind the option, to show how addresses are matched. The last test builds the server with `defaultAction: "deny"`, which refuses every request whose address isn't known:

```ts main.ts active
import { App, FrontMcp } from "@frontmcp/sdk";
import { SearchTickets } from "./search-tickets.tool";

@App({ id: "help-desk", name: "Help Desk", tools: [SearchTickets] })
export class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    ipFilter: { denyList: ["192.0.2.0/24", "2001:db8::/32"] },
  },
})
export default class Server {}
```

```ts search-tickets.tool.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "search_tickets", description: "Search tickets by title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    return { query, tickets: [] };
  }
}
```

```ts send.ts
// Sends one MCP 2026-07-28 request to a fetch handler, as a client would.
export async function send(
  handler: (request: Request) => Promise<Response>,
  method: string,
  params: Record<string, unknown> = {},
  headers: Record<string, string> = {},
) {
  const response = await handler(
    new Request("https://desk.example.com/", {
      method: "POST",
      headers: {
        "content-type": "application/json",
        accept: "application/json, text/event-stream",
        "mcp-protocol-version": "2026-07-28",
        "mcp-method": method,
        ...(typeof params.name === "string" ? { "mcp-name": params.name } : {}),
        ...headers,
      },
      body: JSON.stringify({
        jsonrpc: "2.0",
        id: 1,
        method,
        params: { ...params, _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } },
      }),
    }),
  );
  return { status: response.status, retryAfter: response.headers.get("retry-after"), body: await response.json() };
}
```

```ts ip-filter.test.ts
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance, IpFilter } from "@frontmcp/sdk";
import { HelpDeskApp } from "./main";
import { send } from "./send";

test("the deny list comes first, then the allow list, then defaultAction", () => {
  const filter = new IpFilter({ allowList: ["10.0.0.0/8"], denyList: ["10.0.9.0/24"], defaultAction: "deny" });
  expect(filter.check("10.0.9.4")).toEqual({ allowed: false, reason: "denylisted", matchedRule: "10.0.9.0/24" });
  expect(filter.check("10.2.0.1")).toEqual({ allowed: true, reason: "allowlisted", matchedRule: "10.0.0.0/8" });
  expect(filter.check("203.0.113.5")).toEqual({ allowed: false, reason: "default" });
});

test("an allow list alone refuses nobody", () => {
  const filter = new IpFilter({ allowList: ["10.0.0.0/8"] });
  expect(filter.check("203.0.113.5")).toEqual({ allowed: true, reason: "default" });
});

test("IPv6 ranges, and IPv4 addresses written as IPv6, match too", () => {
  const filter = new IpFilter({ denyList: ["2001:db8::/32", "192.0.2.0/24"] });
  expect(filter.check("2001:db8::17").allowed).toBe(false);
  expect(filter.check("::ffff:192.0.2.9").allowed).toBe(false);
});

test("with defaultAction: 'deny', a request with no known address is refused", async () => {
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDeskApp],
    throttle: { enabled: true, ipFilter: { allowList: ["10.0.0.0/8"], defaultAction: "deny" } },
  });
  const refused = await send(handler, "tools/list");
  expect(refused.status).toBe(403);
  expect(refused.body.error).toEqual({ code: -32001, message: "Forbidden: client IP rejected by ipFilter" });
});
```

On a real server behind a load balancer, set `FRONTMCP_TRUST_PROXY=true` so FrontMCP reads the address from `X-Forwarded-For`; otherwise every request has the load balancer's address, or none.

### Limiting by an argument

`partitionBy` sees the caller, not the arguments. For "each ticket can be escalated once an hour", keep a `GuardManager` of your own and key it by ticket. Fail with `RateLimitError`, so the model reads the same message and code as for any rate limit:

```ts escalate-ticket.tool.ts active
import { createGuardManager, RateLimitError, Tool, ToolContext, z } from "@frontmcp/sdk";

const limits = createGuardManager({ config: { enabled: true } });

@Tool({
  name: "escalate_ticket",
  description: "Page the on-call engineer about a ticket. Each ticket can be escalated once an hour.",
  inputSchema: { id: z.string().describe("Ticket id, like T-1") },
})
export class EscalateTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    const check = await (await limits).checkRateLimit(`escalate:${id}`, { maxRequests: 1, windowMs: 60 * 60_000 });
    if (!check.allowed) this.fail(new RateLimitError(Math.ceil(check.retryAfterMs! / 1000)));
    return { id, escalated: true };
  }
}
```

```ts escalate.test.ts
import { test, expect } from "@frontmcp/testing";

test("each ticket can be escalated once an hour", async ({ mcp }) => {
  expect(await mcp.tools.call("escalate_ticket", { id: "T-1" })).toBeSuccessful();

  const again = await mcp.tools.call("escalate_ticket", { id: "T-1" });
  expect(again).toBeError("RATE_LIMIT_EXCEEDED");
  expect(again.text()).toMatch(/^Rate limit exceeded\. Retry after \d+ seconds$/);

  expect(await mcp.tools.call("escalate_ticket", { id: "T-2" })).toBeSuccessful();
});
```

The manager above keeps its counts in memory, per process. Give it the same `storage` as `throttle` to share them.

### Turning the guard off

`throttle: { enabled: false }` switches off every rate limit and concurrency limit on the server, the tools' own included, and every `throttle` default. A tool's own `timeout` still applies:

```ts main.ts active
import { App, FrontMcp } from "@frontmcp/sdk";
import { ExportTickets, LookupCustomer, SyncCrm } from "./tickets.tools";

@App({ id: "help-desk", name: "Help Desk", tools: [ExportTickets, SyncCrm, LookupCustomer] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: { enabled: false, defaultTimeout: { executeMs: 100 } }, // 🚩 for local testing only
})
export default class Server {}
```

```ts tickets.tools.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";

const slowly = () => new Promise((resolve) => setTimeout(resolve, 300));

@Tool({
  name: "export_tickets",
  description: "Export every ticket as CSV. Once an hour.",
  inputSchema: {},
  rateLimit: { maxRequests: 1, windowMs: 60 * 60_000 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title\nT-1,Cannot log in" };
  }
}

@Tool({ name: "sync_crm", description: "Copy tickets to the CRM", inputSchema: {} })
export class SyncCrm extends ToolContext {
  async execute() {
    await slowly();
    return { synced: true };
  }
}

@Tool({ name: "lookup_customer", description: "Look up a customer", inputSchema: { id: z.string() }, timeout: { executeMs: 100 } })
export class LookupCustomer extends ToolContext {
  async execute({ id }: { id: string }) {
    await slowly();
    return { id, name: "Globex" };
  }
}
```

```ts off.test.ts
import { test, expect } from "@frontmcp/testing";

test("the tool's own rateLimit is ignored", async ({ mcp }) => {
  for (let i = 0; i < 3; i++) expect(await mcp.tools.call("export_tickets", {})).toBeSuccessful();
});

test("so is defaultTimeout", async ({ mcp }) => {
  expect(await mcp.tools.call("sync_crm", {})).toBeSuccessful();
});

test("a tool's own timeout still applies", async ({ mcp }) => {
  expect(await mcp.tools.call("lookup_customer", { id: "C-9" })).toBeError("EXECUTION_TIMEOUT");
});
```

---

## Troubleshooting

### `Rate limit exceeded. Retry after N seconds`, and after N seconds it's refused again

The count is an [estimate over two windows](#ratelimit): when the window ends, the calls from the one before still count, less and less as time goes on. The first call after the wait gets through; the next may not, until enough of the old window has passed. If callers retry exactly on time, use a longer window with a higher `maxRequests`, or tell the model in the description to space its calls out.

### `Invalid input: expected boolean, received undefined` at `throttle.enabled`

The server's class declaration throws this when `throttle` has no `enabled`. It's required, even when you only set defaults: add `enabled: true`. TypeScript reports it too.

### `Too small: expected number to be >0` at `rateLimit.maxRequests`

`maxRequests`, `maxConcurrent`, `windowMs` and `executeMs` must be positive integers; the tool fails when it's declared. To turn a tool off, remove it from its app, or use [`availableWhen`](https://frontmcp.dev/reference/sdk/tool). To lift its limit, remove `rateLimit`.

FrontMCP checks these options with the schemas it exports, so you can check a config the same way:

```ts search-tickets.tool.ts
import { Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "search_tickets", description: "Search tickets by title", inputSchema: { query: z.string() } })
export class SearchTickets extends ToolContext {
  async execute({ query }: { query: string }) {
    return { query, tickets: [] };
  }
}
```

```ts config.test.ts
import { test, expect } from "@frontmcp/testing";
import { guardConfigSchema, rateLimitConfigSchema } from "@frontmcp/sdk";

test("throttle needs enabled", () => {
  const result = guardConfigSchema.safeParse({ defaultRateLimit: { maxRequests: 60 } });
  expect(result.error?.issues[0]).toMatchObject({ path: ["enabled"], message: "Invalid input: expected boolean, received undefined" });
});

test("maxRequests must be above zero", () => {
  const result = rateLimitConfigSchema.safeParse({ maxRequests: 0 });
  expect(result.error?.issues[0]).toMatchObject({ path: ["maxRequests"], message: "Too small: expected number to be >0" });
});

test("defaults are filled in", () => {
  expect(rateLimitConfigSchema.parse({ maxRequests: 60 })).toEqual({ maxRequests: 60, windowMs: 60_000, partitionBy: "global" });
});
```

### The `throttle` defaults have no effect

`defaultRateLimit`, `defaultConcurrency`, `defaultTimeout`, `global`, `globalConcurrency` and `ipFilter` need `enabled: true`. With `enabled: false`, they're all ignored, and so are the tools' own limits except `timeout`. Also check that the tool doesn't set its own limit: that replaces the default.

### Limits aren't shared between server instances

Each process keeps its own counts unless `throttle.storage` points them at a shared store. Check its shape: it's `{ type: "redis", redis: { url } }` (or `redis: { config: { host, port } }`), not `{ provider: "redis", host, port }`. A `storage` without `type` auto-detects from environment variables (`REDIS_URL`, `REDIS_HOST`, `UPSTASH_REDIS_REST_URL`, `KV_REST_API_URL`) and quietly uses memory if none is set, so the second form looks like it works. The top-level `redis` option isn't used by the guard. See [`storage`](#storage-and-keyprefix).

### The server doesn't start: `throttle.storage (redis) is unavailable`

`throttle.storage` points at a store the server can't reach. Rate limits are a security control, so they fail closed: in production, unless `fallback` says otherwise, the server refuses to start instead of running with limits that nobody shares. The error is a `GuardStorageUnavailableError`, with code `GUARD_STORAGE_UNAVAILABLE`, and its message ends with the way out: `Set throttle.storage.fallback: 'memory' to start with per-instance counters instead.` Fix the address or the credentials, or choose the fallback on purpose, knowing that each instance then counts alone. Before 1.8.6, the error was the storage client's own `Failed to connect to Redis`, with nothing about `throttle.storage`.

The tests use a Redis address that nothing listens on:

```ts help-desk.app.ts
import { App, Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({ name: "get_ticket", description: "Get one support ticket by its id", inputSchema: { id: z.string() } })
export class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [GetTicket] })
export class HelpDesk {}
```

```ts storage.test.ts active
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance, GuardStorageUnavailableError } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";

const info = { name: "help-desk", version: "1.0.0" };
const unreachableRedis = { type: "redis", redis: { config: { host: "127.0.0.1", port: 1 } } } as const;

test('`fallback: "error"` stops the server, and the error names `throttle.storage`', async () => {
  const starting = FrontMcpInstance.createDirect({
    info,
    apps: [HelpDesk],
    throttle: { enabled: true, storage: { ...unreachableRedis, fallback: "error" } },
  });
  await expect(starting).rejects.toBeInstanceOf(GuardStorageUnavailableError);
  await expect(starting).rejects.toMatchObject({ code: "GUARD_STORAGE_UNAVAILABLE", statusCode: 503, storageType: "redis" });
  await expect(starting).rejects.toThrow(/^throttle\.storage \(redis\) is unavailable: /);
  await expect(starting).rejects.toThrow("Rate limits fail closed");
});

test('`fallback: "memory"` starts the server with counts of its own', async () => {
  const server = await FrontMcpInstance.createDirect({
    info,
    apps: [HelpDesk],
    throttle: { enabled: true, storage: { ...unreachableRedis, fallback: "memory" } },
  });
  const result = await server.callTool("get_ticket", { id: "T-1" });
  await server.dispose();
  expect(result.isError).toBeFalsy();
});
```

### Calls fail with `GUARD_STORAGE_UNAVAILABLE` while the server runs

The store answered at startup and then stopped, like a Redis that went down. A call that needs a limit check or a concurrency slot can't be counted, so with `fallback: "error"`, the default in production, it is refused: a `GuardStorageUnavailableError`, code `GUARD_STORAGE_UNAVAILABLE`, status `503`. In development, its message says what failed: `throttle.storage (redis) is unavailable: <the store's error>. Rate limits fail closed, so this call was refused. Set throttle.storage.fallback: 'memory' to keep serving with per-instance counters while it is down.` In production a client reads `Service temporarily unavailable: the rate-limit store cannot be reached. Retry shortly.`, which doesn't give away where the store is. With `throttle.global`, the whole request is refused instead, with HTTP `503` and `Retry-After: 1`. With `fallback: "memory"`, the default outside production, the guard switches to counters of its own for each instance, logs it once, tries the store again every 30 seconds, and logs when it answers. Calls made while the store is down are counted per instance, so the limit can be reached on every instance in turn.

The tests pass their own Redis client, one that answers until a test takes it down:

```ts help-desk.app.ts
import { App, Tool, ToolContext, z } from "@frontmcp/sdk";

@Tool({
  name: "get_ticket",
  description: "Get one support ticket by its id. Two calls a minute.",
  inputSchema: { id: z.string() },
  rateLimit: { maxRequests: 2, windowMs: 60_000 },
})
export class GetTicket extends ToolContext {
  async execute({ id }: { id: string }) {
    return { id, title: "Cannot log in", status: "open" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [GetTicket] })
export class HelpDesk {}
```

```ts redis.example.ts
// Stands in for a Redis client, which the Playground can't reach. It answers the few
// commands the rate limiter sends, until `down` is set. A real server doesn't need this file.
export class FakeRedis {
  down = false;
  private values = new Map<string, string>();

  private check() {
    if (this.down) throw new Error("connect ECONNREFUSED 127.0.0.1:6379");
  }
  async ping() {
    this.check();
    return "PONG";
  }
  async mget(...keys: (string | string[])[]) {
    this.check();
    return keys.flat().map((key) => this.values.get(key) ?? null);
  }
  async incr(key: string) {
    this.check();
    const next = Number(this.values.get(key) ?? 0) + 1;
    this.values.set(key, String(next));
    return next;
  }
  async expire() {
    this.check();
    return 1;
  }
  async quit() {
    return "OK";
  }
  on() {
    return this;
  }
  removeListener() {
    return this;
  }
}
```

```ts outage.test.ts active
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance, GuardStorageUnavailableError, createErrorHandler, toMcpError } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";
import { FakeRedis } from "./redis.example";

const start = (redis: FakeRedis, fallback: "error" | "memory") =>
  FrontMcpInstance.createDirect({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDesk],
    throttle: { enabled: true, storage: { type: "redis", redis: { client: redis as never }, fallback } },
  });

test('`fallback: "error"` refuses calls while the store is down, and serves again when it answers', async () => {
  const redis = new FakeRedis();
  const server = await start(redis, "error");
  try {
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });

    redis.down = true;
    const refused = await server.callTool("get_ticket", { id: "T-1" }).catch((error) => error);
    expect(refused).toBeInstanceOf(GuardStorageUnavailableError);
    expect(refused).toMatchObject({ code: "GUARD_STORAGE_UNAVAILABLE", statusCode: 503, storageType: "redis" });
    expect(refused.message).toContain("Rate limits fail closed, so this call was refused.");
    expect(toMcpError(refused).isPublic).toBe(true);
    const production = createErrorHandler({ isDevelopment: false }).handle(refused);
    expect(production.content).toEqual([
      { type: "text", text: "Service temporarily unavailable: the rate-limit store cannot be reached. Retry shortly." },
    ]);

    redis.down = false;
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });
  } finally {
    await server.dispose();
  }
});

test("with `global`, the whole request gets 503 and Retry-After", async () => {
  const redis = new FakeRedis();
  const handler = await FrontMcpInstance.createFetchHandler({
    info: { name: "help-desk", version: "1.0.0" },
    apps: [HelpDesk],
    throttle: { enabled: true, global: { maxRequests: 100, windowMs: 60_000 }, storage: { type: "redis", redis: { client: redis as never }, fallback: "error" } },
  });
  redis.down = true;
  const response = await handler(
    new Request("https://desk.example.com/", {
      method: "POST",
      headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/list" },
      body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/list", params: { _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } } }),
    }),
  );
  expect(response.status).toBe(503);
  expect(response.headers.get("retry-after")).toBe("1");
  expect(await response.json()).toMatchObject({
    error: { code: -32603, message: "Service temporarily unavailable: the rate-limit store cannot be reached", data: { code: "GUARD_STORAGE_UNAVAILABLE" } },
  });
});

test('`fallback: "memory"` keeps serving, with counts of its own', async () => {
  const redis = new FakeRedis();
  const server = await start(redis, "memory");
  try {
    redis.down = true;
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });
    expect((await server.callTool("get_ticket", { id: "T-1" })).structuredContent).toMatchObject({ id: "T-1" });
    await expect(server.callTool("get_ticket", { id: "T-1" })).rejects.toMatchObject({ code: "RATE_LIMIT_EXCEEDED" });
  } finally {
    await server.dispose();
  }
});
```

New in 1.8.7: this runtime case. Starting without the store was already handled in 1.8.6. Changed in 1.9: before, production clients read the development message, with the store's error and address, and with `throttle.global` the request failed with a bare `500`.

### Every request gets `-32001 Forbidden: client IP rejected by ipFilter`

`ipFilter` has `defaultAction: "deny"`, and the requests' address isn't on the allow list: the server is behind a proxy it doesn't trust, so every request has the proxy's address, or it runs under `createFetchHandler()` in a runtime that doesn't report addresses, so it has none. Set `FRONTMCP_TRUST_PROXY=true` (and `FRONTMCP_TRUSTED_PROXY_DEPTH` for more than one proxy) so it reads `X-Forwarded-For`. `ipFilter.trustProxy` in the config does nothing.

### An `allowList` doesn't keep other addresses out

An address on neither list gets `defaultAction`, which is `"allow"`. Add `defaultAction: "deny"`. See [the test above](#filtering-by-ip-address).

### `Tool "…" execution failed: Concurrency limit reached for …`, with `TOOL_EXECUTION_ERROR`

`execute()` threw a guard error of its own, like a `ConcurrencyLimitError` for a slot from a `GuardManager`. Thrown, only `ExecutionTimeoutError` keeps its code; the others are wrapped as `TOOL_EXECUTION_ERROR` and hidden in production. Pass the error to `this.fail()`, which keeps its code and message, or fail with `RateLimitError`, as in [Limiting by an argument](#limiting-by-an-argument):

```ts start-export.tool.ts active
import { ConcurrencyLimitError, createGuardManager, Tool, ToolContext } from "@frontmcp/sdk";

const limits = createGuardManager({ config: { enabled: true } });
const oneAtATime = { maxConcurrent: 1 };

@Tool({ name: "start_export", description: "Export tickets. One export at a time.", inputSchema: {} })
export class StartExport extends ToolContext {
  async execute() {
    const slot = await (await limits).acquireSemaphore("export", oneAtATime);
    if (!slot) throw new ConcurrencyLimitError("export", 1); // 🚩 thrown: TOOL_EXECUTION_ERROR
    await new Promise((resolve) => setTimeout(resolve, 200));
    await slot.release();
    return { exported: true };
  }
}

@Tool({ name: "start_import", description: "Import tickets. One import at a time.", inputSchema: {} })
export class StartImport extends ToolContext {
  async execute() {
    const slot = await (await limits).acquireSemaphore("import", oneAtATime);
    if (!slot) this.fail(new ConcurrencyLimitError("import", 1)); // ✅ keeps CONCURRENCY_LIMIT
    await new Promise((resolve) => setTimeout(resolve, 200));
    await slot.release();
    return { imported: true };
  }
}
```

```ts slots.test.ts
import { test, expect } from "@frontmcp/testing";

test("a thrown guard error is wrapped", async ({ mcp }) => {
  const [, second] = await Promise.all([mcp.tools.call("start_export", {}), mcp.tools.call("start_export", {})]);
  expect(second).toBeError("TOOL_EXECUTION_ERROR");
  expect(second.text()).toMatch(/^Tool "start_export" execution failed: Concurrency limit reached for "export" \(max: 1\)/);
});

test("passed to this.fail(), it keeps its code", async ({ mcp }) => {
  const [, second] = await Promise.all([mcp.tools.call("start_import", {}), mcp.tools.call("start_import", {})]);
  expect(second).toBeError("CONCURRENCY_LIMIT");
  expect(second.text()).toBe('Concurrency limit reached for "import" (max: 1)');
});
```

### One anonymous client used up everyone's limit

Under `"userId"` and `"session"`, all anonymous callers share the key `"anonymous"`. Have callers sign in, so each gets a count of their own, or, on a real server, count anonymous callers by address: `partitionBy: (ctx) => ctx.userId ?? ctx.clientIp ?? "anonymous"`.
