Limiting and Timing Out Calls

IntermediateMCP 2026-07-28

A model can call a tool as often as it likes, and nothing in MCP stops it from calling one in a loop. For most tools that's fine. For an export that reads the whole database, a merge that must never run twice at once, or a lookup that can hang on a slow service, it isn't. Three @Tool options set limits: rateLimit caps how often a tool runs, concurrency caps how many calls run at once, and timeout caps how long one call may take.

You will learn

  • How to cap how often a tool runs, and what a caller gets back when a limit trips
  • Whose calls a limit counts, and what anonymous callers share
  • How to stop two runs of a tool overlapping
  • What a timeout stops, and how to let the work stop with it
  • How to choose limits for expensive and destructive tools

Capping how often a tool runs

Exporting tickets reads every ticket in the database. The help desk wants at most two exports an hour, so the tool sets rateLimit. Open the Tests tab:

Open
import { Tool, ToolContext } from "@frontmcp/sdk";
import { tickets } from "./store";

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV. Reads the whole ticket database, so at most 2 exports an hour.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
  rateLimit: { maxRequests: 2, windowMs: 60 * 60_000 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    const rows = tickets.map((t) => `${t.id},${t.title},${t.status}`);
    return { csv: ["id,title,status", ...rows].join("\n") };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The export above ran once when the example started. Press Call in the Call tab twice more: the second call works, the third is refused. rateLimit takes:

FieldDefaultMeaning
maxRequestsrequiredHow many calls a window allows.
windowMs60_000The window's length, in milliseconds.
partitionBy"global"Whose calls share a count. See below.

What the caller gets back

A call over the limit doesn't reach execute(). The client gets an ordinary tool error:

{
  "content": [{ "type": "text", "text": "Rate limit exceeded. Retry after 2274 seconds" }],
  "isError": true,
  "_meta": { "errorId": "err_…", "code": "RATE_LIMIT_EXCEEDED" }
}
  • The retry hint is in the text, for the model. The HTTP status is still 200 and there's no Retry-After header, so the model is the one that reads when to try again. The message survives in production too.
  • The wait is until the window ends. Windows line up with the clock, not with the first call: an hour-long window runs from one full hour (in UTC) to the next, and a minute-long one from one full minute to the next. The number is the time left in the current window, so it can be anything from a second up to the whole windowMs. Refused at 10:40 UTC by an hourly limit, a caller is told to retry in about 1200 seconds.

Say the limit in the description as well, as export_tickets does ("at most 2 exports an hour"). A model that knows the rule can plan around it, or tell the user, instead of finding out from an error.

Whose calls a limit counts

partitionBy decides which calls share a count:

partitionByOne count perUnder MCP 2026-07-28
"global"The whole serverEvery caller shares one limit.
"userId"Signed-in caller, by authInfo.clientIdEach caller with a key or token has a count of their own. All anonymous callers share one.
"ip"Client IP addressWorks on a real server. The Playground has no IP, so there it behaves like "userId".
"session"MCP sessionThere are no sessions, so it behaves like "userId".
(ctx) => stringWhatever key you returnctx has sessionId, plus userId for a signed-in caller and clientIp when there is one.

"userId" is the natural choice for "two exports per person", and for signed-in callers it does exactly that. Anonymous callers have no user id, so they all share one count:

Open
import { Tool, ToolContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description:
    "Export every support ticket as CSV. At most 2 exports an hour per signed-in caller, and 2 an hour for all anonymous callers together.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
  rateLimit: { maxRequests: 2, windowMs: 60 * 60_000, partitionBy: "userId" },
})
export class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title,status\nT-1,Cannot log in,open" };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Each of these calls arrives as a new anonymous caller with a new id, but FrontMCP doesn't count anonymous callers by those ids, so the third one is refused like the third call from a single person. That's the safe choice: nobody gets around a limit by not signing in. It also means one busy anonymous client can use up the exports for every other anonymous client. If that matters, have callers sign in, so each gets a count of their own, or, on a real server, count anonymous callers by address with a function: partitionBy: (ctx) => ctx.userId ?? ctx.clientIp ?? "anonymous".

Counts are kept in the server's memory unless you say otherwise, so each instance of a server counts on its own: two instances behind a load balancer let through four exports an hour, not two. With throttle.storage on one Redis, the instances share a count: five exports sent to two instances in turn let two through. Running Several Instances sets that up. If that Redis can't be reached, a production server doesn't start, unless throttle.storage has fallback: "memory", which starts it with each instance counting on its own. When Redis goes away while the server runs, a call to a tool with a limit fails with GUARD_STORAGE_UNAVAILABLE, or, with the fallback, is counted in the instance's memory until Redis answers again: the guard reference shows what a caller gets, and what happens at startup.

One at a time

Merging two customer records moves all their tickets from one record to the other. Two merges running at once can move the same tickets twice. concurrency caps how many calls of a tool run at the same moment:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";

export const merges: string[] = [];

@Tool({
  name: "merge_customers",
  description:
    "Merge one customer record into another, moving all its tickets. One merge runs at a time; if another is running, wait a minute and try again.",
  inputSchema: { from: z.string().describe("Customer to merge away, like C-7"), into: z.string().describe("Customer to keep") },
  annotations: { destructiveHint: true },
  concurrency: { maxConcurrent: 1 },
})
export class MergeCustomers extends ToolContext {
  async execute({ from, into }: { from: string; into: string }) {
    await new Promise((resolve) => setTimeout(resolve, 300)); // moving tickets takes a while
    merges.push(`${from} → ${into}`);
    return { merged: from, into };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The slot is taken when a call starts and given back when its execute() ends, whether it succeeded or failed. concurrency takes maxConcurrent, partitionBy (the same choices as for rateLimit, "global" by default), and queueTimeoutMs. By default, a call that finds no free slot is refused at once. With a queue, it waits for a slot, and is only refused if none frees up in time:

concurrency: { maxConcurrent: 1, queueTimeoutMs: 5_000 },

A queued call that runs out of time gets Queue timeout for "merge_customers" after waiting 5000ms for a concurrency slot. For a tool a model is likely to call twice in a row, a short queue is kinder than a refusal.

Every limit answers the same way: an isError tool result, over HTTP 200, with a code of its own in _meta.code and a message that is kept in production too:

Limit_meta.codeMessage
rateLimitRATE_LIMIT_EXCEEDEDRate limit exceeded. Retry after 1200 seconds
concurrencyCONCURRENCY_LIMITConcurrency limit reached for "merge_customers" (max: 1)
concurrency with a queueQUEUE_TIMEOUTQueue timeout for "merge_customers" after waiting 5000ms for a concurrency slot
timeoutEXECUTION_TIMEOUTExecution of "lookup_customer" timed out after 500ms

So the model can tell "busy, try again soon" from "broken". It still helps to put the rule in the tool's description, as merge_customers does, so the model knows what to do about it.

Deadlines

A tool that calls another service waits as long as that service takes. timeout: { executeMs } puts a limit on it. Customer C-9 sends the CRM into a slow path:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";

// Stands in for a CRM that is sometimes very slow to answer.
async function crmLookup(id: string) {
  await new Promise((resolve) => setTimeout(resolve, id === "C-9" ? 2_000 : 50));
  return { id, name: id === "C-9" ? "Globex" : "Initech", plan: "business" };
}

@Tool({
  name: "lookup_customer",
  description: "Look up a customer in the CRM by id. Gives up after half a second if the CRM is slow.",
  inputSchema: { id: z.string().describe("Customer id, like C-1") },
  annotations: { readOnlyHint: true, openWorldHint: true },
  timeout: { executeMs: 500 },
})
export class LookupCustomer extends ToolContext {
  async execute({ id }: { id: string }) {
    return crmLookup(id);
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

After half a second the call fails with Execution of "lookup_customer" timed out after 500ms. Try C-1: it answers in time.

A timeout ends the call, not the work

When a call times out, FrontMCP stops waiting for execute() and answers the client. execute() itself carries on unless it's written to stop (the next section shows how), side effects and all. Its concurrency slot stays taken until it really ends, so runs never overlap, and a call that arrives in the meantime is refused:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";

const merges: string[] = [];
let running = 0;
let mostAtOnce = 0;

@Tool({
  name: "merge_customers",
  description: "Merge one customer record into another, moving all its tickets. One merge runs at a time.",
  inputSchema: { from: z.string(), into: z.string() },
  annotations: { destructiveHint: true },
  concurrency: { maxConcurrent: 1 },
  timeout: { executeMs: 200 },
})
export class MergeCustomers extends ToolContext {
  async execute({ from, into }: { from: string; into: string }) {
    running++;
    mostAtOnce = Math.max(mostAtOnce, running);
    await new Promise((resolve) => setTimeout(resolve, 500)); // takes longer than the timeout
    merges.push(`${from} → ${into}`);
    running--;
    return { merged: from, into };
  }
}

@Tool({ name: "merge_log", description: "List finished merges.", inputSchema: {} })
export class MergeLog extends ToolContext {
  async execute() {
    return { merges, mostAtOnce };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

For a tool that changes things, a timeout shorter than the work is worse than none. The model is told the merge failed while it's still going, then that merges are busy, and it may well run the same merge again once the first one is done. Give such tools no timeout, or one comfortably longer than the work can take, and make them safe to repeat, for example by checking whether the merge already happened before starting it.

Stopping the work

A timeout also aborts this.signal, the call's AbortSignal. Work that listens to it stops with the call. fetch() and most clients for databases and other services take a signal option, and a loop can call this.signal?.throwIfAborted() between steps. (The type says signal may be missing, hence the ?..) This export reads the ticket database a page at a time, and a slow database makes it take longer than its timeout:

Open
import { Tool, ToolContext } from "@frontmcp/sdk";

export const pagesRead: number[] = [];
export const stops: string[] = [];

// Stands in for a slow database: 10 pages, a tenth of a second each.
async function readPage(page: number) {
  await new Promise((resolve) => setTimeout(resolve, 100));
  pagesRead.push(page);
  return [`T-${page},Ticket ${page},open`];
}

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV. Gives up after 350 ms.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
  timeout: { executeMs: 350 },
})
export class ExportTickets extends ToolContext {
  async execute() {
    const rows: string[] = [];
    try {
      for (let page = 1; page <= 10; page++) {
        this.signal?.throwIfAborted(); // ✅ stop between pages once the call has timed out
        rows.push(...(await readPage(page)));
      }
    } catch (error) {
      stops.push(String(this.signal?.reason?.message ?? error));
      throw error;
    }
    return { csv: ["id,title,status", ...rows].join("\n") };
  }
}

@Tool({ name: "export_stats", description: "How many database pages exports have read, and why they stopped.", inputSchema: {} })
export class ExportStats extends ToolContext {
  async execute() {
    return { pagesRead: pagesRead.length, stops };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Take out the throwIfAborted() line and run the test again: the call still times out after 350 ms, but the export goes on to read all ten pages that nobody will see. this.signal.reason is the same timeout error the client got. Stopping is right for work that only reads. For work that changes things, stopping halfway can be worse than finishing, which is why the merge above doesn't listen to the signal.

Turning limits off

The limits in @Tool are enforced by FrontMCP's guard, which starts by itself when a tool declares rateLimit or concurrency. @FrontMcp's throttle option configures the guard, and throttle: { enabled: false } turns it off: every rateLimit and concurrency on the server is ignored. That can be what you want for a local test run that calls the same tool many times:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { ExportTickets } from "./export-tickets.tool";

@App({ id: "help-desk", name: "Help Desk", tools: [ExportTickets] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: { enabled: false }, // 🚩 for local testing only
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The description still says "at most 2 exports an hour", and nothing enforces it, so keep enabled: false out of anything a real client can reach. timeout is the one limit it doesn't turn off.

Limits for every tool

throttle can also set defaults, for tools that don't set their own. Unlike a tool's own limits, throttle needs enabled: true:

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
  throttle: {
    enabled: true,
    defaultRateLimit: { maxRequests: 60, windowMs: 60_000, partitionBy: "ip" },
    defaultConcurrency: { maxConcurrent: 5 },
    defaultTimeout: { executeMs: 30_000 },
  },
})
export default class Server {}

A tool's own rateLimit, concurrency or timeout replaces the default. The @FrontMcp reference shows defaultRateLimit running.

Choosing limits

Kind of toolExampleLimits
Cheap and read-onlysearch_ticketsUsually none of its own. A server-wide defaultRateLimit if strangers can reach the server.
Expensive to runexport_ticketsA rateLimit per caller ("userId"), and a timeout whose this.signal stops the work.
Waits on another servicelookup_customerA timeout shorter than the client is willing to wait, so the model gets an error it can explain, with this.signal passed on to the request.
Changes shared statemerge_customersconcurrency: { maxConcurrent: 1 }, with a short queue. No timeout shorter than the work.
Sends email or spends moneyemail_customer, refund_orderA low rateLimit per caller. A model stuck in a loop shouldn't be able to send a hundred emails. A limit doesn't ask anyone: to have the user agree first, add approval.

Whatever you choose, write it in the description. Limits the model knows about are limits it can work with.

Recap

  • rateLimit, concurrency and timeout on @Tool work as soon as a tool sets them. throttle in @FrontMcp adds defaults for every tool, and throttle: { enabled: false } turns rate limits and concurrency off.
  • A call over a limit gets an isError result, over HTTP 200, with its own _meta.code (RATE_LIMIT_EXCEEDED, CONCURRENCY_LIMIT, QUEUE_TIMEOUT, EXECUTION_TIMEOUT) and a message kept in production. "Retry after N seconds" counts down to the end of a clock-aligned window.
  • partitionBy picks whose calls share a count. With "userId", each signed-in caller has a count of their own, and all anonymous callers share one.
  • concurrency: { maxConcurrent, queueTimeoutMs } stops calls overlapping, and a timed-out call keeps its slot until execute() really ends.
  • timeout: { executeMs } ends the call and aborts this.signal. Pass the signal on so read-only work stops too, and don't give tools that change things a timeout shorter than their work.
  • Every guard option, and what each error looks like to a caller, is in the Guard options reference.

Try some challenges

Each challenge runs hidden checks against your code. Edit the code, then press Check.

Challenge 1 of 4

Cap the export

Nothing stops a model from exporting the ticket database over and over. Allow at most 3 exports an hour, for the whole server.

Open
import { App, FrontMcp, Tool, ToolContext } from "@frontmcp/sdk";

@Tool({
  name: "export_tickets",
  description: "Export every support ticket as CSV.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
})
class ExportTickets extends ToolContext {
  async execute() {
    return { csv: "id,title,status\nT-1,Cannot log in,open\nT-2,Invoice total is wrong,closed" };
  }
}

@App({ id: "help-desk", name: "Help Desk", tools: [ExportTickets] })
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.