Caching Results

IntermediateMCP 2026-07-28

A model asks the same question more than once. It looks up the customer on a ticket, then again before it writes the reply, then again for the next ticket from the same customer. When the lookup waits on a slow CRM, every repeat makes the user wait again for an answer the model already had. @frontmcp/plugin-cache stores what a tool returns and answers the next call with the same arguments from the store, without running the tool. Turning it on is one line in the app and one field on the tool. Deciding what to cache, and for whom, is the rest of this lesson.

You will learn

  • How to cache a tool's results with the cache plugin
  • Why the cache keeps each caller's answers apart, and when to share them
  • What makes two calls the same call
  • How long an entry lives, and why you can't clear one
  • What not to cache

A lookup the model repeats

get_customer asks the CRM for a customer's account. The CRM takes a fifth of a second to answer, and crm.ts counts the requests it gets. Open the Tests tab:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";
import { crm } from "./crm";

@Tool({
  name: "get_customer",
  description: "Get a customer account by its id: the company's name and its plan.",
  inputSchema: { id: z.string().describe("Customer id, like C-1") },
  annotations: { readOnlyHint: true },
})
export class GetCustomer extends ToolContext {
  async execute({ id }: { id: string }) {
    return crm.getCustomer(id);
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Three calls, three requests to the CRM, three waits, for the same answer. A customer's name and plan don't change from one minute to the next, so the second and third calls could have been answered at once.

Turning the cache on

Install the plugin:

npm install @frontmcp/plugin-cache

Then register it on the app with CachePlugin.init({ type: "memory" }), and mark each tool it should cache with cache. ttl is how long an answer is kept, in seconds:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { CachePlugin } from "@frontmcp/plugin-cache";
import { GetCustomer } from "./get-customer.tool";

@App({
  id: "help-desk",
  name: "Help Desk",
  tools: [GetCustomer],
  plugins: [CachePlugin.init({ type: "memory" })],
})
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The cache is on, and the CRM still gets every request. Nothing is wrong with the setup. By default the cache keeps a separate entry for each caller, so that one user is never given an answer that was meant for another. The Playground's client doesn't sign in, and under MCP 2026-07-28 a caller without credentials is a new anonymous caller on every request. Each call stores an entry under a caller that never comes back, and no call ever reads one. There's no error or warning.

Sharing answers between callers

A customer's account looks the same to every support agent who asks, so there's no reason to keep one copy per caller. keyByIdentity: false makes every caller share one entry per tool and arguments:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { CachePlugin } from "@frontmcp/plugin-cache";
import { GetCustomer } from "./get-customer.tool";

@App({
  id: "help-desk",
  name: "Help Desk",
  tools: [GetCustomer],
  plugins: [CachePlugin.init({ type: "memory", keyByIdentity: false })], // ✅ the same answer for everyone
})
class HelpDeskApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDeskApp],
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Now the first call asks the CRM and the second is answered from the store, without running execute(). Every caller shares the entry, so one agent's lookup also saves the next agent's wait.

  • A hit is marked, and its data isn't changed. Its result carries _meta.cache: "hit", next to the data the tool returned. The marker isn't in structuredContent or in the text the model reads (What a cache hit looks like).
  • Only calls that succeed are stored. A tool that throws stores nothing, so C-9 is looked up again each time.
  • The store is in memory. type: "memory" keeps entries in the server's process: they're lost on restart and not shared between instances. With more than one instance, keep them in Redis instead (Using Redis).

Answers that depend on who's asking

Sharing is right for a customer's account. It's wrong for my_open_tickets, which lists the tickets assigned to the caller. A plugin applies to the tools of the app it's registered on, so tools that need different rules go in different apps, each with its own CachePlugin.init(). my_open_tickets keeps the default, one entry per caller.

The Playground's client is anonymous, which would tell us nothing here, so the tests call the server in-process with createDirect() as two signed-in support agents, nour and sam:

Open
import { App, FrontMcp } from "@frontmcp/sdk";
import { CachePlugin } from "@frontmcp/plugin-cache";
import { GetCustomer, MyOpenTickets } from "./tools";

@App({
  id: "customers",
  name: "Customers",
  tools: [GetCustomer],
  plugins: [CachePlugin.init({ type: "memory", keyByIdentity: false })], // the same answer for everyone
})
class CustomersApp {}

@App({
  id: "queue",
  name: "Queue",
  tools: [MyOpenTickets],
  plugins: [CachePlugin.init({ type: "memory" })], // ✅ one entry per caller
})
class QueueApp {}

export const config = {
  info: { name: "help-desk", version: "1.0.0" },
  apps: [CustomersApp, QueueApp],
};

@FrontMcp(config)
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

Each agent's first call searches, and their repeat is answered from their own entry. With the whole app on keyByIdentity: false, the second test shows what goes wrong: Sam asks for his tickets and gets Nour's, marked as a cache hit.

With keyByIdentity on, the cache tells callers apart by what authentication established:

CallerEntries
Signed in through an identity provider (transparent, local or remote)One set per user, by the token's sub, shared by all of that user's sessions and devices
A static keyOne set per key
Anonymous, under MCP 2026-07-28None that are ever read back: a new caller on every request
Anonymous, on a session-based client (before 2026-07-28)One set per session

What makes two calls the same

The cache looks an answer up by a key made of three things: the caller (unless keyByIdentity is false), the tool as <app id>:<tool name>, and the call's arguments after the input schema has validated them. That last part decides how often a model's calls hit. search_articles searches the help center, whose articles are the same for everyone:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";
import { helpCenter } from "./help-center";

@Tool({
  name: "search_articles",
  description: "Search the help center's articles. Returns the best matches, most relevant first.",
  inputSchema: {
    query: z.string().trim().toLowerCase().describe("What to search for, like 'password reset'"),
    limit: z.number().int().min(1).max(10).default(5).describe("How many articles to return"),
  },
  annotations: { readOnlyHint: true },
  cache: { ttl: 600 },
})
export class SearchArticles extends ToolContext {
  async execute({ query, limit }: { query: string; limit: number }) {
    return { articles: await helpCenter.search(query, limit) };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The key is built from what execute() receives, not from what the model sent. Object keys are sorted, arguments the schema doesn't declare are dropped, defaults are filled in, and transforms like trim() and toLowerCase() have already run. So a schema that normalizes its input also makes the cache hit more often: " Rules " and "RULES" are the same search as "rules". Anything left that differs, even limit: 1 against limit: 2, is another entry.

How long an entry lives

An entry lives for its ttl, and then the next call runs the tool again. The plugin has no way to delete an entry sooner. Here update_plan changes a customer's plan in the CRM, and get_customer keeps giving the old one. The tests move the clock forward by replacing Date.now(), which is what the cache's memory store reads:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";
import { crm } from "./crm";

@Tool({
  name: "get_customer",
  description: "Get a customer account by its id: the company's name and its plan. May be up to 5 minutes old.",
  inputSchema: { id: z.string().describe("Customer id, like C-1") },
  annotations: { readOnlyHint: true },
  cache: { ttl: 300 },
})
export class GetCustomer extends ToolContext {
  async execute({ id }: { id: string }) {
    return crm.getCustomer(id);
  }
}

@Tool({
  name: "update_plan",
  description: "Change a customer's plan.",
  inputSchema: { id: z.string(), plan: z.enum(["starter", "business", "enterprise"]) },
})
export class UpdatePlan extends ToolContext {
  async execute({ id, plan }: { id: string; plan: "starter" | "business" | "enterprise" }) {
    return crm.setPlan(id, plan);
  }
}

@Tool({
  name: "get_sla_policy",
  description: "The help desk's response times for each plan.",
  inputSchema: {},
  annotations: { readOnlyHint: true },
  cache: true, // 🚩 kept for a day: the plugin's defaultTTL
})
export class GetSlaPolicy extends ToolContext {
  async execute() {
    crm.requests++;
    return { starter: "2 days", business: "8 hours", enterprise: "1 hour" };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

The model changed the plan, and the next lookup still says starter: the model is told two different things in a row. A client can skip the cache for one call by sending the header x-frontmcp-disable-cache: true, but that call's fresh answer isn't stored, and the old entry stays for everyone else (Letting a client skip the cache). Short of restarting a server whose store is in memory, which empties it, nothing clears an entry early.

So choose ttl as the oldest answer you're willing to give:

  • Say it in the description. "May be up to 5 minutes old" lets the model tell the user, or tell them a change can take a few minutes to show.
  • Keep it short, or don't cache, for data your own tools change. If update_plan is used often, a five-minute-old plan will confuse the model more than a slow lookup.
  • Give every cached tool an explicit ttl. cache: true uses the plugin's defaultTTL, a day unless you set it, so a change can take a day to reach the model. The last test shows get_sla_policy answering from its first call until the day is over, however often it's read. (Before FrontMCP 1.9, each read gave the entry a full day again, so an answer read daily was never refreshed. Reading an entry now keeps it alive only with slideWindow: true: see the reference.)

What not to cache

A cached call doesn't run execute(). For a tool that only reads, that's the point. For a tool that does something, it means the second call doesn't do it:

Open
import { Tool, ToolContext, z } from "@frontmcp/sdk";

export const notes: string[] = [];

@Tool({
  name: "add_note",
  description: "Add an internal note to a support ticket.",
  inputSchema: { ticketId: z.string(), text: z.string() },
  cache: { ttl: 300 }, // 🚩 on a tool that changes something
})
export class AddNote extends ToolContext {
  async execute({ ticketId, text }: { ticketId: string; text: string }) {
    notes.push(`${ticketId}: ${text}`);
    return { ticketId, added: true };
  }
}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

An agent really may add the same note twice, after the customer writes in again. The second call answers "added" from the cache, and nothing is added. Leave these out of the cache:

Don't cacheBecause
Tools that change something: add_note, close_ticket, refund_invoiceA hit skips the change and reports success. Cache only tools that just read, the ones you'd mark readOnlyHint: true.
Answers that depend on the caller, in an app with keyByIdentity: falseThe first caller's answer goes to everyone. Keep them per caller, behind authentication.
Answers the model acts on right awayA ticket's status just before closing it, a balance just before refunding it: a few minutes old can be wrong.

Recap

  • @frontmcp/plugin-cache answers a repeated call from a store without running execute(). Register CachePlugin.init({ type: "memory" }) on the app, and mark each tool with cache: { ttl }, in seconds.
  • By default, each caller has their own entries. Anonymous callers under MCP 2026-07-28 are new on every request, so they're never answered from the cache.
  • keyByIdentity: false shares entries between callers. Use it only for answers that are the same for everyone, in an app of their own; a plugin covers the tools of the app it's registered on.
  • Two calls are the same when the tool and the validated arguments match: key order, undeclared arguments and defaults don't matter, and the schema's transforms have run.
  • An entry lives until its ttl ends, and the plugin can't clear it sooner. Choose ttl as the oldest answer you'll accept, and say so in the description.
  • Cache only tools that read. Failed calls aren't stored, whether the tool throws or returns an error result.
  • Every option, the Redis stores and the bypass header are in the Cache plugin reference.

Try some challenges

Each challenge runs hidden checks against your code. Edit the code, then press Check.

Challenge 1 of 3

Cache the article lookup

get_article reads an article from the help center, which is slow, and the model reads the same few articles over and over. Answer repeated lookups from a cache for 10 minutes. Articles are the same for everyone who asks.

Open
import { App, FrontMcp, Tool, ToolContext, z } from "@frontmcp/sdk";
import { helpCenter } from "./help-center";

@Tool({
  name: "get_article",
  description: "Read a help center article by its slug.",
  inputSchema: { slug: z.string().describe("The article's slug, like reset-your-password") },
  annotations: { readOnlyHint: true },
})
class GetArticle extends ToolContext {
  async execute({ slug }: { slug: string }) {
    return helpCenter.getArticle(slug);
  }
}

@App({ id: "help-center", name: "Help Center", tools: [GetArticle] })
class HelpCenterApp {}

@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpCenterApp],
})
export default class Server {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.