CRM with CodeCall

Intermediate

A help desk's CRM covers more than tickets: customers and their contacts, accounts, orders, invoices, notes, agents, macros and articles, each with its own list, get, create, update and delete. This example is a CRM server with 47 tools, and it doesn't list them. It puts them behind CodeCall: the model is listed six tools of CodeCall's, uses them to search for the tools it needs and read their schemas, and answers a question that spans customers, tickets and accounts with one script that runs on the server. The nine destructive tools stay out of the script's reach.

You will learn

  • What a model is listed instead of 47 tools, and why that helps
  • How to write tools a model can find and use through CodeCall: the words it will search for, examples and output schemas
  • How a question that needs eight tool calls becomes a search, a describe and one script
  • How to keep destructive tools out of scripts, and what that doesn't protect
  • How to test all of it, and how the Playground's sandbox differs from a server's

The server

Here is the whole server. Open the Call tab: codecall:search has looked for tools for four short phrases, the ones a model would send to answer "Which enterprise customers have an open, high-priority ticket that nobody has replied to in two days, and who manages their account?". It found list_customers, list_tickets, list_accounts and get_agent among the 38 tools CodeCall may use. Call codecall:describe with { "toolNames": ["list_tickets"] } to read what the model reads next, and codecall:invoke with { "tool": "close_ticket", "input": { "id": "T-8" } } to make one call without a script. Then open the Tests tab: with no model in the Playground, they play its part.

Open
import { App } from "@frontmcp/sdk";
import { CodeCallPlugin } from "@frontmcp/plugin-codecall";
import { areaTools } from "./area.tools";
import { CrmStore, TicketStore } from "./stores";
import { crmTools } from "./crm.tools";

@App({
  id: "help-desk",
  name: "Help Desk CRM",
  providers: [CrmStore, TicketStore],
  tools: [...areaTools, ...crmTools],
  plugins: [
    CodeCallPlugin.init({
      mode: "codecall_only",
      includeTools: (tool) => !tool.annotations?.destructiveHint,
    }),
  ],
})
export class HelpDeskApp {}

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

codecall:execute runs scripts in the Playground too, with a different sandbox than a server's. A server runs a script in a sandbox built on Node's vm module, which a browser doesn't have, so the Playground runs it in Enclave's browser sandbox: two nested iframes with no network and no eval. The script is checked and rewritten by the same code first, its tool calls go through the same flow to the tools above, and the 14 tests here pass on Node and in the browser. In the Playground lists what differs.

How it fits together

  1. The model is listed CodeCall's six tools, not 47. Their descriptions tell it how to work: search, then describe, then execute or invoke.
  2. It calls codecall:search with a short phrase for each thing it needs. CodeCall ranks the CRM's tools by the words they share with each phrase, and answers with the best ones.
  3. It calls codecall:describe with the names it picked, and gets each tool's input and output schemas, and usage examples.
  4. It writes a script from what it read, and calls codecall:execute. CodeCall checks the script against AgentScript's rules, then runs it in a sandbox on the server.
  5. The script calls the CRM's tools with callTool(): eight calls here, each through the same flow as a call from a client.
  6. Only what the script returns goes back to the model.

This is the script the tests send, as a model would write it, once it has read the schemas:

const twoDaysAgo = new Date(Date.now() - 2 * 24 * 60 * 60 * 1000).toISOString();
const { customers } = await callTool("list_customers", { plan: "enterprise" });
const report = [];
for (const customer of customers) {
  const { tickets } = await callTool("list_tickets", { customerId: customer.id, status: "open", priority: "high" });
  const waiting = tickets.filter((ticket) => (ticket.lastReplyAt ?? ticket.openedAt) < twoDaysAgo);
  if (waiting.length === 0) continue;
  const { accounts } = await callTool("list_accounts", { customerId: customer.id });
  const manager = await callTool("get_agent", { id: accounts[0].managerId });
  report.push({ customer: customer.name, tickets: waiting.map((ticket) => ticket.id), accountManager: manager.name });
}
return report;

Here is the server again, with codecall:execute already called with that script. Its Call tab has the script in its script field, so the answer is on screen when it starts, and you can change the script and send it again. Try plan: "standard" in the first callTool, which answers Globex with T-2 and its manager Dana, or status: "pending", which answers Umbrella with T-9 and Nour. Show code opens the server's files, the same as above.

Open

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

codecall:execute answers:

{
  "status": "ok",
  "result": [
    { "customer": "Acme", "tickets": ["T-1", "T-10"], "accountManager": "Nour" },
    { "customer": "Hooli", "tickets": ["T-5"], "accountManager": "Dana" }
  ]
}

The model reads three lines instead of eight results. Acme has two stale tickets, and T-10 was never answered, so the script counts from when it was opened. Umbrella's only open high-priority ticket was answered five hours ago, and Globex's stale one belongs to a standard customer.

Why put a 47-tool server behind CodeCall? Listed in full, the tools come to about 43 KB of names, descriptions and schemas, sent with every request. CodeCall's six come to about 21 KB, and stay that size as the CRM grows. That's a saving, but not the biggest one. Nine areas with the same five verbs are look-alike tools, and the question above needs eight calls in a row, each result passing through the model. As a script, it's one call, and the model never sees the tickets it filtered out.

The files

stores.ts: two stores

TicketStore holds the help desk's own data: ten tickets, typed with ticketSchema, which the ticket tools reuse as their outputSchema. Their times count back from when the server starts, so the two-day rule holds whenever you run the example: T-1 was last answered 70 hours ago, T-4 five hours ago, T-10 never (lastReplyAt is null). The tickets are chosen so that every part of the question matters: T-2 is stale and urgent but belongs to a standard customer, T-3 is stale but not urgent, T-7 and T-9 are urgent but closed or waiting on the customer.

CrmStore stands in for the CRM's database: nine tables of rows, with list, get, create, update and remove that work on any of them. Both are providers, GLOBAL, so every tool gets the same instance. A real server would read your databases here.

area.tools.ts: 42 tools from nine areas

The nine areas are a list of settings, and toolsForArea writes their tools: list_, get_, create_, update_ and delete_, with only list_ and get_ for the agents, who are read-only. It's the loop the lesson uses to write its 35 tools; a real CRM server would have a file for each area. What matters is what each tool tells the model:

area.tools.ts
@Tool({
  name: `list_${table}`,
  description: `List ${table}: ${about}. Filter by ${filterBy.join(", ")}.`,
  inputSchema: filters,
  outputSchema: { [table]: z.array(rowSchema) },
  annotations: { readOnlyHint: true },
  examples: [{ description: `List ${table}`, input: listExample }],
})
  • The description says what the area holds in the words a model would search for. Search matches words, not meaning, so about says "each customer's contract, with its account manager and renewal date", and accounts and their managers finds list_accounts. See Finding tools: search, then describe.
  • filterBy keeps each list tool's input to the fields worth filtering by, and every field has a .describe(), which codecall:describe passes on.
  • The examples on every list tool matter more than they look. codecall:describe shows a tool's own examples and only those. A tool without any gets one CodeCall writes from its name and schema, and for a list tool that ends return result.items || result: a guess at how the answer is shaped, and these tools don't return items. A model copies the example it sees, so it should be yours. The step 2 test checks that it is.
  • The outputSchema matters most for scripts. codecall:describe returns it, and it's how the model learns that list_customers answers { customers: [...] } and that a ticket's lastReplyAt can be null. Without one, outputSchema is null and the model writes its script blind.
  • The annotations are what the filter in help-desk.app.ts reads: readOnlyHint on the reads, and destructiveHint on delete_….

crm.tools.ts: the tools that aren't create, read, update and delete

Tickets have a workflow of their own, and their own typed store, so their tools are written by hand: list_tickets, get_ticket, assign_ticket and close_ticket. refund_order is the ninth destructive tool, next to the eight delete_… tools.

crm.tools.ts
if (!this.get(CrmStore).get("agents", agentId)) this.fail(notFound("agent", agentId));

notFound is a PublicMcpError with a code and a message written for the model: There's no agent A-9. codecall:invoke returns it as it is, as a test checks, so the model can correct the id. assign_ticket checks the agent before it changes the ticket, so a wrong id leaves the ticket as it was.

help-desk.app.ts and main.ts: CodeCall on the app

help-desk.app.ts
CodeCallPlugin.init({
  mode: "codecall_only",
  includeTools: (tool) => !tool.annotations?.destructiveHint,
}),

"codecall_only" lists CodeCall's six tools instead of the CRM's. includeTools is called for every tool, with its annotations, and CodeCall only uses the tools it returns true for. The eight delete_… tools and refund_order are marked destructiveHint, so 38 of the 47 remain, which is what totalAvailableTools counts. A destructive tool added later is kept out too, without touching this file. Marking them with codecall: { enabledInCodeCall: false } one by one would miss the next one. See Keeping tools away from CodeCall.

A tool kept out isn't found by search, is in notFound for codecall:describe, is refused by codecall:invoke, and gets Access denied in a script. The model can't tell whether it exists.

main.ts is the usual @FrontMcp server. It makes no difference whether CodeCall is on the app or on the server: either way it hides every app's tools from tools/list and searches all of them (Where to register it). The stores are the app's providers, which is enough for tools; jobs and agents need theirs on the server.

crm.test.ts

The tests play the model. Steps 1 to 3 are what it does with the question: search with four phrases, describe what it found, run one script. The others check what surrounds them.

  • The first two tests are the case for CodeCall: the client is listed six tools, and a second copy of the server without the plugin lists all 47 at over 40 KB. Past 40 tools, tools/list comes in pages (Paging long tool lists). The client connect() returns follows every page, so its listTools() gives all 47. listTools() on the server create() returns still gives only the first 40, with a nextCursor, so the copy is made with connect().
  • Step 3 has two tests: one that the script passes CodeCall's checks, and one that it answers the question. The second runs the script, in Node's vm sandbox on Node and in Enclave's browser sandbox in the Playground, and both tests pass in both.
  • codecall:invoke needs no script for one action, and it returns a tool's own error message, which a script doesn't get (Handling a tool's failure in a script).
  • The destructive tools are out of reach: not found, not described, not invoked, and refused in a script. A client that sends tools/call for refund_order by name gets Tool "refund_order" not found, which is the last assertion of that test: in codecall_only, FrontMCP answers a direct call of a tool CodeCall hides as if it didn't exist, and that's every CRM tool. A script that reads first and then calls one, get_customer and then delete_customer, never reaches the refusal: the sandbox's pattern check stops it with DELETE_AFTER_ACCESS (the limits list the patterns), and the last script test shows it.
  • The last two tests call as a signed-in caller with FrontMcpInstance.createDirect(). CodeCall caches codecall:search and codecall:describe for 60 seconds per signed-in caller: the second identical call comes back with _meta.cache: "hit". codecall:invoke and codecall:execute are never cached. Anonymous callers, like the Playground's, are never cached either, which is why no other test sees old answers.

All the tests share one server, in order. The ticket tests change T-8 and T-9 and the generated-tools test adds a contact, and none of them touches what the script reads. Testing Your Server covers the test API.

Running it for real

Nothing in the code changes. Install the plugin and the Cache plugin it needs, on Node 24 or later:

npm install @frontmcp/plugin-codecall @frontmcp/plugin-cache

I ran the server from this page in a project made with frontmcp create, with FrontMCP 1.9.1 and @frontmcp/plugin-codecall 1.9.1, started it with frontmcp dev, and called it over HTTP. tools/list answered CodeCall's six tools, with no page to follow, and codecall:execute ran the script:

curl -s http://127.0.0.1:3000/mcp \
  -H 'content-type: application/json' \
  -H 'mcp-protocol-version: 2026-07-28' \
  -H 'mcp-method: tools/call' \
  -H 'mcp-name: codecall:execute' \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"codecall:execute","arguments":{"script":"…"},"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28"}}}'

With the script above as the string, its structuredContent was the result shown earlier. With FrontMCP 1.9.4 the tests on this page pass unchanged, in Node; I didn't repeat the run over HTTP. What I couldn't check is a real model: the Playground has none, and neither did I. Whether it searches with phrases that find these tools, describes before it runs, and writes a script AgentScript accepts is for your first conversations to show. A script it gets wrong isn't lost: illegal_access and syntax_error name the rule and the line, and nothing in the script has run. What a script may do lists the rules.

Some things behave differently in production:

  • Signed-in callers see cached answers. Edit a tool's description, or add a tool, and a signed-in caller who searched or described in the last minute gets the old answer until it expires. The last two tests in this example show the cache.
  • No tool but CodeCall's can be called directly. A client that sends tools/call for refund_order, or for get_ticket, gets Tool "…" not found, over HTTP too. A widget, or a client of createDirect(), that should call a tool directly needs codecall: { visibleInListTools: true } on it. To decide who may call a tool through CodeCall, use authorities, which apply to CodeCall's calls as to any other. (Before FrontMCP 1.9, a client that knew a hidden tool's name could call it.)
  • Scripts run with the caller's credentials, within CodeCall's limits: 3.5 seconds of computing time, 5,000 tool calls. The limits lists them, and the patterns the sandbox stops, like a delete_… call soon after a read.
  • Tests. In the scratch project, npx frontmcp test ran all fourteen, and they passed, with the Jest 30 that frontmcp create installs and nothing added to frontmcp.config.ts: frontmcp test transpiles the two ES modules CodeCall's Cache plugin loads, @noble/hashes and @noble/ciphers, by default. The spec needed test.use({ server: "./src/main.ts" }) (fixtures) to start the server. Three tests, the size comparison and the two cache tests, start a second server inside the test process. (With FrontMCP 1.8.7 those three failed there with ERR_VM_DYNAMIC_IMPORT_CALLBACK_MISSING_FLAG, unless NODE_OPTIONS=--experimental-vm-modules was set.)

Ideas to try

Each of these is a change to the Playground above. Add a test for each.

  1. Rename list_tickets to search_tickets and remove its examples. Look at what codecall:describe writes for it now: a query the tool doesn't have, and every ticket back if a model copies it. Put the example back, and test that yours is first.
  2. Make the question ask for more: "...and who has an overdue invoice". Give Hooli an overdue invoice, call list_invoices for each customer in the script, and add the invoices to what it returns, and check the result.
  3. Make the CRM read-only for the model: change includeTools so that only tools with readOnlyHint remain, and check that close_ticket is refused by codecall:invoke and in a script.
  4. Let a client call one tool directly: give get_agent codecall: { visibleInListTools: true }, test that it's listed next to CodeCall's six and that a direct tools/call for it works, and that refund_order is still not found.