Research Assistant
Some support questions have no article of their own: "Why do enterprise customers keep hitting login failures after the SSO change?" The answer is spread over a few knowledge-base articles and the tickets that came before. This example is a help desk server whose research agent reads through both and answers with a few claims, each pointing at the article or ticket ids it came from. It splits the work between two agents: research hands the searching to a narrower agent, find_sources, from its own code, checks what comes back, and asks its own model to write only from the sources that checked out. The articles and tickets are also resources, so a client can open whatever a claim cites.
You will learn
- How a coordinating agent hands work to a narrower one, and why it checks what comes back
- How to make an answer checkable: a schema that demands sources, and code that checks the ids are real
- How articles and tickets become resources a client can open from a citation
- How an agent's tools report progress to the client, even when another agent started that agent
- How two scripted models stand in for real ones, and what changes with real models
The server
Here is the whole server. Open the Call tab: invoke_research has answered the question with three claims, each with the ids of its sources, and the log messages the searches sent are above the result. To see what a claim cites, read the resource kb://KB-2 or tickets://T-43 in the Call tab. Then open the Tests tab.
import { Agent, AgentContext, PublicMcpError, z } from "@frontmcp/sdk";
import { KnowledgeBase } from "./knowledge";
import { writeModel } from "./model.example";
export const RESEARCH_INSTRUCTIONS = `You answer a support agent's question using only the sources you are given: articles (KB-...) and closed tickets (T-...).
Answer only with JSON: {"claims": [{"statement": "<one sentence>", "sources": [ids of the sources that say it]}]}.
Put the most important claim first, write at most four, and give every claim at least one source id.`;
type Answer = { claims: { statement: string; sources: string[] }[] };
@Agent({
name: "research",
description:
"Answer a hard support question from the help desk's own articles and closed tickets. Returns a few claims, each with the ids of " +
"the sources it came from: read KB-... ids as the resource kb://<id>, and T-... ids as tickets://<id>. Pass the question.",
systemInstructions: RESEARCH_INSTRUCTIONS,
inputSchema: {
question: z.string().describe("The question, like: Why do enterprise customers keep hitting login failures after the SSO change?"),
},
outputSchema: {
claims: z
.array(
z.object({
statement: z.string(),
sources: z.array(z.string().regex(/^(KB|T)-\d+$/)).min(1),
}),
)
.min(1),
},
execution: { timeout: 60_000 },
llm: { adapter: writeModel },
})
export class ResearchAgent extends AgentContext {
async execute(input: { question: string }) {
const found = await this.callTool("invoke_find_sources", { question: input.question });
const { sourceIds } = found.structuredContent as { sourceIds: string[] };
const knowledgeBase = this.get(KnowledgeBase);
const sources = sourceIds.flatMap((id) => knowledgeBase.find(id) ?? []);
if (sources.length === 0) {
this.fail(new PublicMcpError("The help desk's articles and closed tickets don't cover that.", "NO_SOURCES_FOUND"));
}
const answer = (await super.execute({ question: input.question, sources })) as Answer;
const givenIds = new Set(sources.map((source) => source.id));
// An answer in the wrong format has no claims, and goes on to fail outputSchema.
const citedIds = answer.claims?.flatMap((claim) => claim.sources) ?? [];
const unknownIds = citedIds.filter((id) => !givenIds.has(id));
if (unknownIds.length > 0) {
this.fail(new PublicMcpError(`The answer cites ${unknownIds.join(", ")}, which the search didn't find.`, "UNKNOWN_SOURCE"));
}
return answer;
}
}Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
In the Call tab, ask invoke_research a smaller question, like "How long does a session last?": the answer is one claim, from KB-4 and T-48. Ask "Where do I park at the office?" and the call is refused with NO_SOURCES_FOUND, before the research agent's own model is asked. The Capabilities tab lists the two resource templates, kb://{id} and tickets://{id}.
How it fits together
- A client calls
invoke_researchwith a question. It's the only tool the client sees:find_sourcesis hidden, and the search tools belong to it. research'sexecute()hands the question tofind_sourceswiththis.callTool().find_sourcesruns its own loop. Its model asks forsearch_articlesandsearch_ticketsin one turn, and each search sends the client a log message when it finishes. On the next turn, the model answers with the ids worth reading.- Back in
research, every id is looked up. An id that isn't an article or a ticket is dropped, and if none are left the call fails withNO_SOURCES_FOUND, without askingresearch's own model. researchruns its own model loop, one turn and no tools, with the question and the full text of each source in the message. The model answers with claims, each with the ids it came from.execute()checks that every id the claims cite is one it sent, and fails withUNKNOWN_SOURCEif not. FrontMCP then checks the answer againstoutputSchemaand returns it asstructuredContent.- The client opens what a claim cites by reading a resource:
kb://KB-2for an article,tickets://T-43for a ticket.
The client's model gets one tool and two resource templates. find_sources's model gets two tools, and its answer is only a list of ids. research's model gets no tools at all, only the text it may cite. Each model has one job, and the code between them does the checking.
The files
knowledge.ts: articles, closed tickets and a search
One provider, GLOBAL, so the search tools, the resources and research's execute() all read the same instance. It holds six articles and six closed tickets about signing in: the SSO change (KB-1, KB-2, T-41, T-43), a clock problem (KB-5, T-52), and a few about passwords, sessions and refunds that don't belong to the question, so the search has something to leave out. searchArticles() and searchTickets() match words and return the best three, and find(id) looks an id up in both. In a real server this is where your knowledge base's and ticket system's own search goes: nothing else in the example cares how it works.
main.ts: one app
@App({
id: "help-desk",
name: "Help Desk",
agents: [ResearchAgent, FindSourcesAgent],
resources: [ArticleResource, ClosedTicketResource],
providers: [KnowledgeBase],
})The app lists both agents, both resource templates and the provider. The search tools run inside find_sources, and an agent's tools see the app's providers as the app's own tools do, so the searches, the resources and research all get the same KnowledgeBase: see Sharing a provider with the agent's tools. A client sees one tool, invoke_research, and two templates in resources/templates/list, which the first test checks.
search.tools.ts: two tools only the searcher has
const articles = this.get(KnowledgeBase).searchArticles(query);
await this.notify(`Articles for "${query}": ${articles.length} found`);
return { articles: articles.map(({ id, title }) => ({ id, title })) };They're in find_sources's tools and nowhere else, so only its model can call them. They return ids and titles, not text: finding is one job, reading is another, and the writer gets the text later.
Each one ends with a log message, and that's how the client hears what the search found. find_sources adds one of its own before each search, Calling tool: search_articles, as every agent does for each tool call. They reach the client even though research started find_sources with this.callTool(): a called agent and its tools report to the client of the outer call, as called tools do. The sixth test reads them. They're log messages rather than progress() calls because a tool only knows its own step, and progress has to keep going up across the whole call. Reporting Progress and Logs covers both.
find-sources.agent.ts: the narrower agent
systemInstructionsgive the steps and the exact JSON to answer with, and tell the model never to list an id a search didn't return. The code doesn't rely on that:researchchecks anyway.outputSchema: { sourceIds: z.array(z.string()) }givesresearcha list, not prose. A searcher that answers in words fails its own call, instead of handing prose toresearch.hideFromDiscovery: truetakesfind_sourcesout oftools/list, becauseresearchis the way in. It hides the agent and doesn't lock it: the second test callsinvoke_find_sourcesby name. An agent whose work needs protecting still needs authorization like any tool.executionallows four turns where two are enough, so a searcher that wants a second round of searches has room, and half a minute.
research.agent.ts: the coordinator
The options divide the audience, as they do for triage:
descriptionis for the client's model. It says what comes back, and how to open a cited id, since an id alone doesn't say:kb://<id>for articles,tickets://<id>for tickets.systemInstructionsare forresearch's model: answer only from the sources, in JSON, the most important claim first, and every claim with a source.outputSchemamakes the answer checkable:
outputSchema: {
claims: z
.array(
z.object({
statement: z.string(),
sources: z.array(z.string().regex(/^(KB|T)-\d+$/)).min(1),
}),
)
.min(1),
},The client gets structuredContent in this shape, or an error: at least one claim, at least one source on each, and every source looks like an id. A claim with no source fails the call with INVALID_OUTPUT, as the last test shows. See Returning structured output. What the schema can't know is which ids exist. That's execute()'s job.
The agent has no tools: the sources in its message are all its model has to work from. execute() works in three steps:
const found = await this.callTool("invoke_find_sources", { question: input.question });
const { sourceIds } = found.structuredContent as { sourceIds: string[] };
const knowledgeBase = this.get(KnowledgeBase);
const sources = sourceIds.flatMap((id) => knowledgeBase.find(id) ?? []);- Hand off.
this.callTool("invoke_find_sources", …)runs the other agent as if a client had called it, and the answer is itsstructuredContent, already checked by itsoutputSchema. Handing off in code, rather than offeringfind_sourcestoresearch's model withagents: [...], makes sure every question is searched first: Agents That Call Agents compares the two. Iffind_sourcesfails,this.callTool()throws and so does the research call: When the other agent fails shows what to answer instead, and the second idea below does it. - Check what comes back. A model's list of ids is a claim like any other.
find()returns nothing for an id that isn't an article or a ticket, andflatMapdrops it, as the eighth test shows with an id the searcher makes up. If nothing is left, the call fails withNO_SOURCES_FOUNDbeforeresearch's own model is asked: a question the help desk has nothing on costs the searcher's turns and no more. - Write, then check.
super.execute({ question, sources })runsresearch's own loop. The input has noquery,message,promptorinputfield, so the model is sent it whole, as JSON: the question, and each source's text. Afterwards, every id the claims cite has to be one that was sent. A model that cites another fails the call withUNKNOWN_SOURCE, and doesn't get its answer trimmed: dropping the claim would change what the answer says, and returning it would send a support agent to a source that doesn't exist. An answer in the wrong format has noclaims, so this check lets it through to failoutputSchema.
Adding steps before and after the loop has more on overriding execute().
sources.resources.ts: opening what an answer cites
Two resource templates: kb://{id} returns an article as markdown, and tickets://{id} returns a closed ticket as JSON. They're resources because the person decides to open a citation: the client reads the URI and shows it, and no model has to decide anything. Exposing Data with Resources covers templates. An id that doesn't exist throws ResourceNotFoundError, so the client gets a -32602 "not found" error and not a server failure, as that lesson explains.
They aren't how research reads its sources: research puts each source's text in the model's message itself, so the model reads exactly the sources find_sources found, in one turn, without asking for them. An agent's model can read resources too, but only those the agent declares in its own resources option, with a read_resource tool call per read (since 1.9.4). These templates are the app's, for clients, and no agent's model is offered them.
model.example.ts: two scripted models
const CLAIMS_IT_CAN_WRITE = [
{
needs: ["KB-1", "KB-2"],
statement: "Since release 3.2, SSO sign-in only accepts responses signed by …",
},
// ...
];Each agent has its own llm, so each gets its own stand-in. searchModel picks search words from a short list, asks for both searches in one turn, then lists every id they returned. writeModel writes a claim only when it was sent every source the claim needs, so what it writes depends on what research handed it, as a real model's would. The sentences are ones it has seen, where a real model writes its own.
scenario.slip makes either model make a mistake that real models make, so the tests can check what research does about it. turns keeps what each model was sent. Your First Agent looks at the prompt in detail.
research.test.ts
The tests check what a client sees: one tool and two templates, the answer and its citations, every cited id opening as a resource, and the searches being reported while the agents work. Then they check what each model was sent: two searcher turns and one writer turn, with the writer sent only what was found. The last six are the failures: an id the searcher invents never reaches the writer, a question with nothing on it is refused before the writer is asked, a claim that cites an id the writer wasn't given is refused, a searcher that answers in words fails the call, an answer in a code fence fails outputSchema, and so does a claim with no source.
All the tests share one server, in order, and each slip is switched off again at the end of its test. The notifications test reads what the Playground's client asked for: under frontmcp test, the client has to ask for log messages first, as Checking notifications shows. Testing Your Server covers the test API.
Running it with a real model
Give both agents a real model, and keep everything else:
@Agent({
// ...
llm: { provider: "openai", model: "gpt-5", apiKey: { env: "OPENAI_API_KEY" } },
})yarn add openai
export OPENAI_API_KEY=sk-...find-sources.agent.ts gets the same llm, or another model: each agent has its own. Then delete model.example.ts, its imports, and the tests that use scenario and turns. Tests that check the exact claims go too, since a real model words them its own way. The ones on the tools, the templates and on every cited id opening as a resource still make sense.
The Playground can't run this: it has no network and no key. The server was run with the real openai package and these two llm lines, against a local stand-in for the API that answers like Chat Completions. That checks FrontMCP's side: what it sends, how it runs the tools the stand-in asks for, and how it reads the replies. It doesn't check a real model: how well one picks search words, reads the sources and cites them wasn't tested. Connecting a Real Model covers the providers, and what to know about each:
- Several model calls per question. The searcher is asked at least twice, once for the searches and once for the ids, and again for every extra round of searches, up to
maxIterations: 4.research's model is asked once more, with the text of every source in the message. The local API's log shows the searcher's requests offering its two tools, and the writer's offering none. - Answers vary, and the checks catch it. A searcher that answers in words instead of JSON fails its own
outputSchema, and the research call fails with that same error,INVALID_OUTPUTandTool output validation failed (output does not match outputSchema at sourceIds), as the eleventh test shows. A writer that wraps its JSON in a code fence failsoutputSchematoo, because FrontMCP only reads text that starts with{or[as JSON. Tell the model not to use code fences, or strip them from its reply. - A failed call can be repeated. Everything this agent does is a read, so unlike triage there's nothing to undo when a model cites an id it wasn't sent: the client asks again.
- Every call costs model calls, and their time.
@Agent's ownrateLimitcaps how often clients can run it, so set one before you expose the agent. - The search is yours.
KnowledgeBaseis where the help desk's real search goes. The tools only pass the model's words on to it.
To strip a code fence, override parseAgentResponse() on ResearchAgent:
protected parseAgentResponse(content: string | null) {
return super.parseAgentResponse(content?.replace(/^```(?:json)?\s*|\s*```$/g, "") ?? null);
}Ideas to try
Each of these is a change to the Playground above. Add a test for each.
- Let
researchsearch again when fewer than two sources come back: callinvoke_find_sourcesa second time, and never a third. Have the stand-in use other words the second time, and check how many turns the searcher's model was asked. - Catch a searcher that fails. Add a
slipthat makessearchModelthrow, catch it inexecute(), and fail with aPublicMcpErrorwhose message tells the client's model what to do. Check what a client sees. - Split the search in two: a
find_articlesagent and afind_ticketsagent, each with one tool, called one after the other fromexecute(). Check that the writer is sent both agents' sources. - Let the client narrow by plan: add an optional
plantoresearch's input, pass it tofind_sources, and keep the other plan's tickets out of the results. Check that an enterprise question never cites T-46 or T-55.