Health checks and metrics
Orchestrators and load balancers ask a server two questions: is the process alive, and should it get traffic now? FrontMCP answers the first at /healthz and the second at /readyz, and serves process gauges and your counters at /metrics for Prometheus. What each path answers depends on how the server runs: FrontMCP's Node server runs readiness probes, a serverless function has no /readyz, and a Worker has /healthz and /readyz only when you turn the second on. This page is about wiring them into a deployment; Observability and telemetry has every option of health and metrics.
@FrontMcp({ info, apps, health: { probes, readyz: { timeoutMs } }, metrics: { enabled: true, auth: "token" } })
Reference
What each runtime answers
Node server: node, distributed, cli serve, frontmcp dev | Serverless: vercel, lambda | Fetch handler: cloudflare, createFetchHandler() | |
|---|---|---|---|
GET /healthz | 200 { status: "ok", server, runtime: { platform, runtime, deployment, env }, uptime }. No checks. | The same, with "deployment": "serverless" | 200 { status: "ok", server, transport: "web-fetch" }. No checks. |
GET /health | The same as /healthz | The same as /healthz | The same as /healthz |
GET /readyz | Runs every probe: 200 "status": "ready", or 503 "not_ready" | 404: off in serverless mode | Off on a Worker (404) until health.readyz.enabled: true; in createFetchHandler() on Node it's on. When on, it runs every probe like the Node server, with "transport": "web-fetch" in the answer. |
GET /metrics | With metrics.enabled | Not checked for this page | With metrics.enabled, since 1.9, and on a Worker since 1.9.2, without the process gauges: see Cloudflare Workers. |
health options | Applied | Not checked for this page | Applied: enabled, healthzPath, readyzPath, probes, includeDetails, readyz |
Cache-Control | no-cache, no-transform | no-cache, no-transform | no-store |
health.enabled: false removes /healthz, /health and /readyz on every runtime. A custom healthzPath replaces /healthz only: /health still answers.
On every runtime:
- Health checks don't need credentials. With
static,transparent,localorremoteauth,/healthzand/readyzstill answer200without a token, while an MCP request without one gets401. - They go through the host check. A server that checks
Hostheaders (host checks) answers a health check sent to another host name, like a pod's IP address, with403Invalid Host header. - They don't see missing secrets. A production server without
MCP_SESSION_SECRETanswers/healthzwith200, and theinitializeof clients before MCP 2026-07-28 with500and{"error":"server_misconfigured","code":"SESSION_SECRET_REQUIRED",…}. Only a Worker withoutJWT_SECRETfails its health checks too (production secrets). - On a Worker, they fail when the server can't be built. Every request, health checks included, answers
500 server_misconfiguredor503 server_unavailablewithRetry-Afteruntil a retry builds it: see Cloudflare Workers.
/readyz on the Node server
It runs FrontMCP's own session-store probe, which pings Redis or SQLite when sessions are stored there, and the probes you add, all at once, each with readyz.timeoutMs (5 seconds by default). Any unhealthy probe, one that throws, or one that times out makes it 503; degraded counts as ready.
{"status":"ready","totalLatencyMs":0,"catalog":{"toolsHash":"…","toolCount":1,"resourceCount":0,"promptCount":0,"skillCount":0,"agentCount":0},"probes":{"session-store":{"status":"healthy","latencyMs":0},"tickets-db":{"status":"healthy","latencyMs":2}}}
Each remote app adds a probe too, named remote:<app id> (the id is the app's name unless you gave one), like remote:status. It reports what FrontMCP's own check of the remote, every 30 seconds, last found. Before the first check it's degraded with "state": "unknown", which counts as ready, so /readyz isn't 503 while a server with a remote app starts. Started against a remote that was up, the first /readyz after the server listened already showed it healthy.
In production the probes field is left out, so the status is all a caller learns; includeDetails: true puts it back. catalog.toolsHash is the same on every instance with the same tools: High availability compares it across replicas. When Redis is down, the session-store probe is unhealthy and /readyz answers 503 until it's back: also at startup, where FrontMCP now retries the connection with a growing delay instead of giving up for good. See Redis.
/metrics
metrics: { enabled: true } needs two packages, and the server doesn't start without either:
npm install @frontmcp/observability @opentelemetry/sdk-trace-baseWithout @opentelemetry/sdk-trace-base, startup logs Failed to wire /metrics endpoint Error: Cannot find module '@opentelemetry/sdk-trace-base' and the process exits. With them, GET /metrics answers Prometheus text (Content-Type: text/plain; charset=utf-8; version=0.0.4, Cache-Control: no-store): process and Node gauges, and the counters you create. With auth: "token", a request without Authorization: Bearer <FRONTMCP_METRICS_TOKEN> gets 401, one with the wrong token 403. metrics has the options and the output.
createFetchHandler() serves /metrics the same way since 1.9, with the same 401 and 403, and Content-Type: text/plain; version=0.0.4; charset=utf-8; in 1.8.7 it answered 404. It needs the same two packages: without the second, createFetchHandler() rejects with Cannot find module '@opentelemetry/sdk-trace-base'. A Worker serves it too since 1.9.2, with Cache-Control: no-store, and only the counters: it has no process to measure, so it leaves the frontmcp_process_* gauges out (in 1.9.2 it listed them, all 0).
Caveats
/healthzchecks nothing: an instance that can't reach its database still answers200. Put what the server can't work without inhealth.probes, and use/readyzfor readiness.- A
GETto the MCP endpoint isn't a health check: without a session, FrontMCP's Node server answers it with Express's404page (Cannot GET /), and a fetch handler with406 Not Acceptableunless it sendsAccept: text/event-stream. - The Kubernetes and load-balancer settings below weren't run on Kubernetes. With FrontMCP 1.9.3: the Docker health check of the image
frontmcp createwrites,/healthbeside a customhealthzPath, and/metricson a Worker and fromcreateFetchHandler()in Node. With 1.9.2: the fetch handler's health paths withcreateFetchHandler()(the Playground below) andwrangler dev, and/readyzon the Node server./metricson the Node server was last run with 1.9.1, and the Prometheus scrape (Prometheus 3.5 in Docker, with the token in a credentials file) with 1.8.7.
Usage
Checking health from inside a container
node:24-slim has no curl, so the Dockerfile frontmcp create writes checks with Node itself:
HEALTHCHECK --interval=30s --timeout=5s --start-period=15s --retries=3 \
CMD node -e "fetch('http://127.0.0.1:' + (process.env.PORT || 3000) + '/health').then((r) => process.exit(r.ok ? 0 : 1), () => process.exit(1))"docker inspect reported the container healthy after the first check. It asks /health, which stays put when health.healthzPath moves /healthz: with healthzPath: "/live", /health and /live answered 200 and /healthz 404. (In 1.9.2 the check asked /healthz.) The Nx server generator's Dockerfile has the same check since 1.9.3. With FRONTMCP_ALLOWED_HOSTS set, include 127.0.0.1:3000 in it, or this check gets 403 and the container is reported unhealthy. Node.js and Docker has the whole Dockerfile.
Pointing Kubernetes at the probes
Liveness at /healthz, readiness at /readyz, with the probe's timeout longer than readyz.timeoutMs:
containers:
- name: help-desk
image: registry.example.com/help-desk:1.2.0
env:
- name: FRONTMCP_BIND_ADDRESS
value: all
- name: FRONTMCP_ALLOWED_HOSTS
value: desk.example.com
ports:
- containerPort: 3000
livenessProbe:
httpGet:
path: /healthz
port: 3000
httpHeaders:
- name: Host
value: desk.example.com
periodSeconds: 10
readinessProbe:
httpGet:
path: /readyz
port: 3000
httpHeaders:
- name: Host
value: desk.example.com
periodSeconds: 15
timeoutSeconds: 6The kubelet calls the pod by its IP address, which isn't in FRONTMCP_ALLOWED_HOSTS, so the probes send the public host name instead. Without FRONTMCP_ALLOWED_HOSTS, a server on all interfaces checks no host, and the headers aren't needed.
Scraping metrics with Prometheus
@FrontMcp({
info: { name: "help-desk", version: "1.0.0" },
apps: [HelpDesk],
metrics: { enabled: true, auth: "token" }, // FRONTMCP_METRICS_TOKEN
})
export default class Server {}scrape_configs:
- job_name: help-desk
metrics_path: /metrics
authorization:
type: Bearer
credentials_file: /etc/prometheus/metrics-token
static_configs:
- targets: ["help-desk:3000"]With the same token in FRONTMCP_METRICS_TOKEN and in the credentials file, Prometheus reports the target up, and frontmcp_process_resident_memory_bytes{job="help-desk"} returns the server's memory. /metrics is on the same port as MCP; keep it off the public internet at your proxy, or rely on the token.
Health checks behind authentication, on a Worker
A load balancer or uptime monitor can check a Worker, or any server behind createFetchHandler(), without a token. The host check applies to them, and so do health's options:
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";
const config = {
info: { name: "help-desk", version: "1.0.0" },
apps: [HelpDesk],
http: { entryPath: "/mcp" },
auth: { mode: "static" as const, tokens: ["sk_desk_1"] },
health: { readyz: { enabled: true } }, // off by default on a Worker
};
const get = async (handler: (request: Request) => Promise<Response>, path: string) => handler(new Request(`https://desk.example.com${path}`));
test("health checks need no token, though MCP calls do", async () => {
const handler = await FrontMcpInstance.createFetchHandler(config);
for (const path of ["/healthz", "/health"]) {
const res = await get(handler, path);
expect(res.status).toBe(200);
expect(res.headers.get("cache-control")).toBe("no-store");
expect(await res.json()).toEqual({ status: "ok", server: { name: "help-desk", version: "1.0.0" }, transport: "web-fetch" });
}
const ready = await get(handler, "/readyz");
expect(ready.status).toBe(200);
expect(await ready.json()).toMatchObject({ status: "ready", transport: "web-fetch", catalog: { toolCount: 1 } });
const call = await handler(
new Request("https://desk.example.com/mcp", {
method: "POST",
headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/list" },
body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/list", params: { _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } } }),
}),
);
expect(call.status).toBe(401);
});
test("a failing probe makes `/readyz` answer 503, and `/healthz` stays 200", async () => {
let up = true;
const probes = [{ name: "tickets-db", check: async () => (up ? { status: "healthy" as const } : { status: "unhealthy" as const, error: "no connection" }) }];
const handler = await FrontMcpInstance.createFetchHandler({ ...config, health: { readyz: { enabled: true }, probes, includeDetails: true } });
expect((await get(handler, "/readyz")).status).toBe(200);
up = false;
const down = await get(handler, "/readyz");
expect(down.status).toBe(503);
expect(await down.json()).toMatchObject({ status: "not_ready", probes: { "tickets-db": { status: "unhealthy", error: "no connection" } } });
expect((await get(handler, "/healthz")).status).toBe(200);
});
test("`health` sets the paths, and `enabled: false` removes them", async () => {
const renamed = await FrontMcpInstance.createFetchHandler({ ...config, health: { healthzPath: "/live", readyzPath: "/ready", readyz: { enabled: true } } });
expect((await get(renamed, "/live")).status).toBe(200);
expect((await get(renamed, "/ready")).status).toBe(200);
expect((await get(renamed, "/healthz")).status).toBe(404);
expect((await get(renamed, "/health")).status).toBe(200); // still an alias of the liveness path
const off = await FrontMcpInstance.createFetchHandler({ ...config, health: { enabled: false } });
for (const path of ["/healthz", "/health", "/readyz"]) expect((await get(off, path)).status).toBe(404);
});
test("a check sent to another host name is refused", async () => {
const guarded = await FrontMcpInstance.createFetchHandler({
...config,
http: { entryPath: "/mcp", security: { dnsRebindingProtection: { enabled: true, allowedHosts: ["desk.example.com"] } } },
});
expect((await guarded(new Request("http://10.0.3.7:8787/healthz"))).status).toBe(403);
expect((await guarded(new Request("https://desk.example.com/healthz"))).status).toBe(200);
});Starting FrontMCP in your browser…
FrontMCP starts when this example comes into view.
In production /readyz leaves probes out unless includeDetails is true, as on the Node server; the test sets it, so it passes with any NODE_ENV. A probe that doesn't answer within health.readyz.timeoutMs is unhealthy with Probe timed out after …ms, and /readyz answers 503. On FrontMCP's Node server enabled: false and the paths behave the same way. Answering health checks has every option.
Troubleshooting
/readyz answers 404
The server runs as a Vercel or Lambda function, where readiness checks are off, or as a Worker, where they're off until you set health: { readyz: { enabled: true } }, or health.enabled is false. Use /healthz.
/health or /healthz answers 404
health.enabled is false, which removes every health path, or healthzPath moved the liveness check: /healthz is gone then, and /health answers in its place.
Health checks get 403 {"error":"Forbidden","message":"Invalid Host header"}
The check calls the server by a name that isn't allowed, like 127.0.0.1:3000 or a pod's IP. Add that name to FRONTMCP_ALLOWED_HOSTS, or send the public name in a Host header, as in the Kubernetes probes.
/readyz answers 503, with no details
In production the answer leaves probes out. Set health: { includeDetails: true } to see which probe failed, or check that Redis is reachable.
remote:… is degraded right after startup
The server has a remote app, and FrontMCP hasn't checked it yet. degraded counts as ready, so /readyz answers 200, and the probe turns healthy after the first check if the remote answers. (In 1.8.6 it was unhealthy until the first check and /readyz answered 503 for about 30 seconds.)
Failed to wire /metrics endpoint Error: Cannot find module '@opentelemetry/sdk-trace-base'
metrics.enabled with @frontmcp/observability installed but not its peer. Install @opentelemetry/sdk-trace-base. Without @frontmcp/observability it's Cannot find module '@frontmcp/observability' (more).
The health check passes, but older clients' initialize answers 500
The server is in production without MCP_SESSION_SECRET. Health checks, and clients on MCP 2026-07-28, don't notice. Set the secret.
Prometheus reports the target down with 401 or 403
401: the scrape sends no Authorization header. 403: its token isn't FRONTMCP_METRICS_TOKEN. Check the authorization block and the credentials file.