Health checks and metrics

Orchestrators and load balancers ask a server two questions: is the process alive, and should it get traffic now? FrontMCP answers the first at /healthz and the second at /readyz, and serves process gauges and your counters at /metrics for Prometheus. What each path answers depends on how the server runs: FrontMCP's Node server runs readiness probes, a serverless function has no /readyz, and a Worker has /healthz and /readyz only when you turn the second on. This page is about wiring them into a deployment; Observability and telemetry has every option of health and metrics.

@FrontMcp({ info, apps, health: { probes, readyz: { timeoutMs } }, metrics: { enabled: true, auth: "token" } })

Reference

What each runtime answers

Node server: node, distributed, cli serve, frontmcp devServerless: vercel, lambdaFetch handler: cloudflare, createFetchHandler()
GET /healthz200 { status: "ok", server, runtime: { platform, runtime, deployment, env }, uptime }. No checks.The same, with "deployment": "serverless"200 { status: "ok", server, transport: "web-fetch" }. No checks.
GET /healthThe same as /healthzThe same as /healthzThe same as /healthz
GET /readyzRuns every probe: 200 "status": "ready", or 503 "not_ready"404: off in serverless modeOff on a Worker (404) until health.readyz.enabled: true; in createFetchHandler() on Node it's on. When on, it runs every probe like the Node server, with "transport": "web-fetch" in the answer.
GET /metricsWith metrics.enabledNot checked for this pageWith metrics.enabled, since 1.9, and on a Worker since 1.9.2, without the process gauges: see Cloudflare Workers.
health optionsAppliedNot checked for this pageApplied: enabled, healthzPath, readyzPath, probes, includeDetails, readyz
Cache-Controlno-cache, no-transformno-cache, no-transformno-store

health.enabled: false removes /healthz, /health and /readyz on every runtime. A custom healthzPath replaces /healthz only: /health still answers.

On every runtime:

  • Health checks don't need credentials. With static, transparent, local or remote auth, /healthz and /readyz still answer 200 without a token, while an MCP request without one gets 401.
  • They go through the host check. A server that checks Host headers (host checks) answers a health check sent to another host name, like a pod's IP address, with 403 Invalid Host header.
  • They don't see missing secrets. A production server without MCP_SESSION_SECRET answers /healthz with 200, and the initialize of clients before MCP 2026-07-28 with 500 and {"error":"server_misconfigured","code":"SESSION_SECRET_REQUIRED",…}. Only a Worker without JWT_SECRET fails its health checks too (production secrets).
  • On a Worker, they fail when the server can't be built. Every request, health checks included, answers 500 server_misconfigured or 503 server_unavailable with Retry-After until a retry builds it: see Cloudflare Workers.

/readyz on the Node server

It runs FrontMCP's own session-store probe, which pings Redis or SQLite when sessions are stored there, and the probes you add, all at once, each with readyz.timeoutMs (5 seconds by default). Any unhealthy probe, one that throws, or one that times out makes it 503; degraded counts as ready.

{"status":"ready","totalLatencyMs":0,"catalog":{"toolsHash":"…","toolCount":1,"resourceCount":0,"promptCount":0,"skillCount":0,"agentCount":0},"probes":{"session-store":{"status":"healthy","latencyMs":0},"tickets-db":{"status":"healthy","latencyMs":2}}}

Each remote app adds a probe too, named remote:<app id> (the id is the app's name unless you gave one), like remote:status. It reports what FrontMCP's own check of the remote, every 30 seconds, last found. Before the first check it's degraded with "state": "unknown", which counts as ready, so /readyz isn't 503 while a server with a remote app starts. Started against a remote that was up, the first /readyz after the server listened already showed it healthy.

In production the probes field is left out, so the status is all a caller learns; includeDetails: true puts it back. catalog.toolsHash is the same on every instance with the same tools: High availability compares it across replicas. When Redis is down, the session-store probe is unhealthy and /readyz answers 503 until it's back: also at startup, where FrontMCP now retries the connection with a growing delay instead of giving up for good. See Redis.

/metrics

metrics: { enabled: true } needs two packages, and the server doesn't start without either:

npm install @frontmcp/observability @opentelemetry/sdk-trace-base

Without @opentelemetry/sdk-trace-base, startup logs Failed to wire /metrics endpoint Error: Cannot find module '@opentelemetry/sdk-trace-base' and the process exits. With them, GET /metrics answers Prometheus text (Content-Type: text/plain; charset=utf-8; version=0.0.4, Cache-Control: no-store): process and Node gauges, and the counters you create. With auth: "token", a request without Authorization: Bearer <FRONTMCP_METRICS_TOKEN> gets 401, one with the wrong token 403. metrics has the options and the output.

createFetchHandler() serves /metrics the same way since 1.9, with the same 401 and 403, and Content-Type: text/plain; version=0.0.4; charset=utf-8; in 1.8.7 it answered 404. It needs the same two packages: without the second, createFetchHandler() rejects with Cannot find module '@opentelemetry/sdk-trace-base'. A Worker serves it too since 1.9.2, with Cache-Control: no-store, and only the counters: it has no process to measure, so it leaves the frontmcp_process_* gauges out (in 1.9.2 it listed them, all 0).

Caveats

  • /healthz checks nothing: an instance that can't reach its database still answers 200. Put what the server can't work without in health.probes, and use /readyz for readiness.
  • A GET to the MCP endpoint isn't a health check: without a session, FrontMCP's Node server answers it with Express's 404 page (Cannot GET /), and a fetch handler with 406 Not Acceptable unless it sends Accept: text/event-stream.
  • The Kubernetes and load-balancer settings below weren't run on Kubernetes. With FrontMCP 1.9.3: the Docker health check of the image frontmcp create writes, /health beside a custom healthzPath, and /metrics on a Worker and from createFetchHandler() in Node. With 1.9.2: the fetch handler's health paths with createFetchHandler() (the Playground below) and wrangler dev, and /readyz on the Node server. /metrics on the Node server was last run with 1.9.1, and the Prometheus scrape (Prometheus 3.5 in Docker, with the token in a credentials file) with 1.8.7.

Usage

Checking health from inside a container

node:24-slim has no curl, so the Dockerfile frontmcp create writes checks with Node itself:

ci/Dockerfile
HEALTHCHECK --interval=30s --timeout=5s --start-period=15s --retries=3 \
  CMD node -e "fetch('http://127.0.0.1:' + (process.env.PORT || 3000) + '/health').then((r) => process.exit(r.ok ? 0 : 1), () => process.exit(1))"

docker inspect reported the container healthy after the first check. It asks /health, which stays put when health.healthzPath moves /healthz: with healthzPath: "/live", /health and /live answered 200 and /healthz 404. (In 1.9.2 the check asked /healthz.) The Nx server generator's Dockerfile has the same check since 1.9.3. With FRONTMCP_ALLOWED_HOSTS set, include 127.0.0.1:3000 in it, or this check gets 403 and the container is reported unhealthy. Node.js and Docker has the whole Dockerfile.

Pointing Kubernetes at the probes

Liveness at /healthz, readiness at /readyz, with the probe's timeout longer than readyz.timeoutMs:

deployment.yaml
containers:
  - name: help-desk
    image: registry.example.com/help-desk:1.2.0
    env:
      - name: FRONTMCP_BIND_ADDRESS
        value: all
      - name: FRONTMCP_ALLOWED_HOSTS
        value: desk.example.com
    ports:
      - containerPort: 3000
    livenessProbe:
      httpGet:
        path: /healthz
        port: 3000
        httpHeaders:
          - name: Host
            value: desk.example.com
      periodSeconds: 10
    readinessProbe:
      httpGet:
        path: /readyz
        port: 3000
        httpHeaders:
          - name: Host
            value: desk.example.com
      periodSeconds: 15
      timeoutSeconds: 6

The kubelet calls the pod by its IP address, which isn't in FRONTMCP_ALLOWED_HOSTS, so the probes send the public host name instead. Without FRONTMCP_ALLOWED_HOSTS, a server on all interfaces checks no host, and the headers aren't needed.

Scraping metrics with Prometheus

main.ts
@FrontMcp({
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDesk],
  metrics: { enabled: true, auth: "token" }, // FRONTMCP_METRICS_TOKEN
})
export default class Server {}
prometheus.yml
scrape_configs:
  - job_name: help-desk
    metrics_path: /metrics
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/metrics-token
    static_configs:
      - targets: ["help-desk:3000"]

With the same token in FRONTMCP_METRICS_TOKEN and in the credentials file, Prometheus reports the target up, and frontmcp_process_resident_memory_bytes{job="help-desk"} returns the server's memory. /metrics is on the same port as MCP; keep it off the public internet at your proxy, or rely on the token.

Health checks behind authentication, on a Worker

A load balancer or uptime monitor can check a Worker, or any server behind createFetchHandler(), without a token. The host check applies to them, and so do health's options:

Open
import { test, expect } from "@frontmcp/testing";
import { FrontMcpInstance } from "@frontmcp/sdk";
import { HelpDesk } from "./help-desk.app";

const config = {
  info: { name: "help-desk", version: "1.0.0" },
  apps: [HelpDesk],
  http: { entryPath: "/mcp" },
  auth: { mode: "static" as const, tokens: ["sk_desk_1"] },
  health: { readyz: { enabled: true } }, // off by default on a Worker
};

const get = async (handler: (request: Request) => Promise<Response>, path: string) => handler(new Request(`https://desk.example.com${path}`));

test("health checks need no token, though MCP calls do", async () => {
  const handler = await FrontMcpInstance.createFetchHandler(config);
  for (const path of ["/healthz", "/health"]) {
    const res = await get(handler, path);
    expect(res.status).toBe(200);
    expect(res.headers.get("cache-control")).toBe("no-store");
    expect(await res.json()).toEqual({ status: "ok", server: { name: "help-desk", version: "1.0.0" }, transport: "web-fetch" });
  }
  const ready = await get(handler, "/readyz");
  expect(ready.status).toBe(200);
  expect(await ready.json()).toMatchObject({ status: "ready", transport: "web-fetch", catalog: { toolCount: 1 } });
  const call = await handler(
    new Request("https://desk.example.com/mcp", {
      method: "POST",
      headers: { "content-type": "application/json", "mcp-protocol-version": "2026-07-28", "mcp-method": "tools/list" },
      body: JSON.stringify({ jsonrpc: "2.0", id: 1, method: "tools/list", params: { _meta: { "io.modelcontextprotocol/protocolVersion": "2026-07-28" } } }),
    }),
  );
  expect(call.status).toBe(401);
});

test("a failing probe makes `/readyz` answer 503, and `/healthz` stays 200", async () => {
  let up = true;
  const probes = [{ name: "tickets-db", check: async () => (up ? { status: "healthy" as const } : { status: "unhealthy" as const, error: "no connection" }) }];
  const handler = await FrontMcpInstance.createFetchHandler({ ...config, health: { readyz: { enabled: true }, probes, includeDetails: true } });
  expect((await get(handler, "/readyz")).status).toBe(200);
  up = false;
  const down = await get(handler, "/readyz");
  expect(down.status).toBe(503);
  expect(await down.json()).toMatchObject({ status: "not_ready", probes: { "tickets-db": { status: "unhealthy", error: "no connection" } } });
  expect((await get(handler, "/healthz")).status).toBe(200);
});

test("`health` sets the paths, and `enabled: false` removes them", async () => {
  const renamed = await FrontMcpInstance.createFetchHandler({ ...config, health: { healthzPath: "/live", readyzPath: "/ready", readyz: { enabled: true } } });
  expect((await get(renamed, "/live")).status).toBe(200);
  expect((await get(renamed, "/ready")).status).toBe(200);
  expect((await get(renamed, "/healthz")).status).toBe(404);
  expect((await get(renamed, "/health")).status).toBe(200); // still an alias of the liveness path

  const off = await FrontMcpInstance.createFetchHandler({ ...config, health: { enabled: false } });
  for (const path of ["/healthz", "/health", "/readyz"]) expect((await get(off, path)).status).toBe(404);
});

test("a check sent to another host name is refused", async () => {
  const guarded = await FrontMcpInstance.createFetchHandler({
    ...config,
    http: { entryPath: "/mcp", security: { dnsRebindingProtection: { enabled: true, allowedHosts: ["desk.example.com"] } } },
  });
  expect((await guarded(new Request("http://10.0.3.7:8787/healthz"))).status).toBe(403);
  expect((await guarded(new Request("https://desk.example.com/healthz"))).status).toBe(200);
});

Starting FrontMCP in your browser…

FrontMCP starts when this example comes into view.

In production /readyz leaves probes out unless includeDetails is true, as on the Node server; the test sets it, so it passes with any NODE_ENV. A probe that doesn't answer within health.readyz.timeoutMs is unhealthy with Probe timed out after …ms, and /readyz answers 503. On FrontMCP's Node server enabled: false and the paths behave the same way. Answering health checks has every option.


Troubleshooting

/readyz answers 404

The server runs as a Vercel or Lambda function, where readiness checks are off, or as a Worker, where they're off until you set health: { readyz: { enabled: true } }, or health.enabled is false. Use /healthz.

/health or /healthz answers 404

health.enabled is false, which removes every health path, or healthzPath moved the liveness check: /healthz is gone then, and /health answers in its place.

Health checks get 403 {"error":"Forbidden","message":"Invalid Host header"}

The check calls the server by a name that isn't allowed, like 127.0.0.1:3000 or a pod's IP. Add that name to FRONTMCP_ALLOWED_HOSTS, or send the public name in a Host header, as in the Kubernetes probes.

/readyz answers 503, with no details

In production the answer leaves probes out. Set health: { includeDetails: true } to see which probe failed, or check that Redis is reachable.

remote:… is degraded right after startup

The server has a remote app, and FrontMCP hasn't checked it yet. degraded counts as ready, so /readyz answers 200, and the probe turns healthy after the first check if the remote answers. (In 1.8.6 it was unhealthy until the first check and /readyz answered 503 for about 30 seconds.)

Failed to wire /metrics endpoint Error: Cannot find module '@opentelemetry/sdk-trace-base'

metrics.enabled with @frontmcp/observability installed but not its peer. Install @opentelemetry/sdk-trace-base. Without @frontmcp/observability it's Cannot find module '@frontmcp/observability' (more).

The health check passes, but older clients' initialize answers 500

The server is in production without MCP_SESSION_SECRET. Health checks, and clients on MCP 2026-07-28, don't notice. Set the secret.

Prometheus reports the target down with 401 or 403

401: the scrape sends no Authorization header. 403: its token isn't FRONTMCP_METRICS_TOKEN. Check the authorization block and the credentials file.