All posts
AINestJS

Building an MCP Server for Your SaaS with NestJS

"Can we use your product from Claude?" is quickly becoming a standard question in SaaS sales calls, right next to "Do you have an API?". Customers want their AI assistants and agents to search their data, create records and trigger workflows in the tools they already pay for, without copy-pasting between tabs.

The Model Context Protocol (MCP) is how that works. It is an open standard for connecting AI applications to external tools and data, supported by Claude, ChatGPT, Cursor, VS Code and most agent frameworks. Build one MCP server and your product becomes usable from all of them.

This post shows how to add a production-ready MCP server to an existing NestJS backend: tools shaped around real tasks, OAuth with your existing identity provider, every call running with the user's own permissions, and safe handling of actions that change data. The examples use the official TypeScript SDK v2, released alongside the 2026-07-28 version of the specification.

Why an MCP server and not just your REST API

Your REST API was designed for developers who read documentation, write code against it once and handle errors explicitly. An AI agent is a different kind of client: it discovers what it can do at runtime, decides which call to make based on natural-language descriptions and has to recover from mistakes on its own.

MCP gives agents exactly that. A server exposes three kinds of capabilities:

  • Tools: actions the model can call, such as searching orders or issuing a refund. This is where most of the value is.
  • Resources: read-only data the application can load into context, such as a document or a report.
  • Prompts: reusable templates the user can pick, such as "summarise this customer's account".

The client connects, lists what the server offers, and the model uses the names, descriptions and input schemas to decide what to call. Local servers run as a process on the user's machine over stdio. For a SaaS product you want a remote server over HTTP, so customers connect with a URL and log in with their existing account.

How it fits into an existing NestJS app

The MCP server does not need to be a separate service. The simplest and safest setup is a new module inside your existing NestJS application:

  • A McpController serves the /mcp endpoint.
  • The tools call the same services your REST controllers use, so business rules, validation and permission checks live in one place.
  • Your existing identity provider issues the tokens. The MCP endpoint only verifies them, exactly like the rest of your API.

This keeps the MCP layer thin. It is an adapter that translates between the protocol and your domain, not a second backend.

Design tools around tasks, not endpoints

The most common mistake is generating one tool per REST endpoint. A product with 60 endpoints becomes 60 tools, every tool description is sent to the model on every request, and the model picks worse when it has too many similar options.

Instead, start from the jobs users want to get done and design a small set of tools for them, typically five to fifteen:

  • Combine calls the agent would always make together. A get_order tool can return the order with its items, payments and shipments in one response instead of four separate lookups.
  • Name and describe tools for a reader who has never seen your product. The description is the only documentation the model gets: what the tool does, when to use it and what it returns.
  • Keep inputs small and strictly typed. Enums instead of free text, IDs in a known format, sensible limits. Everything the model sends is validated before your code runs.
  • Return compact results. Summaries and IDs first, with a separate tool for details. Large responses cost tokens and push useful context out of the window.

Authentication with your existing identity provider

A remote MCP server should never accept shared API keys pasted into a config file. The specification uses OAuth: your MCP endpoint acts as a resource server, and the client logs the user in through your authorization server and receives an access token issued specifically for the MCP server.

That means you do not build a login flow. If you already use Auth0, Keycloak, Cognito, Clerk or your own OAuth server, it keeps issuing tokens. The MCP server needs to do three things:

  • Reject requests without a valid token with a 401 challenge that tells the client where to authenticate.
  • Publish protected resource metadata at a well-known URL, which points clients to your authorization server.
  • Verify each token, including that it was issued for this server. A token meant for another API must be rejected, and the MCP server must never forward the client's token to other services.

The SDK handles the protocol details. Your part is a verifier that reuses the JWT validation you already have:

import { type AuthInfo, OAuthError, OAuthErrorCode, type OAuthTokenVerifier } from '@modelcontextprotocol/server';
import { Injectable } from '@nestjs/common';
import { JwtService } from '@nestjs/jwt';

export const MCP_URL = new URL(process.env.MCP_URL ?? 'https://api.example.com/mcp');

interface AccessTokenClaims {
  sub: string;
  azp: string;
  scope: string;
  tenant_id: string;
  exp: number;
}

@Injectable()
export class McpTokenVerifier implements OAuthTokenVerifier {
  constructor(private readonly jwt: JwtService) {}

  async verifyAccessToken(token: string): Promise<AuthInfo> {
    const claims = await this.jwt
      // Only accept tokens issued for this MCP server, never tokens meant for another API.
      .verifyAsync<AccessTokenClaims>(token, { audience: MCP_URL.href })
      .catch(() => {
        throw new OAuthError(OAuthErrorCode.InvalidToken, 'Invalid or expired access token');
      });

    return {
      token,
      clientId: claims.azp,
      scopes: claims.scope.split(' '),
      expiresAt: claims.exp,
      extra: { userId: claims.sub, tenantId: claims.tenant_id },
    };
  }
}

The module wires the verifier into the SDK's bearer auth middleware for the MCP route only, so the rest of your API is unaffected:

import { requireBearerAuth } from '@modelcontextprotocol/express';
import { getOAuthProtectedResourceMetadataUrl } from '@modelcontextprotocol/server';
import { type MiddlewareConsumer, Module, type NestModule } from '@nestjs/common';
import { JwtModule } from '@nestjs/jwt';
import { OrdersModule } from '../orders/orders.module';
import { McpController } from './mcp.controller';
import { McpServerFactory } from './mcp-server.factory';
import { MCP_URL, McpTokenVerifier } from './mcp-token.verifier';

@Module({
  imports: [
    JwtModule.register({
      publicKey: process.env.AUTH_PUBLIC_KEY,
      verifyOptions: { algorithms: ['RS256'], issuer: process.env.AUTH_ISSUER },
    }),
    OrdersModule,
  ],
  controllers: [McpController],
  providers: [McpTokenVerifier, McpServerFactory],
})
export class McpModule implements NestModule {
  constructor(private readonly verifier: McpTokenVerifier) {}

  configure(consumer: MiddlewareConsumer) {
    consumer
      .apply(
        requireBearerAuth({
          verifier: this.verifier,
          // Sent in the 401 challenge, so clients can discover where to log in.
          resourceMetadataUrl: getOAuthProtectedResourceMetadataUrl(MCP_URL),
        }),
      )
      .forRoutes(McpController);
  }
}

Finally, publish the metadata document in main.ts. Clients read it to find your authorization server and the scopes you support:

import { mcpAuthMetadataRouter } from '@modelcontextprotocol/express';
import { NestFactory } from '@nestjs/core';
import { AppModule } from './app.module';
import { MCP_URL } from './mcp/mcp-token.verifier';

async function bootstrap() {
  const app = await NestFactory.create(AppModule);

  // Your identity provider stays the authorization server; this API only publishes where to find it.
  const oauthMetadata = await fetch(`${process.env.AUTH_ISSUER}/.well-known/openid-configuration`).then((res) => res.json());
  app.use(mcpAuthMetadataRouter({ oauthMetadata, resourceServerUrl: MCP_URL, scopesSupported: ['orders:read', 'refunds:write'] }));

  await app.listen(process.env.PORT ?? 3000);
}
void bootstrap();

If you already have authentication and authorization in NestJS set up with JWTs and guards, most of this is reuse rather than new work.

Serving MCP from a NestJS controller

SDK v2 makes stateless servers the default: a handler creates a fresh MCP server for every request. There are no sessions to store or share, so the endpoint scales horizontally like any other route, and the server can be built specifically for the authenticated caller.

import { toNodeHandler } from '@modelcontextprotocol/node';
import { createMcpHandler } from '@modelcontextprotocol/server';
import { All, Controller, Req, Res } from '@nestjs/common';
import type { Request, Response } from 'express';
import { McpServerFactory } from './mcp-server.factory';

@Controller('mcp')
export class McpController {
  // A fresh server per request, built for the authenticated caller. No sessions to share between instances.
  private readonly handler = toNodeHandler(createMcpHandler((ctx) => this.servers.create(ctx.authInfo)));

  constructor(private readonly servers: McpServerFactory) {}

  @All()
  handle(@Req() req: Request, @Res() res: Response) {
    return this.handler(req, res, req.body);
  }
}

The bearer auth middleware attaches the verified token to the request, and the SDK passes it to the factory as ctx.authInfo.

Tools that run as the user

The factory is where tools are defined. Because it receives the verified identity, every tool is bound to the current user and tenant before the model sees it. There is no user ID parameter for the model to fill in, so there is nothing to manipulate.

import { type AuthInfo, McpServer, requireScopes } from '@modelcontextprotocol/server';
import { Injectable } from '@nestjs/common';
import { z } from 'zod';
import { OrdersService } from '../orders/orders.service';
import { RefundsService } from '../orders/refunds.service';

const json = (data: unknown) => ({ content: [{ type: 'text' as const, text: JSON.stringify(data) }] });
const fail = (message: string) => ({ content: [{ type: 'text' as const, text: message }], isError: true });

@Injectable()
export class McpServerFactory {
  constructor(
    private readonly orders: OrdersService,
    private readonly refunds: RefundsService,
  ) {}

  create(auth: AuthInfo | undefined): McpServer {
    if (!auth?.extra) throw new Error('MCP request without verified auth');
    // Every tool runs as this user, through the same services and checks as the REST API.
    const user = { userId: String(auth.extra.userId), tenantId: String(auth.extra.tenantId) };

    const server = new McpServer({ name: 'acme-orders', version: '1.0.0' });

    server.registerTool(
      'search_orders',
      {
        title: 'Search orders',
        description:
          'Find orders of the current customer by status and creation date. Returns at most 20 orders, newest first. Use get_order for full details.',
        inputSchema: z.object({
          status: z.enum(['pending', 'paid', 'shipped', 'delivered', 'cancelled']).optional(),
          from: z.iso.date().optional().describe('Earliest creation date, YYYY-MM-DD'),
          to: z.iso.date().optional().describe('Latest creation date, YYYY-MM-DD'),
        }),
        annotations: { readOnlyHint: true },
      },
      async (query) => {
        const orders = await this.orders.search(user, { ...query, limit: 20 });
        // Compact summaries keep responses small; the model asks for details when it needs them.
        return json(orders.map(({ id, number, status, total, currency }) => ({ id, number, status, total, currency })));
      },
    );

    server.registerTool(
      'get_order',
      {
        title: 'Get order',
        description: 'Full details of one order: items, payments, shipments and refunds.',
        inputSchema: z.object({ orderId: z.uuid() }),
        annotations: { readOnlyHint: true },
      },
      async ({ orderId }) => {
        const order = await this.orders.findOne(user, orderId);
        return order ? json(order) : fail(`Order ${orderId} not found. Use search_orders to find the right ID.`);
      },
    );

    // Write tools, see the next section.

    return server;
  }
}

The Zod schema does triple duty: the SDK turns it into the JSON Schema the model sees, validates arguments before your handler runs and infers the TypeScript types of the handler's input.

Actions that change data

Reading data is low risk. Refunds, cancellations and emails to customers are not, and they deserve extra layers:

  • Scopes: write tools require a separate OAuth scope. A user can connect an assistant with read-only access, and a token without refunds:write gets an insufficient scope error, which prompts the client to ask for more access.
  • Two steps: a prepare_refund tool checks eligibility and returns the exact amount, while confirm_refund executes it. The model has to show the user concrete details before anything happens, and the confirmation ID doubles as an idempotency key.
  • Annotations: readOnlyHint and destructiveHint tell clients which tools are safe to run freely and which should require the user's approval. They are hints for the client, not a security boundary. Enforcement stays on the server.
  • Server-side limits: maximum refund amounts, rate limits per user and an audit log of every call. The same rules your REST API enforces, because the tools call the same services.
server.registerTool(
  'prepare_refund',
  {
    title: 'Prepare refund',
    description:
      'Checks whether an order can be refunded and returns the exact amount and a confirmationId. Does not move money. Show the details to the user before calling confirm_refund.',
    inputSchema: z.object({
      orderId: z.uuid(),
      amount: z.number().positive().optional().describe('Partial amount; omit to refund the full order'),
      reason: z.string().min(3).max(500),
    }),
    annotations: { readOnlyHint: true },
    scopeChallenge: requireScopes('refunds:write'),
  },
  async (input) => json(await this.refunds.preview(user, input)),
);

server.registerTool(
  'confirm_refund',
  {
    title: 'Confirm refund',
    description: 'Executes a refund prepared by prepare_refund. Only call after the user has explicitly agreed.',
    inputSchema: z.object({ confirmationId: z.string() }),
    annotations: { destructiveHint: true, idempotentHint: true },
    scopeChallenge: requireScopes('refunds:write'),
  },
  async ({ confirmationId }) => json(await this.refunds.execute(user, confirmationId)),
);

For the most sensitive actions, MCP also supports elicitation: the server asks the user a question directly through the client, without relying on the model to relay it. Where clients support it, it is the strongest way to get explicit confirmation.

Errors the model can recover from

An agent cannot read a stack trace or open your documentation. When something goes wrong, the error message is its only guidance, so write errors for the model:

  • Return tool errors as results with isError: true and a clear message instead of throwing. The model reads it and tries again.
  • Say what to do next: "Order not found. Use search_orders to find the right ID." works far better than "404".
  • Let the SDK reject invalid arguments. It already returns a descriptive validation error, such as an invalid UUID, before your handler runs.
  • Keep internal details out of messages. The model's output may be shown to the user or logged by the client.

Rate limits, cost and abuse

An agent can call tools far faster than a human clicks buttons, and a confused one can loop. Protect the MCP endpoint like a public API:

  • Rate limit per user and per OAuth client, not just per IP, for example with @nestjs/throttler.
  • Cap result sizes and paginate. Twenty results with a cursor beats a thousand rows in one response.
  • Log every tool call with the user, client, arguments and duration. When a customer asks what their assistant did, you can answer.
  • Monitor tool error rates. A tool that fails often usually has a confusing description or schema.

Testing

Test the server at three levels:

  • Interactively with the MCP Inspector : run npx @modelcontextprotocol/inspector, connect to your local /mcp URL and call each tool by hand.
  • Integration tests with the SDK's client package: sign a test token, connect, list tools and call them against a test database. Cover the auth cases too: missing token, wrong audience and missing scope.
  • With real models: connect Claude or another client, try the tasks your customers care about and watch which tools the model picks. Unclear descriptions show up quickly. For a structured approach, see what actually works when adding AI agents to production apps .

Deploying

Because the MCP server is part of your existing application, there is nothing new to deploy. It ships in the same Docker image, behind the same HTTPS reverse proxy, with the same health checks and logging. The stateless design means you can run as many instances as you like without sticky sessions. If you self-host, a platform like Coolify handles it like any other service.

Two production details are easy to miss: set MCP_URL to the exact public URL clients use, because it must match the token audience and the metadata document, and make sure your reverse proxy does not buffer streamed responses on the /mcp route.

Checklist

  • A small set of task-shaped tools with clear names, descriptions and strict schemas.
  • The MCP module inside your existing app, calling the same services as the REST API.
  • OAuth through your existing identity provider, with protected resource metadata published.
  • Tokens verified for this server's audience and never forwarded to other services.
  • Tools bound to the authenticated user and tenant, with no user ID in the inputs.
  • Separate scopes, a prepare and confirm flow, and server-side limits for write actions.
  • Errors returned as helpful tool results.
  • Rate limits, result caps and an audit log of every call.
  • Tests with the Inspector, integration tests for auth and tools, and trials with real clients.

Conclusion

An MCP server is quickly becoming as expected as a REST API, and for a NestJS product it is a thin layer on top of what you already have: your services, your identity provider and your deployment. The work that matters is in the details: tools designed for how agents think, permissions that follow the user and write actions that cannot surprise anyone.

If you want your product to be usable from Claude, ChatGPT and other AI agents, and would like help designing or building the MCP server, get in touch . I help teams add AI integrations to the backends they already run.