ai-agents #mcp#model context protocol#ai agents#python#typescript#production

Building Production MCP Servers: Python and TypeScript Guide

S

S L Manikanta

Aug 24, 2026 • 7 min read

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

When building AI tools and agents, writing custom glue code for every model provider quickly turns into a mess. Anthropic introduced the Model Context Protocol (MCP) to fix this by giving AI clients a standard way to discover tools, fetch data, and read system prompts.

Cursor, Claude Desktop, Claude Code, and custom agent frameworks now speak MCP natively.

If you want your internal Postgres database, Redis cache, or REST APIs accessible to any AI coding tool without rewriting custom integrations, building an MCP server is the cleanest approach.

Here is how MCP works under the hood, how to build production servers in both Python and TypeScript, and how to handle security, auth, and deployment.

flowchart LR
    subgraph Clients [MCP Clients]
        Claude[Claude Code / Desktop]
        Cursor[Cursor IDE]
        CustomAgent[Custom LangGraph / PydanticAI Agent]
    end

    subgraph Transport [Transport Layer]
        Stdio[Standard I/O: Local Process]
        SSE[Server Sent Events: Remote HTTP]
    end

    subgraph MCPServer [Production MCP Server]
        Router[MCP Protocol Router]
        Tools[Exposed Tools]
        Resources[Read Only Resources]
        Prompts[Prompt Templates]
    end

    subgraph Internal [Internal Systems]
        DB[(PostgreSQL Database)]
        API[Internal REST / GraphQL APIs]
    end

    Clients --> Transport
    Transport --> Router
    Router --> Tools
    Router --> Resources
    Router --> Prompts
    Tools --> DB
    Tools --> API
    Resources --> DB

How MCP Works: Transports and Primitives

MCP uses JSON-RPC 2.0 messages. It defines three core capabilities that your server can expose:

  1. Tools: Functions that the LLM can call with arguments to perform actions or mutate data (e.g. running a SQL query or creating a ticket).
  2. Resources: Read only data streams that the client can attach to context, like file contents, database schemas, or logs.
  3. Prompts: Pre-built prompt templates and workflows that users can trigger inside their AI interface.

Choosing Your Transport: Stdio vs SSE

  • Standard I/O (stdio): The client (e.g. Cursor or Claude Desktop) launches your server as a local child process and talks to it over standard input and standard output. This is best for local developer tools, command line utilities, and private scripts running on your machine.
  • Server Sent Events (SSE): The server runs as an HTTP service. The client opens an SSE connection to receive streaming messages from the server and sends HTTP POST requests to send commands. This is what you need for shared team servers, cloud deployments, and Docker containers.

Building an MCP Server in Python with FastMCP

The official mcp Python SDK includes a high level helper called FastMCP (inspired by FastAPI) that handles schema generation from type hints automatically.

1. Installation

pip install "mcp[cli]" asyncpg pydantic

2. Implementation: Database Tool Server

Here is a complete, production ready MCP server that connects to PostgreSQL and exposes a safe query tool and a schema resource:

import os
from typing import Any, Dict, List
import asyncpg
from mcp.server.fastmcp import FastMCP
from pydantic import BaseModel, Field

# Initialize FastMCP server
mcp = FastMCP("postgres-analytics-server")

DATABASE_URL = os.getenv("DATABASE_URL", "postgresql://user:password@localhost:5432/analytics")
pool: asyncpg.Pool | None = None

@mcp.resource("schema://analytics")
async def get_database_schema() -> str:
    """Returns the DDL table schema for the analytics database."""
    async with pool.acquire() as conn:
        rows = await conn.fetch("""
            SELECT table_name, column_name, data_type 
            FROM information_schema.columns 
            WHERE table_schema = 'public'
            ORDER BY table_name, ordinal_position;
        """)
        
        schema_text = "Database Schema:\n"
        current_table = ""
        for row in rows:
            if row["table_name"] != current_table:
                current_table = row["table_name"]
                schema_text += f"\nTable: {current_table}\n"
            schema_text += f"  - {row['column_name']}: {row['data_type']}\n"
            
        return schema_text

class QueryInput(BaseModel):
    query: str = Field(description="Read-only SELECT SQL query to execute")
    max_rows: int = Field(default=50, ge=1, le=500, description="Maximum number of rows to return")

@mcp.tool()
async def run_analytics_query(params: QueryInput) -> List[Dict[str, Any]]:
    """Execute a read-only SQL query against the analytics database."""
    # Basic safety check to reject write operations
    clean_query = params.query.strip().lower()
    if not clean_query.startswith("select") and not clean_query.startswith("with"):
        raise ValueError("Only SELECT and WITH queries are permitted on this endpoint.")
        
    async with pool.acquire() as conn:
        # Enforce query timeout to prevent runaway table scans
        async with conn.transaction(readonly=True):
            await conn.execute("SET LOCAL statement_timeout = '5000ms'")
            records = await conn.fetch(params.query)
            
            # Slice results to prevent context window explosion
            limited_records = records[:params.max_rows]
            return [dict(record) for record in limited_records]

if __name__ == "__main__":
    import asyncio
    
    async def main():
        global pool
        pool = await asyncpg.create_pool(DATABASE_URL, min_size=2, max_size=10)
        try:
            # Runs using standard I/O for local tools like Claude or Cursor
            await mcp.run_stdio_async()
        finally:
            await pool.close()
            
    asyncio.run(main())

Building an MCP Server in TypeScript

If your team runs Node.js or Bun in production, the official TypeScript SDK (@modelcontextprotocol/sdk) provides full type safety.

1. Installation

npm install @modelcontextprotocol/sdk zod
npm install -D typescript @types/node

2. Implementation: TypeScript Server

import { Server } from "@modelcontextprotocol/sdk/server/index.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import {
  CallToolRequestSchema,
  ListToolsRequestSchema,
} from "@modelcontextprotocol/sdk/types.js";
import { z } from "zod";

const server = new Server(
  {
    name: "github-issue-manager",
    version: "1.0.0",
  },
  {
    capabilities: {
      tools: {},
    },
  }
);

// Define tool argument schema
const CreateIssueSchema = z.object({
  repo: z.string().describe("Repository name in owner/repo format"),
  title: z.string().min(3).describe("Issue title"),
  body: z.string().describe("Detailed description of the issue"),
  labels: z.array(z.string()).optional().describe("Issue labels"),
});

// List available tools
server.setRequestHandler(ListToolsRequestSchema, async () => {
  return {
    tools: [
      {
        name: "create_github_issue",
        description: "Creates a new issue in a GitHub repository",
        inputSchema: {
          type: "object",
          properties: {
            repo: { type: "string", description: "Repository in owner/repo format" },
            title: { type: "string", description: "Issue title" },
            body: { type: "string", description: "Issue description" },
            labels: { type: "array", items: { type: "string" }, description: "Labels" },
          },
          required: ["repo", "title", "body"],
        },
      },
    ],
  };
});

// Handle tool execution
server.setRequestHandler(CallToolRequestSchema, async (request) => {
  if (request.params.name === "create_github_issue") {
    const args = CreateIssueSchema.parse(request.params.arguments);
    
    // Simulate GitHub API call
    const issueUrl = `https://github.com/${args.repo}/issues/101`;
    
    return {
      content: [
        {
          type: "text",
          text: `Issue created successfully: ${issueUrl}\nTitle: ${args.title}`,
        },
      ],
    };
  }
  
  throw new Error(`Tool not found: ${request.params.name}`);
});

async function run() {
  const transport = new StdioServerTransport();
  await server.connect(transport);
}

run().catch((err) => {
  console.error("Fatal error running MCP server:", err);
  process.exit(1);
});

Connecting Your MCP Server to Claude Desktop and Cursor

To use your local server inside Claude Desktop or Cursor, add it to your configuration file.

For Claude Desktop (claude_desktop_config.json):

On macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
On Windows: %APPDATA%\Claude\claude_desktop_config.json

{
  "mcpServers": {
    "postgres-analytics": {
      "command": "python",
      "args": ["-m", "src.mcp_server"],
      "env": {
        "DATABASE_URL": "postgresql://user:secret@localhost:5432/analytics"
      }
    },
    "github-tools": {
      "command": "node",
      "args": ["/absolute/path/to/dist/index.js"],
      "env": {
        "GITHUB_TOKEN": "ghp_xxxxxxxxxxxx"
      }
    }
  }
}

For Cursor IDE:

In Cursor, go to Settings > Features > MCP Servers > Add New MCP Server, select command, and enter the execution command.


Production Best Practices and Security

When taking MCP servers from local development to production setups, keep these rules in mind:

1. Never Trust LLM Generated SQL Directly

Always enforce read only transactions (SET LOCAL default_transaction_read_only = on) and use dedicated database users with GRANT SELECT only. Never give an MCP tool raw DROP or DELETE permissions without human confirmation.

2. Enforce Strict Output Limits

LLMs have finite context windows. If an SQL query returns 10,000 rows, serializing that into JSON will exhaust the model’s memory and cost significant token fees. Always hard cap query results at 50 to 100 rows.

3. Add Request Timeouts

Database queries and third party APIs can hang. Set explicit statement timeouts (e.g. 5 seconds) on every network call so your MCP server never blocks the agent indefinitely.

✉ Newsletter

Want to build production-ready AI?

Subscribe to StackMindset to receive actionable systems engineering checklists and code walkthroughs. No spam, only technical insights.

S

Written by S L Manikanta

AI Engineer specializing in agentic workflows, multi-step LLM validation pipelines, and secure cloud environments. Sharing practical lessons from building software.

Related Articles

ai-agents
The Shift to Agentic AI Workflows in Production

Why engineering teams are moving away from simple copilots to autonomous agentic workflows, and the technical challenges of long-running state management.

ai-agents
AI Agent Memory: Short-Term vs Long-Term Memory

A complete architectural breakdown of how AI agents manage state, covering short-term conversational context and long-term persistent memory systems.

ai-agents
AI Agent Observability: Logs, Traces, and Metrics in Production

A complete technical reference and implementation guide to observing agentic workflows, tracking LLM token costs, logging reasoning trajectories, tracing nested tool calls, and monitoring system metrics in production.