A practical guide for installing MCPlama, adding MCP servers, inviting teammates, and connecting AI clients through one governed gateway.
MCPlama is a self-hosted gateway that sits between your AI clients — Claude Desktop, Cursor, VS Code Copilot, or anything else that speaks MCP — and the actual MCP servers you want them to use. Instead of every teammate juggling their own API keys and configs for every tool, you install servers once in MCPlama, and everyone connects through it.
It gives you one place to see who's connected to what, control which tools people can use, and keep a record of every call — while your AI client only ever sees a single connection link, never the underlying credentials.
Every MCP call is authenticated, checked against policy, logged, and forwarded through MCPlama to the selected MCP server.
These are recommended starting specifications with additional headroom for the database, audit logs, and MCP server containers. Monitor your workload and scale up as concurrency and log volume grow.
| Tier | Concurrent users | Servers | Audit logs | CPU | RAM | Disk |
|---|---|---|---|---|---|---|
| Small | ~10 | up to 20 | < 100 MB | 2 vCPU | 2 GB | 10 GB |
| Team | ~100 | 100 | 1 GB | 4 vCPU | 4 GB | 20 GB |
| Growing | ~250 | up to 500 | 5 GB | 4–8 vCPU | 8 GB | 30 GB |
| Large | 500+ | 500+ | 10 GB+ | 8 vCPU | 16 GB | 50 GB+ |
MCPlama ships as a single Docker image — backend, database, and dashboard, all bundled. No separate install steps, and nothing to generate yourself.
The community image is published in the MCPlama registry as mcplama/mcplama:latest. Use a private mirror only when your deployment requires one.
docker pull mcplama/mcplama:latest
The docker run below pulls it automatically if it isn't present locally — running the pull yourself first just makes the first run faster to watch.
MCPlama creates and manages its Docker network automatically. Mount the Docker socket so MCPlama can start MCP server containers, and keep the named volume so your database and generated keys survive restarts and upgrades. Choose the example that matches how users and MCP clients will reach the gateway.
Use this when MCPlama is accessed from the same host:
docker run -d --name mcplama \ -p 127.0.0.1:8080:8080 \ -v /var/run/docker.sock:/var/run/docker.sock \ -v mcplama_pgdata:/var/lib/postgresql \ -e GATEWAY_URL=http://localhost:8080 \ mcplama/mcplama:latest
Use this when a reverse proxy, load balancer, or ingress provides the public HTTPS endpoint:
docker run -d --name mcplama \
-p 127.0.0.1:8080:8080 \
-v /var/run/docker.sock:/var/run/docker.sock \
-v mcplama_pgdata:/var/lib/postgresql \
-e GATEWAY_URL=https://[yourdomain].com \
mcplama/mcplama:latest
In the public example, the proxy must forward the public HTTPS URL to http://127.0.0.1:8080. Do not change GATEWAY_URL to the internal address.
Use this only when you intentionally want MCPlama reachable directly without a reverse proxy. The host firewall must allow the selected port:
docker run -d --name mcplama \
-p 8080:8080 \
-v /var/run/docker.sock:/var/run/docker.sock \
-v mcplama_pgdata:/var/lib/postgresql \
-e GATEWAY_URL=http://[yourdomain].com:8080 \
mcplama/mcplama:latest
To use another host port, for example 8043, map it to the container's port 8080 and include it in GATEWAY_URL: -p 8043:8080 with http://[yourdomain].com:8043.
GATEWAY_URL is important. This example uses http://localhost:8080 for local access. Replace it with the exact address users and MCP clients will use, such as http://[yourdomain].com:8080 or https://[yourdomain].com. Include a custom port when the public URL uses one.For direct access or public exposure through HTTPS, choose the appropriate option in Public access through a proxy.
mcplama_pgdata volume on first boot. Back up this volume together with the database; losing it can make saved credentials and OAuth connections unreadable.For this local example, visit http://localhost:8080. The setup wizard walks you through creating the admin account, naming your gateway, and optionally turning on email.
For a local test, the defaults are enough. For a public deployment, set GATEWAY_URL to the URL users and MCP clients actually use. The other variables are optional overrides.
| Variable | Required? | Default | Purpose |
|---|---|---|---|
| GATEWAY_URL | Required for public use Optional locally | http://localhost:8080 | The address your AI clients and browser actually use to reach this instance |
| ALLOWED_PRIVATE_WEBHOOK_HOSTS | Optional | empty | Comma-separated exact hostnames or IPs that webhook policies may call on a private network |
| BROKER_MAX_CPU_LIMIT | Optional | 4 | Configured maximum CPUs for managed MCP containers. The broker also caps this automatically to the CPUs available on the host |
| BROKER_MEM_LIMIT | Optional | 1g | Configured maximum memory for managed MCP containers; the broker also caps it to memory available on the host |
| RUNNER_IMAGE | Optional | mcplama/runner:latest | Runner image used for npx, uvx, and stdio-based MCP servers; useful for mirrors or pinned tags |
The free community edition includes up to 10 users and 100 installed servers.
Advanced/private installs can also set RUNNER_IMAGE to point npx, uvx, and stdio-based servers at a registry mirror or pinned runner tag. Most installs should leave it unset.
Per-server CPU and memory limits are configured in the Add server or Edit server form. For managed Docker, npx, uvx, and stdio servers, use values such as CPU 1 and memory 256m. The broker clamps these values to both the configured limits and the resources available on the host, including after a VM resize.
Private webhook destinations are blocked by default for SSRF protection. For a trusted LAN-only webhook, add its exact host or IP, for example -e ALLOWED_PRIVATE_WEBHOOK_HOSTS=192.168.1.42. This applies only to policy webhooks; private MCP server URLs remain blocked.
For a public deployment, put an HTTPS reverse proxy, load balancer, or ingress in front of MCPlama. The proxy owns the public TLS endpoint; MCPlama listens on a private host port.
https://[yourdomain].com
↓
HTTPS proxy :443
↓
127.0.0.1:8080
↓
MCPlama container
For a proxy on the same host, bind MCPlama to localhost:
-p 127.0.0.1:8080:8080
For intentional direct access, bind all host interfaces and choose the host port users will call:
-p 8080:8080 # or -p 8043:8080
The first port is the host port; the second is MCPlama's container port.
GATEWAY_URL must exactly match the address users and MCP clients use. Its default is http://localhost:8080. For a proxy, use the public URL such as https://[yourdomain].com. For direct access, include the host port, such as http://[yourdomain].com:8080 or http://[yourdomain].com:8043.
Terminate TLS at the proxy and forward to the private MCPlama address. Preserve the original Host, X-Forwarded-For, and X-Forwarded-Proto headers. MCP connections use streaming, so disable response buffering and allow long-lived read and send timeouts.
Use a trusted certificate for the public hostname. Allow only the proxy's public ports through the firewall, and keep the MCPlama host port private when using a proxy. FRONTEND_URL and CORS_ORIGINS are optional overrides for a separate frontend origin; they are not required when the dashboard and gateway share one public origin.
Almost always the host/network firewall, or GATEWAY_URL still pointing at localhost — see Public access through a proxy above.
Stopping or restarting the container does not normally remove data. The -v mcplama_pgdata:/var/lib/postgresql volume is what makes your data persistent when the container is recreated or upgraded. Without it, everything lives inside the container's writable layer and is gone when the container is removed.
MCPlama stores its generated encryption keys in the same mcplama_pgdata volume as the database. Restore the database and volume together; restoring only one side can leave encrypted credentials unreadable.
If the host running MCPlama can't reach the registry that contains the server image, installation will fail. Make sure the host is logged in to the registry, or mirror the image somewhere it can reach.
Those server types use the MCPlama runner image. MCPlama pulls it automatically when needed, but the host still needs registry access. In restricted networks, mirror the runner image and set RUNNER_IMAGE to that mirrored image.
The broker automatically caps CPU and memory requests to the resources available on the host. If Docker still reports a resource error, edit the server and lower the matching limit (practical starting points are CPU 1 and memory 256m). Use Docker memory formats such as 256m or 1g, not a bare number or 6mb.
MCPlama includes a built-in catalog page for ready-to-install MCP servers. Use the search, category filters, and details drawer to find the server you need, then install it from the dashboard. Don't see what you need? Add a custom server by hand — see below.
Custom servers let you put internal tools, private packages, and third-party servers behind the same MCPlama controls as the catalog entries. Choose the runtime, add required environment variables or user credential fields, then publish it for your team.
| Type | What it means |
|---|---|
| Remote URL | Proxy to an HTTP MCP server you already run somewhere else — MCPlama just fronts it. |
| Docker image | Point MCPlama at any Docker MCP image; it pulls it, starts the container, and manages its lifecycle for you. |
| NPX package | Run an npm-published MCP package directly — no separate install or build step on your end. |
| UVX package | Same idea for Python — run a uv-published MCP package with no local Python setup required. |
For Docker, NPX, and UVX servers, the Add server and Edit server forms also let you set optional CPU and memory limits. If a server fails to start because Docker rejects its resources, edit the server and lower those values. Whichever type you pick, it ends up behind the same /connect/{token} link as everything else — your AI client doesn't need to know or care how a server actually runs.
When you install a server, its Authorization tab asks you to pick one of two credential modes before anyone can connect through it:
| Mode | How it works | Use it for |
|---|---|---|
| Shared | An admin authorizes once (an API key, or an OAuth login). Every connection to that server reuses that same authorization. | Team-wide tools where identity doesn't matter — a shared Slack bot, a shared search API key. |
| Per-user | Each user completes their own authorization before their connection works. Credentials are kept isolated per user. | Anything scoped to an individual account — GitHub, Notion, email — where one person shouldn't see another's data. |
For servers that support OAuth, MCPlama discovers the authorization server automatically and runs the full login flow — users click "Authorize" and sign in with the real provider directly. MCPlama never sees their password, and your AI client only ever holds a connection link you can revoke in one click.
From Users in the dashboard, an admin can invite someone by email and pick their role — admin (full dashboard access) or member (Member Portal only, see below). Each invite is a link that's valid for 7 days.
If email (SMTP) is configured, MCPlama can send the invite for you. If not, just copy the invite link and send it yourself — either way works.
Regular users never touch the admin dashboard. They get their own Member Portal instead — a simple, self-service way to get connected:
Policies let admins control exactly what a connection is allowed to do, beyond just "approved or not." You can scope a policy to a specific server, user, or role.
| Policy type | What it does |
|---|---|
| Tool block | Deny specific tools on a server (e.g. block a "delete" tool while allowing everything else). |
| Tool allow | The inverse — only the listed tools are permitted, everything else is denied. |
| Rate limit | Cap how many calls a user or role can make in a given time window. |
| Time restrict | Only allow calls during set hours or days. |
| Webhook | Ask a trusted HTTP endpoint to allow, deny, or rewrite a tool call before it reaches the MCP server. |
MCPlama sends the webhook a JSON payload before each tool call. The endpoint must return a JSON object with allowed set to true or false. A false response blocks the call and can include a user-facing reason.
Request
{
"server_id": 7,
"user_id": 42,
"tool_name": "delete_file",
"request": { "jsonrpc": "2.0", "method": "tools/call", "params": { "name": "delete_file" } }
}
Allow: { "allowed": true }
Deny: { "allowed": false, "reason": "Destructive tools require approval" }
You can also return mutated_request with a rewritten JSON-RPC request. Webhook failures, timeouts, non-2xx responses, and malformed responses fail closed and block the call. The policy’s general “Action when violated” setting does not apply to webhook policies—the webhook response is the decision.
ALLOWED_PRIVATE_WEBHOOK_HOSTS.A blocked call comes back to the AI client as a normal failed tool call, not a broken connection — so the assistant can explain it rather than just hanging.
Optional, but recommended. Configuring an outbound mail server lets MCPlama send email for:
You can turn it on during the setup wizard, or any time afterward from Settings → Email / SMTP. All it needs is a host, port, and credentials for any standard SMTP provider (Gmail, SES, Postmark, your own mail server, etc.).
Use this short checklist before putting a deployment into regular use. The detailed URL and HTTPS guidance is in Public access through a proxy; variable meanings and defaults are in Environment variables.
| Public URL | Set GATEWAY_URL to the exact address clients use. |
| Transport security | Use HTTPS with a trusted certificate for public access. |
| Persistence | Keep and back up mcplama_pgdata; it stores the database and generated encryption keys. |
| Docker access | Treat access to the Docker socket as host-level administrative access. |
tool_block or tool_allow rule is cheap insurance against an AI client calling something it shouldn't.