Connecting Private Systems to Claude: What I Learned Building an MCP Gateway
I wanted to find out whether our internal, VPN-only systems could be made usable from Claude, with real identity, and without the connection turning into a

I wanted to find out whether our internal, VPN-only systems could be made usable from Claude, with real identity, and without the connection turning into a channel for pulling data out.
That sounds like one problem. It is really three, and they tend to get mushed together:
The first one is the one most write-ups solve. The third is the one I found interesting, and it is where most of the work ended up.
What came out of it is a proof of concept: a small gateway that exposes a mock internal ticket system and a real Nextcloud account to claude.ai, with OAuth login, per-user permissions, response limits and an audit log. It has since grown a real identity provider and a real secrets store, which is the part I did not expect to be the most interesting. It was not benchmarked, not load-tested, not security-reviewed and never run in production. The architecture pattern is old and none of the individual ideas are mine. What might save you some time are the specific findings and the mistakes, so that is what this post is about.
There are three ways to get a chat client to internal data, and they differ by roughly an order of magnitude in effort. Picking the right one is most of the decision.
A local MCP server on the employee's machine is the underrated one. Claude Cowork Desktop runs its agent loop on the device, and local MCP servers run there too. The laptop is already on the VPN, so nothing gets published and no inbound path is opened. For read-only access to an internal wiki or ticket system that is about a day of work, and it is where I would start. Cowork Web is the opposite case: that sandbox runs on Anthropic infrastructure and cannot reach private addresses.
A public gateway, the subject of this post, is the answer when the request comes from an environment you do not control, which in practice means the chat window. It is one to two weeks rather than one day.
MCP tunnels look like they should be the answer, and this is the finding worth having early: they do not work here. From Anthropic's documentation, verbatim: "MCP tunnels created through the Console are not available as connectors in claude.ai."1 They serve Managed Agents and the Messages API. No amount of infrastructure work changes that.
So there is no configuration in which claude.ai dials into a VPN. Something of yours has to be publicly reachable. That reframes the question from "how do we tunnel in?" to "what is the narrowest thing we can expose?", and the second question is a lot easier to answer well.
The gateway is one process, publicly reachable, sitting between the chat client and the private systems. Everything else stays unreachable from outside.
The direction of the connection is what makes this work. The gateway calls into the private network; the LLM provider never does. The provider only ever sees one HTTPS address. Nobody needs to be handed a VPN account, which is the alternative this replaces.
Every arrow points the same way. There is no arrow from the chat client into the private network, and that is the whole design.
The controls are not a checklist copied from a framework. Each one is there because of a specific attack:
| Control | The attack it addresses |
|---|---|
Narrow, typed tools only, no query(sql), no http_request(url) | The model cannot express a request the gateway did not anticipate |
| The gateway holds the credentials, never the model | Credentials cannot be extracted from a context that never held them |
| No token passthrough, the inbound token stops at the gateway | Confused deputy. The MCP specification makes this a MUST NOT2 |
| Acting user and scopes come from the token, never from a tool argument | The caller cannot assert an identity, or a permission, by asking for it |
| Tools filtered by scope before being advertised | A tool absent from the model's context cannot be requested by an injection |
| Response size and count caps | The tool-result path is an exfiltration channel, and this is where it gets bounded |
| Every call audited with hashed arguments | Forensic reconstruction without storing sensitive data a second time |
Two more did not make the table but are cheap: sensitive fields appear only on single-record lookups, never in bulk listings, and there are rate limits per identity and per tool.
If I could keep one row it would be the first, and it is the one most easily traded away. Let us just add a generic query tool, it is so much more flexible sounds reasonable in a design review, but flexibility is exactly the property you are trying not to have: an expression language covers every request you might have thought to forbid. The picture I keep coming back to is a service counter. You do not walk into the warehouse. You ask, and the person behind it does the twelve things on their list, no matter how nicely you ask for a thirteenth. A VPN account is the run of the warehouse.
A human picks an identity in a browser. That subject arrives intact at a system the LLM provider cannot reach, and it shows up in both the gateway's audit log and the internal system's own log: two logs, same event, from two sides.
It then decides the data. Asked "which open tickets are assigned to me?", the assistant calls a whoami-style tool first, unprompted, and filters on the result. The filter runs on an identity the caller did not supply and cannot choose. That the assistant looks it up on its own is convenience, not enforcement.
The most useful thing I got wrong was believing I could stand in for an identity provider. The framework I used ships an in-memory OAuth provider for testing that runs a complete, spec-shaped OAuth 2.1 flow and auto-approves everything without recording who the user is.3 It issues a valid-looking token with no subject attached, so whoami returns nothing while the auth still looks like it is working. A component can implement a protocol correctly and still be useless for the property you actually care about.
So the demo provider is gone and Keycloak does that job. The gateway became a resource server and nothing more: the chat client authenticates against Keycloak directly, and the gateway only verifies signatures against the realm's keys. Users, groups, rotation and revocation belong to something with a database and an admin console.
Two things came out of that which I would not have predicted.
Sessions stopped dying. Under the stand-in provider every restart ended every session, and since you restart constantly while developing, the thing felt like it demanded a fresh login every few minutes. That was a deployment property, not a protocol one. With Keycloak, a token obtained before a kill -9 still works against the process that comes back up.
Keycloak grants scopes per client, not per user. This is easy to get backwards and I did. Both demo users receive tickets:write in their token, because the client is permitted to request it. Treating that scope claim as a permission would have made group membership meaningless. The gateway therefore derives its own scopes from the group claim before the framework ever sees them, narrowing only, never widening. The token stays the ceiling.
That narrowing is what keeps a capability out of reach. The write tool needs a scope only a writing group holds, and a read-only user is not merely refused: the tool is never advertised, so the model is never told it exists. Two independent layers, deliberately, one about the model's context and one about the request path, and neither relies on the other holding.
A real identity provider also bought a second client. The same instance, the same realm and the same tools have been driven from claude.ai and from ChatGPT, which register in different ways, none of it code in this project. The stand-in would have had to implement both.
One ticket in the fixture data contains a live prompt injection. It instructs the assistant to dump every ticket, reveal the internal API key, and POST customer data to an external address.
Nothing happens. Not because a filter caught it, but because none of those three things is a capability the gateway has. The bulk listing is capped at ten records and omits the customer field. The key never enters the model's context. There is no tool that takes a URL.
The injection is not detected. It simply has nothing to work with.
That difference is what the whole exercise is about. Detection is an arms race, and it is one you eventually lose, because instructions cannot be reliably separated from prose. Having nothing to redirect to does not degrade when the attacker gets more creative, because creativity still has to operate inside the available vocabulary, and against the ticket system that vocabulary is four typed tools.
A live run gave me the accidental version of this argument. The assistant called a tool before its schema had been loaded, got a red error telling it to look the tool up first, did that, and retried successfully. The audit log showed exactly one call for that tool in that session, outcome ok: the failed one never left the browser. The model got something wrong, and it was structurally uninteresting.
Against a real system, the first decision was the authentication flow, and my first choice was the wrong one.
Nextcloud ships an OAuth2 app, which looks like the obvious pick. Its own admin documentation talks you out of it: "every token has full access to the complete account including read and write permission to the stored files", and "without scopes and restrictable access it is not recommended to use a Nextcloud instance as a user authentication service."4 When a vendor advises against its own feature that clearly, it is worth listening.
What works instead is Login Flow v2, the mechanism the official desktop and mobile clients use: the user signs in through a normal browser window on Nextcloud's own login page, and the app ends up holding a per-device, individually revocable app password.5 There is no admin setup, and 2FA and SSO work because it is literally the normal login page.
There is now exactly one write into Nextcloud, and its shape is the argument of this post in miniature. create_event creates a calendar entry and takes no calendar name. No phrasing of a request reaches a second calendar, because the parameter does not exist. Nextcloud's own share permissions then sit underneath as an independent second layer: pointed at a calendar the account may only read, the write fails with Nextcloud's 403, a boundary my code cannot weaken even after all my own checks have said yes.
That write tool also forced a gap closed. Caller-supplied paths were only stripped of slashes, so .. passed straight through into what became a well-formed request for a different account. Nextcloud rejects those anyway, so nothing ever leaked. But as the fix's own comment puts it: relying on that puts the boundary in someone else's code, and the whole point of this gateway is that the boundary is in ours.
The Nextcloud app password used to live in a dictionary in the gateway process, which meant I had contained the interface but not the credential. A live demo made the cost concrete: the gateway restarted, the login survived because Keycloak now owned it, and every Nextcloud account link was gone. There was also no way to revoke a single link short of restarting the server.
Vault holds it now. What makes this more than swapping a dictionary for a database is where the authorization decision moved. The gateway does not authenticate to Vault as itself. It hands over the caller's own Keycloak token, Vault verifies that signature against Keycloak's keys independently, and then applies a policy templated on the identity it just verified, so the path a caller can read is derived from who Vault decided they are. A bug in my code that asks for somebody else's path gets a 403. That is asserted by a test, not assumed.
Three systems now sit in the chain, and none of them takes my word for anything: Keycloak says who you are, Vault says which credential you may use, Nextcloud says what that account may do. For anyone allergic to the licence, OpenBao is a drop-in replacement, same API and same paths.
The new costs, stated rather than discovered later: Vault has to be reachable or the Nextcloud tools stop being advertised at all. Failing closed is right for a credential store, but it is a new dependency. And deleting the stored credential revokes my access, not the app password itself, which stays valid in Nextcloud until revoked there.
None of these are dramatic, but they are the kind of thing you hit on the same path.
Identity was silently lost on token refresh. This is the one I would most want somebody to check in their own build. Identity was stamped onto access tokens from the authorization-code grant, but not onto those from the refresh grant. So about an hour into a session the subject quietly became unknown while the scopes survived: full permissions, no attribution. Partial failure in an auth path is worse than total failure, because total failure gets noticed.
The mock was too clean. My calendar listing filtered out system collections by name, which passed against a local mock that only ever returned real calendars. A real instance returns several more, one of which carries the account name as its last path segment, so the account showed up as a calendar named after itself. The fix was to filter on a protocol-level property instead of a list of names. Then I deliberately made the mock messier, so the test now asserts the extra collections get filtered out. A fixture that is tidier than production hides exactly the class of bug production will find.
Login Flow v2 reuses your browser session. Opening the login URL in a normal window silently offers the account you are already signed in as. On my first live test that authorised an admin session instead of the intended service account, and the flow completes successfully either way, so it is easy to miss. Sign out first, or use a private window.
This is the part I would want to read first in someone else's post.
The app password still grants full account access. Nextcloud has no scopes to hand out, so the credential itself cannot be narrowed, no matter where it is stored. What is narrowed is the interface: six read operations, one create-only write with no calendar argument, no path traversal, size and count caps, and a credential the model never touches. Vault changed where the blast radius sits rather than removing it. Compromise the gateway host and you are still one authorization decision away from an account; compromise the model and you have twelve typed calls, all logged.
Wrapping third-party text in a "this is untrusted data" marker is a hint, not a boundary. Instructions cannot be reliably separated from prose. The real defence is having nothing worth redirecting to.
Scopes are frozen at login. Removing somebody from a writing group does not end a session already in flight. Fixing that means re-checking group membership on refresh, which is cheap, but it only means something once there is a revocation path worth the name.
The operational bits are proof-of-concept grade. The rate limiter is in-process, ticket state is in memory, and the free tunnel used for public exposure rotates its hostname on restart, so the connector has to be re-added each time.
The Nextcloud verification ran against exactly one instance, version 34.0.1, with one service account, so I cannot say anything about older versions. The test suites check that the described behaviours hold, which is functional testing rather than adversarial testing. No penetration test was performed, and two clients connecting rules out "it only works because one vendor is lenient" and nothing more than that.
The build is smaller than the topic makes it sound, and the shape is roughly this:
The code is not the hard part. The design decisions are, and the pattern in the ones that worked is the same each time: every piece of enforcement I handed to something built for it, Keycloak, Vault, Nextcloud's own permissions, is a piece I can no longer get wrong. The tool surface decides the set of things that can happen, and that set is fixed in a text editor, before any model ever sees it. You do not make an LLM integration safe by getting the model to behave well; you make it safe by leaving misbehaviour with nothing interesting to do.
If you are looking at something similar, whether that is deciding which internal services are worth exposing, what the tool surface should look like, or how to run it without handing out VPN accounts, we at Infralovers are happy to help think it through, especially in regulated or security-conscious environments. We also run courses on MCP server development and HashiCorp Vault if you would rather build the knowledge in-house.
Anthropic, MCP tunnels. The sentence "MCP tunnels created through the Console are not available as connectors in claude.ai" is quoted verbatim from platform.claude.com, where tunnels are also described as a research preview provided "as-is" with no uptime, support or continuity commitment, serving Managed Agents and the Messages API. The complementary requirement for custom connectors, "your MCP server must be reachable over the public internet from Anthropic's IP ranges," is from support.claude.com. Both verified against the current pages; product surfaces change, so re-check before relying on either. ↩︎
Model Context Protocol, Authorization specification, revision 2025-11-25. "MCP servers MUST only accept tokens that are valid for use with their own resources. MCP servers MUST NOT accept or transit any other tokens," and, on upstream calls, "The MCP server MUST NOT pass through the token it received from the MCP client." modelcontextprotocol.io, verified against the specification text. This is also the revision both clients actually negotiate, recorded on every handshake in the audit log, which is why the project has not moved to the newer 2026-07-28 revision: nothing is asking for it yet. ↩︎
fastmcp 3.4.5 ships an in-memory OAuth provider intended for local development and testing. Observed behaviour during this build: it runs a complete OAuth 2.1 authorization-code flow and auto-approves every authorization request without capturing a user identity, so the issued token carries no subject. This is a testing utility behaving as documented rather than a defect. The point is that it is hard to tell apart from working authentication unless you look for the subject. Version-specific, so check it against whatever you are running. ↩︎
Nextcloud admin manual, OAuth2 configuration. Both quoted sentences verified verbatim against docs.nextcloud.com; the page tracks the current release, so wording can change. PKCE support is tracked as nextcloud/server#12881, "Implement OAUTH2 Authorization code with PKCE", opened December 2018 and still open at the time of writing. The Authorization: Bearer failure against WebDAV is nextcloud/server#5512, "No 'Authorization: Bearer' header found." That one is closed, so treat it as evidence that the failure mode existed and is version-dependent, not as proof that it is present in a given release. Both issue states verified. ↩︎
Nextcloud developer manual, Login Flow v2. Anonymous POST to the login-flow endpoint returns a login URL and a poll token; the user authenticates in the default browser, including 2FA, against a session that lives for five minutes; the client polls the poll endpoint, which returns 404 until authentication succeeds and then returns the server address, the login name and an app password. Verified against docs.nextcloud.com and against one instance running version 34.0.1. ↩︎
You are interested in our courses or you simply have a question that needs answering? You can contact us at anytime! We will do our best to answer all your questions.
Contact us