Dev.to AI 🤖 Ai 👁 0 📖 3 min read

How Uber exposes existing APIs to AI agents through its MCP Gateway

onnecting an AI agent to one backend service is manageable. Doing it across hundreds of teams raises questions about ownership, permissions, protocol compatibility, and how agents find the right tools. Uber Engineering

How Uber exposes existing APIs to AI agents through its MCP Gateway

onnecting an AI agent to one backend service is manageable. Doing it across hundreds of teams raises questions about ownership, permissions, protocol compatibility, and how agents find the right tools.

Uber Engineering describes its approach in Designing MCP Gateway: Uber's MCP Management Platform. At the time of publication, the platform hosted more than 800 MCP servers and 5,000 tools.

This article summarizes the design choices that caught my attention as a backend developer. The implementation and diagrams are Uber's.

1. Put existing APIs behind a shared gateway

Uber already has internal services using HTTP, gRPC, and TChannel. Its gateway makes those APIs accessible through the Model Context Protocol (MCP), without requiring changes to the downstream services.

The architecture separates two responsibilities:

  • The MCP Registry, or control plane, stores tool definitions, ownership, and enablement settings.
  • The Proxy Gateway, or data plane, handles requests and translates between MCP and downstream protocols.

For each virtual MCP server, the gateway exposes a /<service-name>/mcp endpoint. Handlers translate requests into the downstream format and convert responses back into MCP-compatible results. Native MCP servers can also sit behind the gateway.

Uber's MCP Gateway routes calls from agents and other clients to backend services using different protocols.

Figure 1 from Uber Engineering's original article: MCP Gateway, APIs as tools.

The part I find useful here is API reuse. Teams can expose capabilities they already maintain, while the shared gateway handles MCP integration.

2. Generate tool definitions from service contracts

Manually maintaining thousands of tool definitions would create a substantial maintenance burden. Uber uses AutoCrawler, a discovery system built on Cadence workflows.

For services defined with Protobuf or Thrift, it reads the interface definitions, extracts methods and schemas, and generates MCP tool definitions. An LLM helps write descriptions using the schemas and documentation comments.

Native MCP servers follow a different path: AutoCrawler detects their heartbeat signals and calls listTools to retrieve their tools and schemas.

Uber's AutoCrawler workflow connects service discovery, tool generation, owner review, and the MCP Registry.

Figure 2 from Uber Engineering's original article: AutoCrawler.

Service contracts provide a starting point for the integration. Owners can then review the generated descriptions before agents use them.

3. Require owner approval before enabling tools

Every discovered server and tool starts disabled. The team responsible for the service must review and explicitly enable it.

Changes to tool descriptions produce configuration diffs that require owner approval. Previous configurations remain available for rollback.

At execution time, the gateway applies authorization policies with tool-level granularity and redacts sensitive data from responses. Its runtime periodically refreshes configuration in memory, so configuration changes can take effect without a service restart.

This is a detail I would carry into a smaller implementation: generating a tool definition and granting access to that tool should be separate operations.

4. Discover tools progressively and keep responses small

Thousands of tools create another problem: their descriptions and schemas can consume the model's context window.

Uber addresses this with Omni MCP, a single proxy server that supports progressive discovery through four tools:

  • discover_server
  • discover_tools
  • get_tool_schema
  • invoke_tool

An agent can search for a relevant server, inspect its tools, and retrieve a schema when needed.

The gateway also supports response projection. A caller specifies the fields it needs, and the gateway trims the response accordingly. That distinction matters: the article describes filtering at the gateway, so this should not be assumed to reduce the downstream service's work.

For coding agents, Uber offers Code Mode through its aifx CLI. Agents can discover and call tools, write results to files, and inspect selected content before bringing it into model context. Uber reports that this is its default approach for MCP use in coding agents.

What I would take into a smaller project

I would start with a few existing APIs, explicit tool ownership, and a review step before enablement. I would also measure how much context the tool schemas and responses consume before expanding the catalog.

Uber's scale explains the investment in automated discovery and a centralized platform. For a smaller backend, these individual design choices are still useful to evaluate.

Read the full Uber Engineering article for the runtime diagrams, registry screenshots, and third-party integration details.

Jatniel - Full-stack Developer
jatniel.dev

📰 Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.