Postman has published an account, on the AWS Machine Learning Blog, of how it runs Agent Mode, its AI assistant inside the Postman application, for what it calls 40 million developers. The post is co-written by a Postman engineer and an AWS solutions architect and sits in AWS's customer-solutions series, so it is vendor-published: it describes design choices and lessons but gives no cost, latency, accuracy or adoption figures.

Too many tools made the agent worse

Postman says it first built highly atomic tools: open a request, update one field, fetch one piece of metadata. Long workflows then needed long chains of calls, each returning to the model, and users watched the agent step through what they thought of as one action. In Postman's testing, tool-selection errors rose once the visible toolset passed roughly 40 tools. The agent called tools that did not exist, passed wrong arguments despite valid schemas, or picked plausible but wrong tools. Larger and newer models reduced this but did not remove it.

The current design embeds the tool catalogue in a vector database. A root agent narrows more than 170 tools to about 15 for the request and hands them to a context-isolated sub-agent.

Context, not capability, was the bottleneck

Postman says it expected missing tools to be the main blocker. In practice, missing or wrong context caused more failures. Serialising the application's own data model did not work, because those objects were shaped for rendering and transfer rather than reasoning. Postman built a dedicated handler per entity type that distils what the agent needs, and says truncation of open-ended fields such as OpenAPI specs and request payloads became the next problem.

Two related findings. Many tools were coupled to interface state, so the agent had to open a tab to read a request. Postman says it is decoupling tools from tabs, and the agent can now send requests in the background, with approval still required. For analytics products, it replaced several narrow read tools with one query tool that lets the agent write SQL against documented ClickHouse schemas.

What runs on Bedrock

  • Model routing: Postman says it routes workloads across supported Claude models, a faster one for high-volume interactions and a larger one for complex reasoning. It describes switching models as mainly a configuration change.
  • Cross-Region inference: it chooses a global inference profile for maximum throughput, or a geographic profile when processing must stay within a defined geography. AWS notes this does not mean inference runs in Postman's own AWS environment.
  • Prompt caching: a one-hour checkpoint on the near-fixed prefix (system prompt, agent instructions, core tool definitions) and a five-minute checkpoint on variable context. The post says Bedrock requires the longer-lived checkpoint to come first.
  • Controls: user approval before actions that change application state, Bedrock Guardrails to redact personal data before it reaches the model (an admin setting), and zero data retention on supported models. AWS says availability is model-dependent.

What the post leaves out

There is no spend, cache hit rate, error rate before and after the tool changes, or share of traffic per model. The 40-tool threshold is Postman's observation on its own workload, not a general limit. The implementation is proprietary, and the post says there is no public sample repository. Treat the numbers as a starting hypothesis to test against your own tool set.

DiscussHow many tools does your production agent see at once, and have you measured tool-selection errors as that number grows?