llm-catalog-archive

Change

2d94d4d

2d94d4d9d479e7f0b8d72be555caa50ec6fc7956 · commit on GitHub

aws-blog-feed: changed (529755 bytes, HTTP 200)

raw/aws-blog-feed/response.xml modified

Lines added
+200
Lines removed
-604
Stored bytes at this commit
529,755
Timestamp
observed
Raw artifact at this commit
raw/aws-blog-feed/response.xml
Recorded headers
observed_at2026-10-10T05:53:15.944Z
origin_datenull
status200
final URLhttps://aws.amazon.com/blogs/machine-learning/feed/
etagnull
last-modifiedFri, 09 Oct 2026 22:09:31 GMT
dateSat, 10 Oct 2026 05:53:15 GMT
agenull
cache-controlnull
cf-cache-statusnull
content-encodingnull
content-lengthnull
@@@ -5,7 +5,7 @@
<atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/>
<link>https://aws.amazon.com/blogs/machine-learning/</link>
<description>Official Machine Learning Blog of Amazon Web Services</description>
- <lastBuildDate>Thu, 08 Oct 2026 18:33:29 +0000</lastBuildDate>
+ <lastBuildDate>Fri, 09 Oct 2026 15:38:39 +0000</lastBuildDate>
<language>en-US</language>
<sy:updatePeriod>
hourly </sy:updatePeriod>
@@@ -13,6 +13,205 @@
1 </sy:updateFrequency>
<item>
+ <title>ICYMI: What landed for AI builders in September 2026</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/icymi-what-landed-for-ai-builders-in-september-2026/</link>
+
+ <dc:creator><![CDATA[Prachi Mishra]]></dc:creator>
+ <pubDate>Fri, 09 Oct 2026 15:38:39 +0000</pubDate>
+ <category><![CDATA[Amazon Bedrock]]></category>
+ <category><![CDATA[Amazon Bedrock AgentCore]]></category>
+ <category><![CDATA[Foundational (100)]]></category>
+ <category><![CDATA[Thought Leadership]]></category>
+ <guid isPermaLink="false">47f6ae2bd79e8fc92d958051f1fc9a5345f3e76e</guid>
+
+ <description>A monthly recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026: broader model choice, faster serverless agents with built-in evaluation, and automated knowledge base syncing with native enterprise connectors.</description>
+ <content:encoded>&lt;p&gt;&lt;em&gt;A recap of the latest Amazon Bedrock, Amazon Bedrock AgentCore, and Strands updates from September 2026&lt;/em&gt;&lt;/p&gt;
+&lt;p&gt;At AWS, we believe the fastest path to enterprise AI is giving builders real choice at every layer: model, runtime, and tooling. That belief shapes how we invest across Amazon Bedrock, AgentCore, and Strands. &lt;a href="https://aws.amazon.com/bedrock/" target="_blank" rel="noopener"&gt;Ama…
+&lt;p&gt;As AI models become more capable and model choice expands, the industry conversations are shifting. Model performance is no longer the only question. Customers now weigh cost against benefit for their specific use case, and those decisions matter when delegating work to agents in production…
+&lt;p&gt;In September, updates across Amazon Bedrock, AgentCore, and Strands strengthened those foundations. New capabilities expanded model choice, made agent operations more efficient and measurable, and created more direct ways to connect AI applications with current business information.&lt;/p&g…
+&lt;h2 id="run-faster-agents-and-evaluate-what-matters"&gt;Run faster agents and evaluate what matters&lt;/h2&gt;
+&lt;p&gt;&lt;strong&gt;Build agents optimized for OpenAI models.&lt;/strong&gt; &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/09/bedrock-managed-agents-preview/" target="_blank" rel="noopener"&gt;Amazon Bedrock Managed Agents&lt;/a&gt;, powered by OpenAI, is now available in public pre…
+&lt;p&gt;&lt;strong&gt;Agents now start faster with less infrastructure overhead.&lt;/strong&gt; The latest &lt;a href="https://aws.amazon.com/blogs/machine-learning/the-new-agentcore-runtime-elastic-optimized-and-consistently-fast-starts/" target="_blank" rel="noopener"&gt;AgentCore runtime&lt;/a&g…
+&lt;p&gt;&lt;strong&gt;Lower your token cost with frontier performance.&lt;/strong&gt; &lt;a href="https://strandsagents.com/blog/introducing-strands-harness/" target="_blank" rel="noopener"&gt;Strands harness&lt;/a&gt; is a new open source agent harness that matches popular harnesses on accuracy wh…
+&lt;p&gt;&lt;strong&gt;Let your models make big decisions safely, locally.&lt;/strong&gt; &lt;a href="https://strandsagents.com/blog/introducing-strands-decider/" target="_blank" rel="noopener"&gt;Strands Decider 2B&lt;/a&gt; is a small, open source, 2B-parameter decision model that picks between pr…
+&lt;h2 id="find-the-right-model-for-each-workload"&gt;Find the right model for each workload&lt;/h2&gt;
+&lt;p&gt;&lt;strong&gt;Choose from a broader set of OpenAI models for your everyday work.&lt;/strong&gt; OpenAI’s Astra, Sol, and Luna models are now generally available on Amazon Bedrock, with options for demanding projects, recurring complex work, and high-volume tasks.&lt;/p&gt;
+&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/machine-learning/take-on-your-most-ambitious-work-with-gpt-6-astra-on-amazon-bedrock/" target="_blank" rel="noopener"&gt;GPT-6 Astra&lt;/a&gt; is the flagship model for your most ambitious work, bringing deeper reasoning to complex decisions, documen…
+&lt;p&gt;Sol brings near-Astra intelligence to coding, computer use, and professional workloads that run frequently. Both &lt;a href="https://aws.amazon.com/blogs/machine-learning/bring-near-astra-intelligence-to-everyday-work-with-gpt-6-1-sol-on-amazon-bedrock/" target="_blank" rel="noopener"&gt;GP…
+&lt;p&gt;&lt;strong&gt;Take on long-running coding and research with Claude&lt;/strong&gt;. Model choices for Claude now include the latest versions to advance your coding and scientific work. &lt;a href="https://aws.amazon.com/blogs/machine-learning/introducing-claude-fable-5-1-on-aws/" target="_bl…
+&lt;p&gt;&lt;strong&gt;Reason through complexity with a broader model portfolio.&lt;/strong&gt; Moonshot AI’s Kimi K3 is now available on Amazon Bedrock. According to Moonshot AI, &lt;a href="https://aws.amazon.com/blogs/machine-learning/introducing-kimi-k3-on-amazon-bedrock/" target="_blank" rel="n…
+&lt;p&gt;xAI’s frontier models &lt;a href="https://aws.amazon.com/blogs/machine-learning/xais-grok-4-6-is-now-available-in-amazon-bedrock/" target="_blank" rel="noopener"&gt;Grok 4.6&lt;/a&gt; and &lt;a href="https://aws.amazon.com/blogs/machine-learning/grok-4-7-is-now-available-on-amazon-bedrock/"…
+&lt;p&gt;Together, these additions give you more flexibility when matching models to your requirements such as context length, latency, throughput, and cost.&lt;/p&gt;
+&lt;h2 id="get-more-control-over-your-enterprise-ai-agents"&gt;Get more control over your enterprise AI agents&lt;/h2&gt;
+&lt;p&gt;&lt;strong&gt;Keep agent knowledge current.&lt;/strong&gt; With new &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/09/amazon-bedrock-managed-knowledge-base-automatic-sync-scheduling-data-source-connectors/" target="_blank" rel="noopener"&gt;automatic sync scheduling&lt;/a&gt; i…
+&lt;p&gt;&lt;strong&gt;Connect more enterprise sources without maintaining custom ingestion pipelines.&lt;/strong&gt; Amazon Bedrock Managed Knowledge Base now supports &lt;a href="https://aws.amazon.com/about-aws/whats-new/2026/09/amazon-bedrock-managed-knowledge-base-servicenow-native-data-source-…
+&lt;h2 id="get-started"&gt;Get started&lt;/h2&gt;
+&lt;p&gt;Explore &lt;a href="https://aws-samples.github.io/sample-amazon-bedrock-central/" target="_blank" rel="noopener"&gt;Amazon Bedrock&lt;/a&gt;, deploy agents with the &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/agentcore-get-started-cli.html" target="_blank" rel=…
+&lt;p&gt;Interested in learning how Amazon Bedrock can support your team? &lt;a href="https://pages.awscloud.com/Amazon-Bedrock-Contact-Us.html" target="_blank" rel="noopener"&gt;Connect with us&lt;/a&gt; to start the conversation.&lt;/p&gt;
+&lt;p style="clear: both"&gt;&lt;/p&gt;
+&lt;hr style="width: 100%"&gt;
+&lt;h2&gt;About the author&lt;/h2&gt;
+&lt;footer&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;img class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-22094-1.jpg" alt="Prachi Mishra" width="100" height="133"&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Prachi Mishra&lt;/h3&gt;
+ &lt;p style="overflow: hidden"&gt;Prachi is a Senior Product Marketing Manager for Amazon Bedrock AgentCore at Amazon Web Services (AWS), where she helps bring the power of production-ready AI agents to developers and enterprises worldwide.&lt;/p&gt;
+ &lt;/div&gt;
+&lt;/footer&gt;</content:encoded>
+
+
+
+ </item>
+ <item>
+ <title>How Postman runs Agent Mode for 40 million developers on Amazon Bedrock</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/how-postman-runs-agent-mode-for-40-million-developers-on-amazon-bedrock/</link>
+
+ <dc:creator><![CDATA[Srinivas Kini]]></dc:creator>
+ <pubDate>Fri, 09 Oct 2026 15:35:02 +0000</pubDate>
+ <category><![CDATA[Advanced (300)]]></category>
+ <category><![CDATA[Amazon Bedrock]]></category>
+ <category><![CDATA[Customer Solutions]]></category>
+ <guid isPermaLink="false">b70b3126633995a55ec3140e570c01681b054d89</guid>
+
+ <description>Building an AI agent that works in a demo is a different problem from running one for 40 million developers. Postman and AWS share the architectural patterns behind Agent Mode: controlling tool sprawl, exposing schema-based reads, and treating context as the real bottleneck, plus h…
+ <content:encoded>&lt;p&gt;Building an AI agent for a demo and operating one for &lt;a href="https://www.postman.com/company/about-postman/" target="_blank" rel="noopener"&gt;40 million developers&lt;/a&gt; are different engineering problems. Postman set out to build &lt;a href="https://lea…
+&lt;p&gt;In this post, Postman and AWS describe the architectural patterns that emerged while making a mature product legible to an AI agent. These patterns include controlling tool sprawl, exposing schema-based reads, and treating context rather than capability as the primary bottleneck.&lt;/p&gt;
+&lt;p&gt;We also explain how &lt;a href="https://learning.postman.com/docs/getting-started/basics/about-agent-mode/" target="_blank" rel="noopener"&gt;Agent Mode&lt;/a&gt; uses &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html" target="_blank" rel="noopener"&gt;Am…
+&lt;h2 id="why-postman-built-agent-mode"&gt;Why Postman built Agent Mode&lt;/h2&gt;
+&lt;p&gt;&lt;a href="https://learning.postman.com/docs/getting-started/basics/about-agent-mode/" target="_blank" rel="noopener"&gt;Agent Mode&lt;/a&gt; is Postman’s portal for working with the product in an AI-native way across testing, documentation, discovery, and implementation. Postman has evolv…
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-1.png" alt="Postman Agent Mode open…
+ &lt;p class="wp-caption-text"&gt;Figure 1: Agent Mode works directly against the Postman application. In this example, it opens a pull request and proposes next steps without requiring the user to navigate through the interface&lt;/p&gt;
+&lt;/div&gt;
+&lt;p&gt;&lt;a href="https://learning.postman.com/docs/getting-started/basics/about-agent-mode/" target="_blank" rel="noopener"&gt;Agent Mode&lt;/a&gt; runs on &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html" target="_blank" rel="noopener"&gt;Amazon Bedrock&lt;/…
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-2.png" alt="Postman Agent Mode arch…
+ &lt;p class="wp-caption-text"&gt;Figure 2: Postman Agent Mode combines client-side tools, agent orchestration, purpose-built context, and Amazon Bedrock model inference. Tools are scoped for each task, and user approval remains part of actions that modify application state&lt;/p&gt;
+&lt;/div&gt;
+&lt;p&gt;Human oversight is part of the production design. Agent Mode requires user approval before actions that modify application state. Postman also scopes available tools to the task, selects purpose-built context, and applies model-dependent data-retention settings. These controls reduce uninte…
+&lt;h2 id="handling-tool-sprawl"&gt;Handling tool sprawl&lt;/h2&gt;
+&lt;p&gt;In Agent Mode, tools define how the agent acts inside Postman. Early on, the team leaned toward highly atomic tools: small, precise actions such as opening a request, updating one field, or fetching a specific piece of metadata. That approach supported correctness and control in early itera…
+&lt;p&gt;Many real-world workflows require long sequences of tool calls. Even when each step was fast, the overall experience felt slow, because every action had to return to the model before the next one could begin. Users watched the agent step through actions they had mentally grouped as a single…
+&lt;p&gt;In Postman’s testing, tool-selection errors increased once the visible toolset exceeded approximately 40 tools. The agent could call nonexistent tools, pass incorrect arguments despite valid schemas, or select tools that seemed semantically reasonable but were wrong in context. Larger or ne…
+&lt;p&gt;Past a certain toolset size, exposing more tools can reduce agent effectiveness. The current architecture selects tools based on need and context and isolates individual execution threads. The model sees only the tools relevant to the current task. Figure 3 illustrates this dynamic selectio…
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-3.png" alt="Root agent narrowing mo…
+ &lt;p class="wp-caption-text"&gt;Figure 3: The root agent queries a vector database of tool embeddings and narrows more than 170 tools to approximately 15 relevant to the request. It then hands those tools to a context-isolated sub-agent, so the model sees only the tools needed for the task&lt;/p&g…
+&lt;/div&gt;
+&lt;p&gt;A subtler problem was that many client APIs were implicitly coupled to interface state. Tools that modified requests needed certain elements to be open, while other tools opened new tabs as side effects. The agent had to open a request tab to read it, mimicking interface interactions instea…
+&lt;p&gt;&lt;strong&gt;Builder takeaway:&lt;/strong&gt; Treat your tool catalog as part of the context budget. Dynamically scope the tools exposed to the model per task, and decouple “what the agent can do” from “what the UI happens to have open.”&lt;/p&gt;
+&lt;h2 id="exposing-schema-based-reads"&gt;Exposing schema-based reads&lt;/h2&gt;
+&lt;p&gt;For products such as the API Catalog, Postman consolidated multiple narrow views into a single query tool. These products expose structured data such as service uptime, test results, and endpoint response times across many services.&lt;/p&gt;
+&lt;p&gt;Given the schemas of the underlying ClickHouse tables, the agent can generate complex queries with joins and WHERE clauses. This substantially reduces the number of distinct tools needed to answer an analysis question:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-sql"&gt;SELECT toString(service_id) AS service_id,
+ countMerge(total_events_state) AS total_requests,
+ countMerge(error_events_state) AS total_errors,
+ round(countMerge(error_events_state) * 100.0
+ / countMerge(total_events_state), 4) AS error_rate_pct,
+ avgMerge(avg_latency_state) AS avg_latency_ms,
+ quantileMerge(0.95)(p95_latency_state) AS p95_latency_ms
+FROM http_events_summary_1d
+WHERE service_id IN ('...list of service IDs')
+ AND bucket_1d &amp;gt;= today() - 7
+GROUP BY service_id
+HAVING p95_latency_ms &amp;lt; 100
+ AND total_requests &amp;gt; 0
+ORDER BY error_rate_pct DESC;&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;With this approach, the engineering job shifts from &lt;em&gt;building a tool per question&lt;/em&gt; to &lt;em&gt;modeling the data well once&lt;/em&gt;. The agent can then generate a far wider variety of queries than the team could ever have enumerated as individual tools.&lt;/p&gt;
+&lt;p&gt;&lt;strong&gt;Builder takeaway:&lt;/strong&gt; Where you have well-structured data, give the agent schema-aware read access to a query engine instead of a proliferation of single-purpose read tools. You trade tool count for data modeling, which produces a better scaling curve.&lt;/p&gt;
+&lt;h2 id="context-was-the-real-bottleneck"&gt;Context was the real bottleneck&lt;/h2&gt;
+&lt;p&gt;Postman initially assumed missing &lt;em&gt;tools&lt;/em&gt; would be the biggest blocker. In practice, missing or incomplete &lt;em&gt;context&lt;/em&gt; caused more failures than missing capabilities.&lt;/p&gt;
+&lt;p&gt;Context is the agent’s understanding of where the user is in Postman, which entities are active, and what state has already been established. When that context was wrong or absent, even correct tools became ineffective. Figure 4 distinguishes the two forms of context supplied to the agent.&…
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-4.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-4.png" alt="Background context gath…
+ &lt;p class="wp-caption-text"&gt;Figure 4: Two kinds of context feed the agent. Broad, shallow background context is gathered automatically and minified for the prompt. Deep, focused selected context is chosen by the user and routed through a dedicated handler for each entity type. Each handler dis…
+&lt;/div&gt;
+&lt;p&gt;The challenge was structural. Over 11 years, developers and users learned to find information through the interface. Re-engineering that awareness for an agent required multiple iterations to determine what mattered for each workflow and what was noise. Serializing the existing interface da…
+&lt;p&gt;As more objects gained handlers, truncation became the next problem. Many fields contain open-ended user-generated data, including request descriptions, OpenAPI specifications, and request payloads. This data can crowd the context window. Managing the context budget carefully is essential a…
+&lt;p&gt;&lt;strong&gt;Builder takeaway:&lt;/strong&gt; Don’t feed the model your rendering data model. Build purpose-shaped context handlers and treat the context window as a scarce, actively managed budget. Noise crowds out signal long before the model reaches its limit.&lt;/p&gt;
+&lt;h2 id="putting-it-all-together"&gt;Putting it all together&lt;/h2&gt;
+&lt;p&gt;As Agent Mode evolved, it became clear the system had to aggregate three distinct components, each solving a different problem.&lt;/p&gt;
+&lt;ol type="1"&gt;
+ &lt;li&gt;Client-side tools live in the Postman application and represent the final actions the agent can take, such as opening requests, modifying settings, running collections, and inspecting authentication. Agent Mode also uses server-side tools for functions such as web search and agent-loop ma…
+ &lt;li&gt;Generic agent instructions define system-level behavior, including how proactive Agent Mode should be, how it communicates uncertainty, and what baseline product knowledge it carries.&lt;/li&gt;
+ &lt;li&gt;A knowledge base uses a Retrieval Augmented Generation (RAG) approach. Postman has a large product surface that includes multiple request protocols, mock servers, monitors, documentation, the API Network, workspace governance, variables, helpers, code generation, request settings, and col…
+&lt;/ol&gt;
+&lt;p&gt;Encoding all of this in static prompts was not feasible, and most of it is irrelevant to a given query. For initial seeding, the team used Postman’s Learning Center to generate concise feature-specific articles. At runtime, Agent Mode selects knowledge articles based on the incoming query a…
+&lt;h2 id="running-agent-mode-on-amazon-bedrock"&gt;Running Agent Mode on Amazon Bedrock&lt;/h2&gt;
+&lt;p&gt;The three components previously described resolve to the same runtime action: an inference call to a foundation model (FM). At Postman’s scale, traffic is bursty and developer-driven. Routing, caching, and geographic processing controls help Postman accommodate traffic bursts, manage infere…
+&lt;h3 id="model-flexibility-across-the-claude-family"&gt;Model flexibility across the Claude family&lt;/h3&gt;
+&lt;p&gt;Agent Mode isn’t tied to one model. Through &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/welcome.html" target="_blank" rel="noopener"&gt;Amazon Bedrock model inference APIs&lt;/a&gt;, Postman can access supported Anthropic Claude models and route each workload to an a…
+&lt;h3 id="cross-region-inference-for-high-throughput"&gt;Cross-Region inference for high throughput&lt;/h3&gt;
+&lt;p&gt;Developer traffic is spiky, and provisioning for peak demand in one AWS Region can be costly. Agent Mode uses &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/cross-region-inference.html" target="_blank" rel="noopener"&gt;Amazon Bedrock cross-Region inference&lt;/a&gt; to au…
+&lt;ol type="1"&gt;
+ &lt;li&gt;&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/geographic-cross-region-inference.html" target="_blank" rel="noopener"&gt;Geographic inference profiles&lt;/a&gt; route requests only among supported Regions within a defined geography, such as the United States or European …
+ &lt;li&gt;Global inference profiles can route requests among supported destination Regions worldwide to provide additional throughput during traffic bursts. They are appropriate only when the workload does not require a geographically constrained processing boundary.&lt;/li&gt;
+&lt;/ol&gt;
+&lt;p&gt;Postman can select the inference profile per workload: a global profile for maximum available throughput or a geographic profile when processing must remain within the profile’s defined geography. This choice is explicit in the modelId used for each Bedrock inference request.&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;# Schematic Converse request
+response = bedrock_runtime.converse(
+ modelId="&amp;lt;geographic-inference-profile-id-or-arn&amp;gt;",
+ messages=messages,
+ system=system_blocks,
+)&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;h3 id="data-residency-and-enterprise-controls"&gt;Data residency and enterprise controls&lt;/h3&gt;
+&lt;p&gt;For enterprise customers, permitted processing geography can be as important as throughput. Geographic inference profiles constrain Bedrock routing to the profile’s supported destination Regions within the selected geography. This does not mean that inference runs inside Postman’s own AWS e…
+&lt;h3 id="prompt-caching-to-keep-costs-in-check"&gt;Prompt caching to keep costs in check&lt;/h3&gt;
+&lt;p&gt;A production agent resends substantial stable context on each turn, including system instructions, generic agent behavior, a core tool set, selected knowledge, and conversation context. Reprocessing the unchanged prefix on every request adds avoidable latency and cost.&lt;/p&gt;
+&lt;p&gt;Agent Mode uses &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html" target="_blank" rel="noopener"&gt;Amazon Bedrock prompt caching&lt;/a&gt; to reuse stable prompt prefixes. The near-immutable core, including the system prompt, agent instructions, and core…
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;# Schematic cache checkpoints in Converse content blocks
+{"cachePoint": {"type": "default", "ttl": "1h"}} # stable core
+{"cachePoint": {"type": "default", "ttl": "5m"}} # variable layer&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;&lt;strong&gt;Builder takeaway:&lt;/strong&gt; Treat inference as a routing-and-caching problem, not only a model-selection decision. Select the Claude model per workload, choose the appropriate cross-Region inference profile, and cache the stable prompt prefix with TTLs that match how freq…
+&lt;h2 id="best-practices-for-scaling-agents-in-production"&gt;Best practices for scaling agents in production&lt;/h2&gt;
+&lt;p&gt;Distilled from Postman’s journey, for builders working on Amazon Bedrock:&lt;/p&gt;
+&lt;ol type="1"&gt;
+ &lt;li&gt;&lt;strong&gt;Budget tools as carefully as tokens.&lt;/strong&gt; Dynamically select the tools exposed per task. In Postman’s testing, tool-selection errors increased as the visible toolset became large.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Prefer schema-aware reads over tool proliferation.&lt;/strong&gt; Model your data well and let the agent query it.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Decouple agent actions from interface state.&lt;/strong&gt; If a tool requires an open tab, the agent is navigating the interface rather than reasoning directly over data.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Engineer context deliberately.&lt;/strong&gt; Purpose-built context handlers beat serializing your rendering model every time.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Manage the context window as a scarce resource.&lt;/strong&gt; Truncation and expansion strategy is a first-class design problem, not an afterthought.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Ship docs with features.&lt;/strong&gt; A RAG knowledge base only stays useful if it evolves in lockstep with the product.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Route and cache on Bedrock.&lt;/strong&gt; Match each workload to the appropriate Claude model, choose cross-Region inference based on throughput and geographic requirements, and apply tiered caching to stable prompt prefixes.&lt;/li&gt;
+&lt;/ol&gt;
+&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
+&lt;p&gt;Building Agent Mode required Postman to confront the gap between large language model capabilities and the structure of mature products: interface assumptions, coupled clients, sprawling tool catalogs, and knowledge distributed across documentation and teams. Dynamic tool selection, schema-…
+&lt;p&gt;Whether you’re building your first agent or scaling an existing one, these patterns can help teams avoid common agent-integration and scaling challenges.&lt;/p&gt;
+&lt;p&gt;To learn more, see the &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/what-is-bedrock.html" target="_blank" rel="noopener"&gt;Amazon Bedrock documentation&lt;/a&gt;, including guidance for &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/cross-region-infere…
+&lt;p&gt;Postman’s production implementation is proprietary and isn’t available as a public sample repository.&lt;/p&gt;
+&lt;p&gt;&amp;nbsp;&lt;/p&gt;
+&lt;p style="clear: both"&gt;&lt;/p&gt;
+&lt;hr style="width: 100%"&gt;
+&lt;h2&gt;About the authors&lt;/h2&gt;
+&lt;footer&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-5.jpg" alt="Srinivas Kini" width="100" height="133"&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Srinivas Kini&lt;/h3&gt;
+ &lt;p style="overflow: hidden"&gt;Srinivas is a Senior Engineer on Postman’s AI team, building enterprise agents at the intersection of distributed systems and AI infrastructure. His focus is core agent architecture that keeps agents reliable and accurate at scale.&lt;/p&gt;
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ML-21683-6.jpg" alt="Shubham Gupta" width="100" height="133"&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Shubham Gupta&lt;/h3&gt;
+ &lt;p style="overflow: hidden"&gt;Shubham is a Solutions Architect at AWS based in Bengaluru, India, supporting independent software vendors (ISVs). He works with engineering and leadership teams to design, build, and run their products on AWS, from first architecture to production, with a deep fo…
+ &lt;/div&gt;
+&lt;/footer&gt;</content:encoded>
+
+
+
+ </item>
+ <item>
<title>Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments</title>
<link>https://aws.amazon.com/blogs/machine-learning/pay-per-inference-for-ai-agents-how-blockrun-and-incarna-use-amazon-bedrock-agentcore-payments/</link>
@@@ -3271,609 +3470,6 @@ export CLAUDE_CODE_USE_MANTLE=1&lt;/code&gt;&lt;/pre&gt;
- </item>
- <item>
- <title>Agentic retrieval with LangChain and Amazon Bedrock Knowledge Bases</title>
- <link>https://aws.amazon.com/blogs/machine-learning/agentic-retrieval-with-langchain-and-amazon-bedrock-knowledge-bases/</link>
-
- <dc:creator><![CDATA[Manideep Reddy Gillela]]></dc:creator>
- <pubDate>Mon, 05 Oct 2026 15:53:56 +0000</pubDate>
- <category><![CDATA[Advanced (300)]]></category>
- <category><![CDATA[Amazon Bedrock Knowledge Bases]]></category>
- <category><![CDATA[Technical How-to]]></category>
- <guid isPermaLink="false">9ef1f1b789cbe6fd45278d748b65ebf4500d6c5e</guid>
-
- <description>Build a Retrieval Augmented Generation (RAG) application on Amazon Bedrock Managed Knowledge Base with LangChain, and see how agentic retrieval handles the multi-part questions that single-shot retrieval answers poorly. Run the same query through both paths, read the trace events, …
- <content:encoded>&lt;p&gt;When a user asks the support assistant, a Retrieval Augmented Generation (RAG) application built with LangChain to compare two products across three dimensions, they’re effectively posing six questions simultaneously. Similarity search uses a single query vector t…
-&lt;p&gt;In this post, we showcase a RAG application on Amazon Bedrock Managed Knowledge Base with &lt;a href="https://github.com/langchain-ai/langchain-aws" target="_blank" rel="noopener"&gt;LangChain&lt;/a&gt;. We run the same multi-part question through standard and agentic retrieval, and read th…
-&lt;p&gt;Agentic retrieval is available on Amazon Bedrock Managed Knowledge Base. Instead of one search, Amazon Bedrock Managed Knowledge Base plans the retrieval. It breaks the question into sub-queries, runs them, judges whether it has enough evidence, and searches again if it doesn’t. The &lt;cod…
-&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
-&lt;p&gt;&lt;a href="https://aws.amazon.com/blogs/aws/introducing-amazon-bedrock-managed-knowledge-base-for-faster-more-accurate-enterprise-ai-applications/" target="_blank" rel="noopener"&gt;Amazon Bedrock Managed Knowledge Base&lt;/a&gt;, the fully managed RAG capability in Amazon Bedrock, removes…
-&lt;p&gt;Amazon Bedrock Managed Knowledge Bases provides two APIs. We briefly discuss those differences in this post. The &lt;a href="https://docs.aws.amazon.com/bedrock/latest/APIReference/API_agent-runtime_Retrieve.html" target="_blank" rel="noopener"&gt;Retrieve&lt;/a&gt; API runs one hybrid sear…
-&lt;p&gt;The following diagram shows the solution architecture. The application queries Amazon Bedrock Knowledge Bases using either the Retrieve API (standard, single-shot) or the AgenticRetrieveStream API (multi-step planning loop). Both paths return document chunks from the knowledge base, which t…
-&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
- &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/29/ML-21691-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/29/ML-21691-1.png" alt="Application querying Am…
- &lt;p class="wp-caption-text"&gt;Figure 1: Solution architecture for querying Amazon Bedrock Knowledge Bases with the Retrieve and AgenticRetrieveStream APIs&lt;/p&gt;
-&lt;/div&gt;
-&lt;h2 id="implementation-walkthrough"&gt;Implementation walkthrough&lt;/h2&gt;
-&lt;p&gt;The following sections walk you through creating a knowledge base, querying it with both retrieval methods, and reading the trace events the agentic planner produces.&lt;/p&gt;
-&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
-&lt;p&gt;To follow along you need:&lt;/p&gt;
-&lt;ul&gt;
- &lt;li&gt;An AWS account with access to Amazon Bedrock in a Region where Amazon Bedrock Managed Knowledge Bases and agentic retrieval are available. This walkthrough uses the US East (N. Virginia) Region (&lt;code&gt;us-east-1&lt;/code&gt;), and the code assumes it throughout. Check the AWS &lt;a h…
- &lt;li&gt;Two AWS Identity and Access Management (IAM) identities, described in the next section: a service role the knowledge base assumes, and permissions on the identity you call the APIs from.&lt;/li&gt;
- &lt;li&gt;&lt;a href="https://www.python.org/downloads/" target="_blank" rel="noopener"&gt;Python 3.12&lt;/a&gt; or later.&lt;/li&gt;
- &lt;li&gt;An S3 bucket holding the sample documents. The corpus needs several documents that cover overlapping topics so that a comparative question has somewhere to go. A single flat document cannot demonstrate query planning.&lt;/li&gt;
-&lt;/ul&gt;
-&lt;p&gt;Install the packages. The &lt;a href="https://boto3.amazonaws.com/v1/documentation/api/latest/index.html" target="_blank" rel="noopener"&gt;Boto3&lt;/a&gt; version matters: &lt;code&gt;agentic_retrieve_stream&lt;/code&gt; did not exist before 1.43.32.&lt;/p&gt;
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-text"&gt;langchain-aws&amp;gt;=1.6.3
-langchain&amp;gt;=1.0
-boto3&amp;gt;=1.43.32&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;h2 id="permissions"&gt;Permissions&lt;/h2&gt;
-&lt;p&gt;Two identities are involved and separating them is worth doing deliberately. The knowledge base assumes a service role to read your documents and call the embedding model. Your application uses an AWS Security Token Service (AWS STS) caller identity to query. Neither needs the other’s permi…
-&lt;p&gt;Amazon Bedrock creates the service role for you if you let it. To supply your own, give it a trust policy that lets Amazon Bedrock assume it. Scope it with &lt;code&gt;aws:SourceAccount&lt;/code&gt; and &lt;code&gt;aws:SourceArn&lt;/code&gt; so that another account can’t use it as a confuse…
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-json"&gt;{
- "Version": "2012-10-17",
- "Statement": [{
- "Effect": "Allow",
- "Principal": {"Service": "bedrock.amazonaws.com"},
- "Action": "sts:AssumeRole",
- "Condition": {
- "StringEquals": {"aws:SourceAccount": "111122223333"},
- "ArnLike": {
- "aws:SourceArn": "arn:aws:bedrock:us-east-1:111122223333:knowledge-base/*"
- }
- }
- }]
-}&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;p&gt;The service role also needs &lt;code&gt;s3:ListBucket&lt;/code&gt; on your bucket and &lt;code&gt;s3:GetObject&lt;/code&gt; on its contents, both conditioned on &lt;code&gt;aws:ResourceAccount&lt;/code&gt;. Scope the &lt;code&gt;knowledge-base/*&lt;/code&gt; wildcard character down to speci…
-&lt;p&gt;The AWS STS caller identity needs a different set. &lt;code&gt;bedrock:AgenticRetrieveStream&lt;/code&gt; and &lt;code&gt;bedrock:InvokeModelWithResponseStream&lt;/code&gt; can’t be scoped to a knowledge base Amazon Resource Name (ARN). &lt;code&gt;bedrock:Retrieve&lt;/code&gt; and &lt;code…
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-json"&gt;{
- "Version": "2012-10-17",
- "Statement": [
- {
- "Sid": "AgenticRetrievalAndPlannerModel",
- "Effect": "Allow",
- "Action": [
- "bedrock:AgenticRetrieveStream",
- "bedrock:InvokeModelWithResponseStream"
- ],
- "Resource": "*"
- },
- {
- "Sid": "RetrieveAndFullDocumentExpansion",
- "Effect": "Allow",
- "Action": ["bedrock:Retrieve", "bedrock:GetDocumentContent"],
- "Resource": "arn:aws:bedrock:&amp;lt;region&amp;gt;:111122223333:knowledge-base/&amp;lt;knowledge-base-id&amp;gt;"
- },
- {
- "Sid": "GenerateAnswersInTheChains",
- "Effect": "Allow",
- "Action": ["bedrock:InvokeModel", "bedrock:Converse", "bedrock:ConverseStream"],
- "Resource": "*"
- }
- ]
-}&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;p&gt;&lt;code&gt;bedrock:GetDocumentContent&lt;/code&gt; is often overlooked. Agentic retrieval calls it when a &lt;code&gt;FullDocumentExpansion&lt;/code&gt; step decides a passage lacks the context to answer. A policy with only &lt;code&gt;bedrock:Retrieve&lt;/code&gt; works until the planner …
-&lt;p&gt;To create and manage the knowledge base itself, the calling role additionally needs &lt;code&gt;bedrock:CreateKnowledgeBase&lt;/code&gt; on &lt;code&gt;*&lt;/code&gt;, and the &lt;code&gt;GetKnowledgeBase&lt;/code&gt;, &lt;code&gt;UpdateKnowledgeBase&lt;/code&gt;, &lt;code&gt;DeleteKnowledg…
-&lt;p&gt;Running this walkthrough might incur costs for document storage and ingestion in the knowledge base, retrieval calls, and foundation model (FM) inference.&lt;/p&gt;
-&lt;p&gt;For more information about pricing, see the Knowledge Bases section of &lt;a href="https://aws.amazon.com/bedrock/pricing/" target="_blank" rel="noopener"&gt;Amazon Bedrock pricing&lt;/a&gt;.&lt;/p&gt;
-&lt;p&gt;Delete the resources when you complete this experiment.&lt;/p&gt;
-&lt;h2 id="creating-and-populating-the-knowledge-base"&gt;Creating and populating the knowledge base&lt;/h2&gt;
-&lt;p&gt;Create the knowledge base with a &lt;code&gt;managedKnowledgeBaseConfiguration&lt;/code&gt;. Setting &lt;code&gt;embeddingModelType&lt;/code&gt; to &lt;code&gt;MANAGED&lt;/code&gt; uses the service-managed embedding model.&lt;/p&gt;
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-python"&gt;import boto3
-import os
-
-REGION = os.environ["AWS_REGION"]
-bedrock_agent = boto3.client("bedrock-agent", region_name=REGION)
-
-response = bedrock_agent.create_knowledge_base(
- name=KB_NAME,
- roleArn=KB_ROLE_ARN,
- knowledgeBaseConfiguration={
- "type": "MANAGED",
- "managedKnowledgeBaseConfiguration": {
- "embeddingModelType": "MANAGED",
- },
- },
-)
-KB_ID = response["knowledgeBase"]["knowledgeBaseId"]&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;p&gt;There’s no &lt;code&gt;storageConfiguration&lt;/code&gt; in that request. For a self-managed knowledge base you would pass one describing your vector store. Amazon Bedrock Managed Knowledge Base does not take one, which is the clearest signal in the API that Amazon Bedrock owns the storage …
-&lt;p&gt;Attach the S3 bucket as a data source, then start an ingestion job. Ingestion is asynchronous, so poll until the job reaches a terminal state rather than sleeping for a fixed interval and hoping.&lt;/p&gt;
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-python"&gt;import time
-
-SUCCESS_STATES = frozenset({"COMPLETE"})
-FAILURE_STATES = frozenset({"FAILED", "STOPPED"})
-
-def wait_for_ingestion(kb_id, ds_id, job_id, timeout_s=1800):
- """Poll an ingestion job until it reaches a terminal state."""
- deadline = time.time() + timeout_s
- while time.time() &amp;lt; deadline:
- job = bedrock_agent.get_ingestion_job(
- knowledgeBaseId=kb_id,
- dataSourceId=ds_id,
- ingestionJobId=job_id,
- )["ingestionJob"]
- status = job["status"]
- if status in SUCCESS_STATES:
- return job
- if status in FAILURE_STATES:
- reasons = job.get("failureReasons") or ["no reason reported"]
- raise RuntimeError(f"Ingestion job {job_id} finished as {status}: " + "; ".join(reasons))
- time.sleep(15)
- raise TimeoutError(f"Ingestion job {job_id} did not finish in {timeout_s}s")&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;p&gt;The full data source configuration and error handling are in the &lt;a href="https://github.com/aws-samples/sample-rag-bedrock-langchain-blog" target="_blank" rel="noopener"&gt;sample repository&lt;/a&gt;.&lt;/p&gt;
-&lt;h2 id="querying-with-the-langchain-retriever"&gt;Querying with the LangChain retriever&lt;/h2&gt;
-&lt;p&gt;&lt;a href="https://reference.langchain.com/python/langchain-aws/retrievers/bedrock/AmazonKnowledgeBasesRetriever" target="_blank" rel="noopener"&gt;AmazonKnowledgeBasesRetriever&lt;/a&gt; wraps the Retrieve API and behaves like any other LangChain retriever. For Amazon Bedrock Managed Know…
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-python"&gt;from langchain_aws.retrievers import AmazonKnowledgeBasesRetriever
-
-SIMPLE_QUERY = "What is the restore time objective for the checkout service?"
-
-retriever = AmazonKnowledgeBasesRetriever(
- knowledge_base_id=KB_ID,
- region_name=REGION,
- retrieval_config={
- "managedSearchConfiguration": {
- "numberOfResults": 5,
- }
- },
-)
-docs = retriever.invoke(SIMPLE_QUERY)&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;p&gt;Each result comes back as a LangChain &lt;code&gt;Document&lt;/code&gt;. The relevance score is in &lt;code&gt;metadata["score"]&lt;/code&gt;, and the source document’s own metadata is under &lt;code&gt;metadata["source_metadata"]&lt;/code&gt;, renamed so it does not collide. If you want to…
-&lt;p&gt;For a question with one clear intent, this is the right tool. It is one call. The latency is the lowest of the two options, and you keep full control of how the answer gets generated. Most of the queries a production assistant sees are this shape, and reaching for a planning loop to answer …
-&lt;h2 id="where-single-shot-retrieval-runs-out"&gt;Where single-shot retrieval runs out&lt;/h2&gt;
-&lt;p&gt;Now give the same retriever a question with several parts:&lt;/p&gt;
-&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-python"&gt;COMPLEX_QUERY = (
- "Compare the checkout and inventory services across on-call escalation, backup and "
- "restore targets, and deployment rollback procedure. Where do they differ?"
-)
-
-docs = retriever.invoke(COMPLEX_QUERY)&lt;/code&gt;&lt;/pre&gt;
-&lt;/div&gt;
-&lt;p&gt;Five chunks come back, ranked by hybrid score against one embedding of that whole question.&lt;/p&gt;
-&lt;p&gt;That question contains six intents: two services across three dimensions. Scoring the retrieved text for evidence of each one gives a concrete measure of what a single embedding recovers.&lt;/p&gt;
-&lt;table border="1px" width="100%" cellpadding="10px"&gt;
- &lt;tbody&gt;
- &lt;tr&gt;
- &lt;td&gt;&lt;strong&gt;numberOfResults&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Chunks&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Share of corpus&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Sub-intents covered&lt;/strong&gt;&lt;/td&gt;

Diff display stops at 400 lines. The line counts above are from the whole diff. 67 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.