llm-catalog-archive

Change

c81a97c

c81a97c312ff128feb7def592b372d82610131d4 · commit on GitHub

aws-blog-feed: changed (642110 bytes, HTTP 200)

raw/aws-blog-feed/response.xml modified

Lines added
+4,623
Lines removed
-5,529
Stored bytes at this commit
642,110
Timestamp
observed
Raw artifact at this commit
raw/aws-blog-feed/response.xml
Recorded headers
observed_at2026-10-07T05:59:54.365Z
origin_datenull
status200
final URLhttps://aws.amazon.com/blogs/machine-learning/feed/
etagnull
last-modifiedTue, 06 Oct 2026 22:43:44 GMT
dateWed, 07 Oct 2026 05:59:54 GMT
agenull
cache-controlnull
cf-cache-statusnull
content-encodingnull
content-lengthnull
@@@ -5,7 +5,7 @@
<atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/>
<link>https://aws.amazon.com/blogs/machine-learning/</link>
<description>Official Machine Learning Blog of Amazon Web Services</description>
- <lastBuildDate>Mon, 05 Oct 2026 23:25:17 +0000</lastBuildDate>
+ <lastBuildDate>Tue, 06 Oct 2026 22:43:32 +0000</lastBuildDate>
<language>en-US</language>
<sy:updatePeriod>
hourly </sy:updatePeriod>
@@@ -13,354 +13,287 @@
1 </sy:updateFrequency>
<item>
- <title>Introducing GLM 5.3 on Amazon Bedrock</title>
- <link>https://aws.amazon.com/blogs/machine-learning/introducing-glm-5-3-on-amazon-bedrock/</link>
+ <title>Building a context-aware AI assistant on AgentCore and OpenClaw</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/building-a-context-aware-ai-assistant-on-agentcore-and-openclaw/</link>
- <dc:creator><![CDATA[Alex Thewsey]]></dc:creator>
- <pubDate>Mon, 05 Oct 2026 23:25:17 +0000</pubDate>
+ <dc:creator><![CDATA[Thiago Verney]]></dc:creator>
+ <pubDate>Tue, 06 Oct 2026 19:19:15 +0000</pubDate>
<category><![CDATA[Advanced (300)]]></category>
- <category><![CDATA[Amazon Bedrock]]></category>
- <category><![CDATA[Announcements]]></category>
- <guid isPermaLink="false">6f3957c3db7756d44ae371ed56bdd96fe96046a0</guid>
+ <category><![CDATA[Amazon Bedrock AgentCore]]></category>
+ <category><![CDATA[Technical How-to]]></category>
+ <guid isPermaLink="false">923519c839c137326cec150f20986e906ba708bc</guid>
- <description>GLM 5.3 from Z.ai is now available on Amazon Bedrock: a 753B-parameter mixture-of-experts model built for coding and long-horizon agentic tasks. Learn how to invoke it with the OpenAI-compatible APIs, cut cost and latency with prompt caching, and run an authorized security test wit…
- <content:encoded>&lt;p&gt;Coding and agentic workloads are asking more of AI models than ever: refactor a repository spanning hundreds of files, sustain a multi-hour agentic workflow without losing context, and reason through complex systems problems with tool use at every step. Meeting th…
-&lt;p&gt;&lt;a href="https://z.ai/blog/glm-5.3" target="_blank" rel="noopener"&gt;GLM 5.3 from Z.ai&lt;/a&gt; (Zhipu AI) is now available on &lt;a href="https://aws.amazon.com/bedrock/" target="_blank" rel="noopener"&gt;Amazon Bedrock&lt;/a&gt;. GLM 5.3, as published &lt;a href="https://huggingface.…
-&lt;p&gt;In this post, we show you how to invoke GLM 5.3 on Amazon Bedrock using the OpenAI-compatible APIs and reduce cost and latency with prompt caching. We then put the model to work in a realistic agentic workflow: running an authorized security test of your own application with Strix, an open-…
-&lt;h2 id="whats-new-compared-to-glm-5"&gt;What’s new compared to GLM 5&lt;/h2&gt;
-&lt;p&gt;GLM 5 arrived on Amazon Bedrock earlier this year. GLM 5.3 builds on the same lineage, with a range of important gains:&lt;/p&gt;
+ <description>Off-the-shelf AI assistants forget you between conversations. This post shows how to build a personal assistant that accumulates context using OpenClaw on Amazon Bedrock AgentCore runtime, with AgentCore memory turning disposable chats into durable, structured knowledge you can ret…
+ <content:encoded>&lt;p&gt;Off-the-shelf AI assistants answer individual questions well, but they fall short on a different axis: continuity. Ask a stateless assistant about your garden today and it has no idea that you mentioned your fast-draining raised beds three weeks ago, that you only…
+&lt;p&gt;The problem isn’t the quality of the answers, but that the assistant has no memory of you. This post shows how to build a personal assistant that accumulates context using &lt;a href="https://openclaw.ai/" target="_blank" rel="noopener"&gt;OpenClaw&lt;/a&gt;, an open source agentic system, …
+&lt;p&gt;Our running example is Sprout, a gardening assistant, but the architecture is domain-agnostic. Swap the persona and the skills manifest, and the same pipeline serves a support bot, a fitness coach, or an internal help desk. The entire system lives in a single AWS CloudFormation template, de…
+&lt;h2 id="solution-overview"&gt;Solution overview&lt;/h2&gt;
+&lt;p&gt;AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. The following diagram shows the end-to-end request flow, from an inbound Telegram webhook through the AgentCore runtime, and its supporting AWS services.&lt;/p&gt;
+&lt;div id="attachment_140908" style="width: 4750px" class="wp-caption alignnone"&gt;
+ &lt;img aria-describedby="caption-attachment-140908" class="size-full wp-image-140908" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/06/ml21227_.drawio.png" alt="" width="4740" height="2270"&gt;
+ &lt;p id="caption-attachment-140908" class="wp-caption-text"&gt;Figure 1: Telegram webhooks and Amazon EventBridge schedules both invoke the same AgentCore runtime agent, which coordinates the OpenClaw gateway, AgentCore memory, and Amazon Bedrock&lt;/p&gt;
+&lt;/div&gt;
+&lt;p&gt;Two entry points converge on one agent. Telegram messages arrive through Amazon API Gateway and a webhook AWS Lambda function, while scheduled jobs such as morning watering reminders arrive through Amazon EventBridge Scheduler and a cronjob Lambda function. Both call the &lt;code&gt;InvokeA…
+&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
+&lt;p&gt;To deploy your own version using the Launch Stack button or scripts/&lt;code&gt;deploy.sh&lt;/code&gt; (described in the Grow your own section), you will need:&lt;/p&gt;
&lt;ul&gt;
- &lt;li&gt;&lt;strong&gt;Stronger coding:&lt;/strong&gt; &lt;a href="https://z.ai/blog/glm-5.3" target="_blank" rel="noopener"&gt;Z.ai claims&lt;/a&gt; competitive performance on a range of coding benchmarks including DeepSWE, Terminal Bench 3.0, and FrontierSWE. They also report a 50% improvement o…
- &lt;li&gt;&lt;strong&gt;Emergent cyber security capabilities:&lt;/strong&gt; Reported benchmark performance on security tasks stands out, which makes the model a natural fit for defensive security workflows. For example, Z.ai &lt;a href="https://z.ai/blog/glm-5.3" target="_blank" rel="noopener"&gt;…
- &lt;li&gt;&lt;strong&gt;Broader Amazon Bedrock integration:&lt;/strong&gt; Cross-Region inference profiles, implicit and explicit prompt caching, and improved feature parity of the OpenAI-compatible Responses and Chat Completions APIs alongside Invoke and Converse.&lt;/li&gt;
+ &lt;li&gt;Amazon Bedrock AgentCore access, including AgentCore runtime and AgentCore memory.&lt;/li&gt;
+ &lt;li&gt;Model access granted for the models you plan to route to: Claude Haiku 4.5 for text and Claude Sonnet 4.5 for vision (or the equivalents available in your account).&lt;/li&gt;
+ &lt;li&gt;Docker with &lt;code&gt;linux/arm64&lt;/code&gt; build support, plus the AWS Command Line Interface (AWS CLI) configured. This is needed only if you plan to build and push your own image.&lt;/li&gt;
+ &lt;li&gt;A Telegram bot token (from BotFather) to serve as the assistant’s front door.&lt;/li&gt;
+ &lt;li&gt;Basic familiarity with agent orchestration concepts and CloudFormation.&lt;/li&gt;
&lt;/ul&gt;
-&lt;h2 id="key-capabilities"&gt;Key capabilities&lt;/h2&gt;
+&lt;h2 id="the-architecture-a-serverless-agent-on-agentcore-runtime"&gt;The architecture: A serverless agent on AgentCore runtime&lt;/h2&gt;
+&lt;p&gt;Every component lives in a single CloudFormation template, and no build tooling is required to launch. The following sections walk through the load-bearing decisions.&lt;/p&gt;
+&lt;h3 id="agentcore-runtime-pay-only-for-active-compute"&gt;AgentCore runtime: Pay only for active compute&lt;/h3&gt;
+&lt;p&gt;The agent lives in a container on AgentCore runtime, which uses consumption-based pricing. You’re billed for the compute your agent actively consumes, not for wall-clock uptime, and you don’t pay for the time when waiting for I/O such as model response. For a personal assistant used in shor…
+&lt;p&gt;The runtime enforces a minimal container contract: listen on port 8080, and expose &lt;code&gt;GET /ping&lt;/code&gt; for health and &lt;code&gt;POST /invocations&lt;/code&gt; as the agent entry point. Our container is &lt;code&gt;linux/arm64&lt;/code&gt;, built multi-stage from the officia…
+&lt;h3 id="openclaw-as-the-agent-substrate"&gt;OpenClaw as the agent substrate&lt;/h3&gt;
+&lt;p&gt;OpenClaw provides the agent loop, tool use, and a skills system. It runs a wrapper (&lt;code&gt;server.py&lt;/code&gt;) that adapts it to &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/runtime-http-protocol-contract.html" target="_blank" rel="noopener"&gt;AgentCor…
&lt;ul&gt;
- &lt;li&gt;&lt;strong&gt;Frontier coding and agentic performance.&lt;/strong&gt; GLM 5.3 is designed for complex systems engineering and long-horizon agentic tasks. These include multi-step reasoning, tool-augmented workflows, and sustained context across large code bases.&lt;/li&gt;
- &lt;li&gt;&lt;strong&gt;Flexible API access.&lt;/strong&gt; You can invoke GLM 5.3 through the OpenAI-compatible Responses and Chat Completions APIs, or the Amazon Bedrock Invoke and Converse APIs.&lt;/li&gt;
- &lt;li&gt;&lt;strong&gt;Prompt caching.&lt;/strong&gt; GLM 5.3 supports implicit (automatic) prompt caching by default, and explicit cache controls (recommended) on the Responses and Chat Completions APIs. For agentic workloads that resend large system prompts or repository context every turn, cach…
- &lt;li&gt;&lt;strong&gt;Cross-Region inference.&lt;/strong&gt; GLM 5.3 is available through US cross-Region inference (&lt;code&gt;us.zai.glm-5.3&lt;/code&gt;) and Global cross-Region inference (&lt;code&gt;global.zai.glm-5.3&lt;/code&gt;) profiles. You send requests to the “source” AWS Region of y…
- &lt;li&gt;&lt;strong&gt;Service tiers.&lt;/strong&gt; Choose Flex to optimize cost for less-time-sensitive workloads, Priority to prioritize latency-critical requests in return for a higher price, or Standard for the default balance between price and speed.&lt;/li&gt;
+ &lt;li&gt;On container start, &lt;code&gt;server.py&lt;/code&gt; launches &lt;code&gt;openclaw gateway run&lt;/code&gt; as a subprocess and health-checks it.&lt;/li&gt;
+ &lt;li&gt;&lt;code&gt;GET /ping&lt;/code&gt; returns healthy quickly, so the AgentCore readiness probe passes.&lt;/li&gt;
+ &lt;li&gt;&lt;code&gt;POST /invocations&lt;/code&gt; does the real work: parse the payload, retrieve memory, assemble context, forward the turn to the gateway, and persist the result. One callout: AgentCore can thaw a frozen container whose subprocess has exited. So invocation path doesn’t assume t…
+&lt;/ul&gt;
+&lt;p&gt;This wrapper pattern generalizes to other use cases. Any agent framework that runs as a local process can be adapted to the AgentCore runtime the same way, without modifying the framework itself.&lt;/p&gt;
+&lt;h3 id="two-models-routed-by-task"&gt;Two models, routed by task&lt;/h3&gt;
+&lt;p&gt;Text chat and image understanding have different cost and quality tradeoffs, so the assistant routes them to different Claude models on &lt;a href="https://aws.amazon.com/bedrock/" target="_blank" rel="noopener"&gt;Bedrock&lt;/a&gt;:&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;strong&gt;Claude Haiku 4.5 for text:&lt;/strong&gt; Fast and cheap for the high-volume conversational turns that dominate daily use.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Claude Sonnet 4.5 for vision:&lt;/strong&gt; Stronger multimodal reasoning for the less frequent but harder task of diagnosing a plant from a photo.&lt;/li&gt;
+&lt;/ul&gt;
+&lt;p&gt;Text turns flow through the OpenClaw gateway, which brings skills and session state. Image turns call the large language model (LLM) from Bedrock directly from &lt;code&gt;server.py&lt;/code&gt;, passing the image bytes as multimodal content blocks. We route images around the gateway delibe…
+&lt;p&gt;The model IDs are environment variables (&lt;code&gt;MODEL_ID&lt;/code&gt;, &lt;code&gt;VISION_MODEL_ID&lt;/code&gt;), so you can swap models per deployment without rebuilding the image.&lt;/p&gt;
+&lt;h3 id="skills-as-the-reusable-capability-unit"&gt;Skills as the reusable capability unit&lt;/h3&gt;
+&lt;p&gt;Capabilities are declared as skills in a &lt;code&gt;community-skills.json&lt;/code&gt; manifest. A deploy-time script materializes them into the container and registers them in the OpenClaw config before the image is built. Sprout ships with weather, reminders, and plant notes skills at th…
+&lt;h3 id="telegram-as-the-serverless-front-door"&gt;Telegram as the serverless front door&lt;/h3&gt;
+&lt;p&gt;Telegram is a practical channel for a personal assistant since it’s webhook-based, and it keeps everything serverless. It requires no client development, works on every device the user already owns, and supports text, images, and rich formatting through a straightforward bot API. BotFather …
+&lt;p&gt;One formatting lesson to note: Telegram’s legacy markdown model is unforgiving about unescaped characters and a single stray underscore in a model response can make the whole message fail to send. Rendering replies as HTML is reliable so the assistant converts model output to Telegram-safe …
+&lt;h2 id="memory-turning-disposable-chats-into-durable-knowledge"&gt;Memory: Turning disposable chats into durable knowledge&lt;/h2&gt;
+&lt;p&gt;The architecture described so far is a capable, cheap, serverless agent, but on its own it still forgets you between conversations. Memory is what changes that. Imagine mentioning weeks ago that you garden organically, and today the assistant recommends a treatment and adds, on its own, tha…
+&lt;h3 id="the-mental-model-short-term-events-long-term-extraction"&gt;The mental model: Short-term events, long-term extraction&lt;/h3&gt;
+&lt;p&gt;AgentCore memory has two layers. &lt;strong&gt;Short-term memory&lt;/strong&gt; stores every conversation turn as an event through &lt;code&gt;CreateEvent&lt;/code&gt;, keyed by &lt;code&gt;actorId&lt;/code&gt; (the Telegram chat ID) and &lt;code&gt;sessionId&lt;/code&gt;. This is the raw t…
+&lt;ul&gt;
+ &lt;li&gt;&lt;code&gt;USER_PREFERENCE&lt;/code&gt;: explicit choices the gardener stated (“I only use organic fertilizer”).&lt;/li&gt;
+ &lt;li&gt;&lt;code&gt;SEMANTIC&lt;/code&gt;: inferred facts (“grows Mexican petunias in a Corten steel raised bed”).&lt;/li&gt;
+ &lt;li&gt;&lt;code&gt;SUMMARIZATION&lt;/code&gt;: episodic session summaries (“discussed yellowing lower leaves during a heat wave”).&lt;/li&gt;
+&lt;/ul&gt;
+&lt;h3 id="namespaces-one-garden-per-gardener"&gt;Namespaces: One garden per gardener&lt;/h3&gt;
+&lt;p&gt;Sprout files records into per-user namespaces, so no two chats ever mix:&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;code&gt;sprout/{chat_id}/long_term&lt;/code&gt;: preferences and semantic facts.&lt;/li&gt;
+ &lt;li&gt;&lt;code&gt;sprout/{chat_id}/episodic/{session_id}&lt;/code&gt;: session summaries.&lt;/li&gt;
+&lt;/ul&gt;
+&lt;p&gt;The chat ID is the only variable segment, which makes isolation straightforward to reason about and to test: each unique gardener maps to exactly one namespace, and no two gardeners collide.&lt;/p&gt;
+&lt;h3 id="the-retrieval-assembly-and-injection-pipeline"&gt;The retrieval, assembly, and injection pipeline&lt;/h3&gt;
+&lt;p&gt;On every turn, the agent retrieves the relevant long-term records, ranks them, and injects them into the system prompt. Here is what happens on every single message, inside &lt;code&gt;server.py&lt;/code&gt;:&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;strong&gt;Retrieve.&lt;/strong&gt; Call &lt;code&gt;RetrieveMemoryRecords&lt;/code&gt; against &lt;code&gt;sprout/{chat_id}/long_term&lt;/code&gt;, using the user’s message as the search query, capped at 50 results, under a 3-second budget. If retrieval times out or errors, we degrade…
&lt;/ul&gt;
-&lt;h2 id="prerequisites"&gt;Prerequisites&lt;/h2&gt;
-&lt;p&gt;For the following usage examples, you need:&lt;/p&gt;
-&lt;ol type="1"&gt;
- &lt;li&gt;An AWS account with access to Amazon Bedrock.&lt;/li&gt;
- &lt;li&gt;AWS Identity and Access Management (IAM) &lt;a href="https://docs.aws.amazon.com/service-authorization/latest/reference/list_bedrock.html" target="_blank" rel="noopener"&gt;permissions&lt;/a&gt; to call the base model and the target inference profile: &lt;code&gt;bedrock:InvokeModel&lt;/c…
- &lt;li&gt;(For the code-based demos) Python 3.10 or later.&lt;/li&gt;
- &lt;li&gt;(For the optional security-testing demo only) install Docker and &lt;a href="https://docs.strix.ai/llm-providers/bedrock" target="_blank" rel="noopener"&gt;Strix with the bedrock extra&lt;/a&gt;.&lt;/li&gt;
-&lt;/ol&gt;
-&lt;h2 id="try-glm-5.3-on-the-amazon-bedrock-console"&gt;Try GLM 5.3 on the Amazon Bedrock console&lt;/h2&gt;
-&lt;p&gt;You can start sending prompts to GLM 5.3 on the AWS Management Console, with no need to write code or install developer tools. To get started, navigate to &lt;a href="https://console.aws.amazon.com/bedrock/" target="_blank" rel="noopener"&gt;Amazon Bedrock&lt;/a&gt; and then choose &lt;stro…
-&lt;p&gt;From this playground interface you can select GLM 5.3 from the model list and send your first prompts through the chat UI, as shown in the following screenshot:&lt;/p&gt;
-&lt;div id="attachment_140841" style="width: 2090px" class="wp-caption alignnone"&gt;
- &lt;img aria-describedby="caption-attachment-140841" class="wp-image-140841 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/Screenshot-2026-10-05-at-6.11.23 PM.png" alt="[Amazon Bedrock Playground screenshot showing chat interface with GLM 5…
- &lt;p id="caption-attachment-140841" class="wp-caption-text"&gt;Figure 1: Chatting with GLM 5.3 on the Amazon Bedrock console&lt;/p&gt;
-&lt;/div&gt;
-&lt;h2 id="get-started-with-the-responses-api"&gt;Get started with the Responses API&lt;/h2&gt;
-&lt;p&gt;Programmatically, you can call the model through the &lt;code&gt;bedrock-runtime&lt;/code&gt; endpoint. This supports both the OpenAI-compatible &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/inference-responses-api.html" target="_blank" rel="noopener"&gt;Responses&lt;/a&g…
-&lt;p&gt;Amazon Bedrock does support &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/api-keys-generate.html" target="_blank" rel="noopener"&gt;generating API keys&lt;/a&gt; for OpenAI-compatible integrations that require them. However, we strongly recommend preferring short-lived cr…
-&lt;p&gt;In the following example, we will call the Responses API from Python using the OpenAI Python SDK, and the &lt;a href="https://pypi.org/project/aws-bedrock-token-generator/" target="_blank" rel="noopener"&gt;aws-bedrock-token-generator&lt;/a&gt; library to generate short-term tokens from you…
-&lt;ol type="1"&gt;
- &lt;li&gt;Install the required packages.
- &lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-bash"&gt;pip install -U openai aws-bedrock-token-generator&lt;/code&gt;&lt;/pre&gt;
- &lt;/div&gt;&lt;/li&gt;
- &lt;li&gt;Save the following code as &lt;code&gt;bedrock-request.py&lt;/code&gt;.
- &lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-python"&gt;from aws_bedrock_token_generator import provide_token
-from openai import OpenAI
-
-region = "us-west-2" # Your source AWS Region
-
-client = OpenAI(
- api_key=provide_token(region=region),
- base_url=f"https://bedrock-runtime.{region}.amazonaws.com/openai/v1",
-)
-
-resp = client.responses.create(
- input="Refactor this Python function to be iterative instead of recursive: ...",
- model="global.zai.glm-5.3",
-)
-
-print(resp.output_text)&lt;/code&gt;&lt;/pre&gt;
- &lt;/div&gt;&lt;/li&gt;
- &lt;li&gt;Run the script, which will display the model’s output.
- &lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-bash"&gt;python bedrock-request.py&lt;/code&gt;&lt;/pre&gt;
- &lt;/div&gt;&lt;/li&gt;
-&lt;/ol&gt;
-&lt;h3 id="optimize-inference-with-explicit-prompt-caching"&gt;Optimize inference with explicit prompt caching&lt;/h3&gt;
-&lt;p&gt;Long-running coding and knowledge workflows often resend stable context across multiple conversation turns, such as system prompts, tool definitions, or repository files.&lt;/p&gt;
-&lt;p&gt;GLM 5.3 on Amazon Bedrock supports implicit prompt caching by default, which helps reduce response latency and input token costs for repeated calls sharing the same initial prompt prefix.&lt;/p&gt;
-&lt;p&gt;With &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html" target="_blank" rel="noopener"&gt;explicit prompt caching&lt;/a&gt; mode you specifically identify the reusable prompt prefixes, which can further improve cache hit rate (and therefore latency and cos…
-&lt;p&gt;To use explicit prompt caching with GLM 5.3, as shown in the following example:&lt;/p&gt;
-&lt;ol type="1"&gt;
- &lt;li&gt;Select the explicit caching mode through &lt;code&gt;prompt_cache_options&lt;/code&gt; on your request.&lt;/li&gt;
- &lt;li&gt;Add one or more &lt;code&gt;prompt_cache_breakpoint&lt;/code&gt; markers on input content blocks to indicate the end (inclusive) of reusable prompt prefixes. Each breakpoint must contain at least 1,024 tokens to be eligible for caching.&lt;/li&gt;
-&lt;/ol&gt;
&lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-python"&gt;resp = client.responses.create(
- model="global.zai.glm-5.3",
- # Enable explicit caching mode:
- extra_body={"prompt_cache_options": {"mode": "explicit"}},
- input=[
- {
- "type": "message",
- "role": "system",
- "content": [
- {
- "type": "input_text",
- "text": SYSTEM_PROMPT,
- # A long, static system prompt is a great target for caching:
- "prompt_cache_breakpoint": {"mode": "explicit"},
- },
- ]
- },
- {
- "type": "message",
- "role": "user",
- "content": [
- {
- "type": "input_text",
- "text": USER_INPUT,
- # Multiple breakpoints can also be defined, for layered cache:
- "prompt_cache_breakpoint": {"mode": "explicit"},
- },
- ],
+ &lt;pre&gt;&lt;code class="language-python"&gt;try:
+ records = memory_client.retrieve_memory_records(
+ memoryId=MEMORY_ID,
+ namespace=f'sprout/{chat_id}/long_term',
+ searchCriteria={
+ 'searchQuery': user_message,
+ 'topK': 50,
+ 'metadataFilters': []
},
+ ) # 3s timeout
+except Exception:
+ records = [] # fall back to answering without memory&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;&lt;em&gt;Snippet 1: Retrieving long-term records for the current turn (representative. See the repo for full source).&lt;/em&gt;&lt;/p&gt;
+&lt;p&gt;Assemble function adds additional custom logic. We want the explicit preferences to rank ahead of inferred facts, order is stable within each class, and the result is capped before injection:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;def assemble(records, cap=50):
+ explicit = [r for r in records if r.type == 'USER_PREFERENCE']
+ inferred = [r for r in records if r.type != 'USER_PREFERENCE']
+ # explicit beats inferred; stable order within each class
+ ordered = explicit + inferred
+ return ordered[:cap]&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;&lt;em&gt;Snippet 2: The assembly step ranks explicit preferences before inferred facts.&lt;/em&gt;&lt;/p&gt;
+&lt;h3 id="metadata-subgrouping-memories-inside-a-namespace"&gt;Metadata: Subgrouping memories inside a namespace&lt;/h3&gt;
+&lt;p&gt;Namespaces answer whose memory a record is, but metadata answers what it’s about. Inside &lt;code&gt;sprout/{chat_id}/long_term&lt;/code&gt;, a semantic search for “my petunias are wilting”, would return everything that is close in meaning. For a gardener, that means a fertilizer preference…
+&lt;p&gt;One rule shapes every decision here. A metadata key is only filterable server-side if you declare it as an indexed key. You can read more in &lt;a href="https://aws.amazon.com/blogs/machine-learning/structured-memory-filtering-with-metadata-in-agentcore-memory/" target="_blank" rel="noopene…
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-yaml"&gt;IndexedKeys: # on the AWS::BedrockAgentCore::Memory resource
+ - Key: type # seperate the kinds of records
+ Type: STRING
+ - Key: section # which bed or area it describes
+ Type: STRING
+ - Key: plants # what is growing there
+ Type: STRINGLIST&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;Each entry names a key, which must match an indexed key to be filterable, and sets &lt;code&gt;extractionType&lt;/code&gt; to either &lt;code&gt;STRICTLY_CONSISTENT&lt;/code&gt;, passed through from the event, or &lt;code&gt;LLM_INFERRED&lt;/code&gt;, extracted from the conversation. For in…
+&lt;h3 id="persisting-the-turn-and-closing-the-loop"&gt;Persisting the turn and closing the loop&lt;/h3&gt;
+&lt;p&gt;After the model responds, &lt;code&gt;server.py&lt;/code&gt; calls &lt;code&gt;CreateEvent&lt;/code&gt; with both the user turn and the assistant turn. That new event feeds the extraction strategies, which enrich the long-term store for next time.&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;memory.create_event(
+ memoryId=MEMORY_ID,
+ actorId=chat_id,
+ sessionId=session_id,
+ payload=[
+ {'role': 'user', 'content': user_message},
+ {'role': 'assistant', 'content': reply},
],
-)
-
-if resp.usage.input_tokens_details.cached_tokens:
- print("Hit cache!")&lt;/code&gt;&lt;/pre&gt;
+) # feeds USER_PREFERENCE / SEMANTIC / SUMMARIZATION extraction; errors are logged, never fatal&lt;/code&gt;&lt;/pre&gt;
&lt;/div&gt;
-&lt;p&gt;For more information, refer to the &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html#prompt-caching-openai" target="_blank" rel="noopener"&gt;prompt caching section&lt;/a&gt; of the Amazon Bedrock User Guide.&lt;/p&gt;
-&lt;h2 id="example-agentic-workload-authorized-security-testing-with-strix"&gt;Example agentic workload: Authorized security testing with Strix&lt;/h2&gt;
-&lt;p&gt;One workload that benefits directly from GLM 5.3’s strengths is automated security testing of your own applications. &lt;a href="https://github.com/usestrix/strix" target="_blank" rel="noopener"&gt;Strix&lt;/a&gt; is an open-source AI penetration testing agent that runs your code dynamicall…
-&lt;p&gt;&lt;strong&gt;Only test applications you own or have explicit written permission to test.&lt;/strong&gt; Unauthorized security testing of systems you don’t own is illegal in most jurisdictions and violates the AWS Acceptable Use Policy. In this walkthrough, the target is &lt;a href="https:/…
-&lt;p&gt;If you want fully managed, continuous security testing beyond running open-source agents yourself, &lt;a href="https://aws.amazon.com/security-agent/" target="_blank" rel="noopener"&gt;AWS Continuum&lt;/a&gt; provides on-demand penetration testing and other security analyses as a managed se…
-&lt;h3 id="to-run-an-authorized-security-test"&gt;To run an authorized security test&lt;/h3&gt;
-&lt;ol type="1"&gt;
- &lt;li&gt;Start the example Juice Shop target application locally.
- &lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-bash"&gt;docker run --rm -p 3000:3000 bkimminich/juice-shop&lt;/code&gt;&lt;/pre&gt;
- &lt;/div&gt;&lt;/li&gt;
- &lt;li&gt;Configure Strix to use GLM 5.3 on Amazon Bedrock. Strix uses &lt;a href="https://docs.litellm.ai/docs/providers/bedrock" target="_blank" rel="noopener"&gt;LiteLLM&lt;/a&gt; under the hood so (as described in &lt;a href="https://docs.strix.ai/llm-providers/bedrock" target="_blank" rel="noo…
- &lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-bash"&gt;# Fill in the REGION and ACCOUNT_ID placeholders below before running!
-export STRIX_LLM="bedrock/converse/arn:aws:bedrock:{AWS_REGION}:{AWS_ACCOUNT_ID}:inference-profile/global.zai.glm-5.3"&lt;/code&gt;&lt;/pre&gt;
- &lt;/div&gt;&lt;/li&gt;
- &lt;li&gt;Run Strix against the local target.
- &lt;div class="hide-language"&gt;
- &lt;pre&gt;&lt;code class="language-bash"&gt;strix --target http://localhost:3000&lt;/code&gt;&lt;/pre&gt;
- &lt;/div&gt;&lt;/li&gt;
- &lt;li&gt;Wait for the root Strix agent to complete, then review the findings.&lt;/li&gt;
-&lt;/ol&gt;
-&lt;p&gt;Strix spins up a team of sub-agents to map the threat surface, explore a range of potential vulnerability categories, and attempt to validate each finding with a working proof of concept. This helps minimize time spent triaging false positives. A successful run will generate a report includ…
-&lt;p&gt;The following video shows the end-to-end journey of setting up and running Strix against the example application, and exploring the results:&lt;/p&gt;
-&lt;div style="width: 640px;" class="wp-video"&gt;
- &lt;video class="wp-video-shortcode" id="video-140838-1" width="640" height="360" preload="metadata" controls="controls"&gt;&lt;source type="video/mp4" src="https://d2908q01vomqb2.cloudfront.net/artifacts/DBSBlogs/ML-22047/Strix+Demo+Video+Censored.mp4?_=1"&gt;&lt;/video&gt;
+&lt;p&gt;&lt;em&gt;Snippet 3: Persisting the turn so the extraction strategies can enrich long-term memory asynchronously.&lt;/em&gt;&lt;/p&gt;
+&lt;p&gt;Extraction is asynchronous, so a fact mentioned in this session typically becomes retrievable in a later one. Design for that delay: short-term session events cover the current conversation, and long-term records cover everything before it.&lt;/p&gt;
+&lt;h3 id="putting-it-together-a-personalized-watering-plan"&gt;Putting it together: A personalized watering plan&lt;/h3&gt;
+&lt;p&gt;Here is where the full pipeline works end-to-end. Over a few conversations you catalog your whole garden, one plant at a time, in plain language. Each mention becomes an event. The extraction strategies extract information about the plant, its location, and its sun exposure into &lt;code&gt…
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-2.jpeg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-2.jpeg" alt="Telegram chat where S…
+ &lt;p class="wp-caption-text"&gt;Figure 2: Sprout answers a question about the garden by recalling the stored plant inventory and growing conditions&lt;/p&gt;
&lt;/div&gt;
-&lt;p&gt;Figure 2: Running an example security test with GLM 5.3 and Strix&lt;/p&gt;
-&lt;blockquote&gt;
- &lt;p style="text-align: left"&gt;&lt;/p&gt;
-&lt;/blockquote&gt;
+&lt;p&gt;Using the scheduler skill on the Amazon EventBridge → Cron path, Sprout can also turn that plan into proactive reminders (“skip the herbs, the soil is still damp from yesterday”) and adjusts them against the weather skill when rain or a heat wave is coming.&lt;/p&gt;
+&lt;p&gt;Memory and vision also compound each other. When the user sends a photo of a wilting plant, the image goes to Claude Sonnet 4.5 while the system prompt still carries everything the memory layer knows. The assistant matches the photo to the Mexican petunias already in the user’s saved invent…
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-3.png" alt="Telegram chat where Spr…
+ &lt;p class="wp-caption-text"&gt;Figure 3: Vision and memory working together. The photo goes to the vision model while the system prompt carries the user’s stored garden context&lt;/p&gt;
+&lt;/div&gt;
+&lt;p&gt;Vision models aren’t infallible. In an earlier exchange without the inventory context, the same plant was confidently identified as a morning glory, a species with similar trumpet-shaped purple flowers. Grounding the vision model with the user’s own stored inventory is what turned a plausib…
+&lt;h3 id="keeping-inference-costs-low-with-prompt-caching"&gt;Keeping inference costs low with prompt caching&lt;/h3&gt;
+&lt;p&gt;Injecting memory into every turn makes the system prompt large, and a naive implementation would pay for those tokens on every request. Prompt caching on Amazon Bedrock addresses this. The assistant structures its prompt so that the stable prefix, the persona and the assembled memory block,…
+&lt;p&gt;The ordering rule matters more than any single setting: put stable content first, volatile content last, and keep the memory block’s internal ordering deterministic (which the preceding assembly function facilitates) so the prefix actually matches between requests.&lt;/p&gt;
+&lt;h2 id="design-guidelines-to-build-on-agentcore-and-openclaw"&gt;Design guidelines to build on AgentCore and OpenClaw&lt;/h2&gt;
+&lt;p&gt;Sprout is one assistant, but the decisions behind it generalize. If you’re building your own assistant on this stack, the following guidelines are the ones we would carry to any domain.&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;strong&gt;Wrap, don’t fork.&lt;/strong&gt; Adapt your agent framework to the AgentCore container contract with a thin HTTP wrapper rather than modifying the framework. The contract is small, port 8080 with &lt;code&gt;/ping&lt;/code&gt; and &lt;code&gt;/invocations&lt;/code&gt;, and a…
+ &lt;li&gt;&lt;strong&gt;Design namespaces before you store anything.&lt;/strong&gt; Memory namespaces are your isolation boundary. Make the user ID the only variable segment, and choose it from a channel-native ID you already trust, such as the chat ID. Multi-tenant designs get audits and deletion …
+ &lt;li&gt;&lt;strong&gt;Treat memory as an enhancement, never a dependency.&lt;/strong&gt; Every memory operation should be allowed to fail gracefully. Retrieval failures should produce a memoryless answer without blocking the reply. Users forgive a forgetful turn far more readily than a failed one…
+ &lt;li&gt;&lt;strong&gt;Route models by task.&lt;/strong&gt; Use a fast, cost-effective model for high-volume text and reserve a stronger multimodal model for the turns that need it. Keep model IDs in environment variables so routing changes are configuration, not code.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Order prompts for the cache.&lt;/strong&gt; Stable persona and memory first, volatile user input last, deterministic ordering throughout. This one structural habit is where most of the inference savings come from.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Plan for extraction latency.&lt;/strong&gt; Long-term memory is extracted asynchronously, so don’t promise same-session recall of new facts. Let short-term session events cover the current conversation and long-term records cover prior ones.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Put a budget on it from day one.&lt;/strong&gt; A consumption-based agent is inexpensive until a retry loop or a chatty user makes it otherwise. An AWS Budgets alert at 80 percent and 100 percent of a monthly cap costs nothing and catches surprises early.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Keep skills small and single-purpose.&lt;/strong&gt; A skill should do one thing a user would name in a sentence, such as check the weather or set a reminder. Small skills are independently testable, independently swappable, and easy for the model to select correctly. A do-e…
+&lt;/ul&gt;
+&lt;h2 id="grow-your-own"&gt;Grow your own&lt;/h2&gt;
+&lt;p&gt;Two ways to plant it, same garden:&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;strong&gt;Single-step Launch Stack:&lt;/strong&gt; the CloudFormation template points at a public Amazon Elastic Container Registry (Amazon ECR) image, so it deploys nothing but a Telegram bot token.&lt;/li&gt;
+ &lt;li&gt;Build your own: The scripts/&lt;code&gt;deploy.sh&lt;/code&gt; script validates the template, builds and pushes your own ARM64 image to your private Amazon ECR repository, deploys the stack, and registers the Telegram webhook, for a fully customizable build.&lt;/li&gt;
+&lt;/ul&gt;
+&lt;p&gt;Light personal use runs about $5–9/month as of July 2026 (roughly $2 infrastructure, $1–3 Haiku text, $2 Sonnet vision), with a built-in AWS Budget that alerts at 80 percent and 100 percent of a cap you set.&lt;/p&gt;
+&lt;p&gt;The full source code is available in the &lt;a href="https://github.com/aws-samples/sample-agentcore-memory-openclaw" target="_blank" rel="noopener"&gt;sample-agentcore-memory-openclaw GitHub repository&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="clean-up"&gt;Clean up&lt;/h2&gt;
-&lt;p&gt;Stop the Juice Shop container with &lt;strong&gt;Ctrl+C&lt;/strong&gt; in the terminal where it’s running, or run &lt;code&gt;docker ps&lt;/code&gt; to find the container ID and stop it with &lt;code&gt;docker stop &amp;lt;container-id&amp;gt;&lt;/code&gt;. Amazon Bedrock inference is pay-p…
-&lt;h2 id="availability"&gt;Availability&lt;/h2&gt;
-&lt;p&gt;Give GLM 5.3 a try on the &lt;a href="https://console.aws.amazon.com/bedrock" target="_blank" rel="noopener"&gt;Amazon Bedrock console&lt;/a&gt;, use it through coding assistants like OpenCode as shown in our &lt;a href="https://aws.amazon.com/blogs/machine-learning/use-open-weight-models-a…
-&lt;p&gt;&lt;em&gt;Interested in how Amazon Bedrock can support your team?&lt;/em&gt; &lt;a href="https://pages.awscloud.com/Amazon-Bedrock-Contact-Us.html" target="_blank" rel="noopener"&gt;&lt;em&gt;Connect with us&lt;/em&gt;&lt;/a&gt; &lt;em&gt;to start the conversation.&lt;/em&gt;&lt;/p&gt;
+&lt;p&gt;When you are done experimenting, tear everything down to avoid ongoing charges. Because the whole system is one CloudFormation stack, cleanup is mostly a single delete:&lt;/p&gt;
+&lt;ol type="1"&gt;
+ &lt;li&gt;Delete the CloudFormation stack. This removes the AgentCore runtime agent, API Gateway, the Lambda functions, the Amazon EventBridge schedule, and the associated AWS Identity and Access Management (IAM) roles.&lt;/li&gt;
+ &lt;li&gt;Delete the AgentCore memory store (and its namespaces) so no user records are retained.&lt;/li&gt;
+ &lt;li&gt;Delete any images you pushed to your private ECR repository, and the repository itself if it’s no longer needed.&lt;/li&gt;
+ &lt;li&gt;Remove the AWS Budget alert if you created one outside the stack.&lt;/li&gt;
+ &lt;li&gt;Revoke Telegram’s webhook (or delete the bot through BotFather), and revoke Bedrock model access if you no longer need it.&lt;/li&gt;
+&lt;/ol&gt;
+&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
+&lt;p&gt;The reusable core of this solution is a serverless agent on Amazon Bedrock AgentCore with a skills system and managed memory. AgentCore memory removes the need to build custom vector stores and extraction pipelines while leaving you full control over what the agent remembers and forgets, co…
+&lt;p&gt;To go further, start with a single domain such as watering reminders and expand memory scope incrementally, explore episodic memory so the agent can reference specific past conversations (“last time we discussed the fig tree, you decided to hold off on fertilizer”), or fork the &lt;a href="…
+&lt;p&gt;To learn more, refer to the &lt;a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/what-is-bedrock-agentcore.html" target="_blank" rel="noopener"&gt;AgentCore documentation&lt;/a&gt;. The following related posts cover the building blocks in more depth:&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/machine-learning/amazon-bedrock-agentcore-memory-building-context-aware-agents/" target="_blank" rel="noopener"&gt;Amazon Bedrock AgentCore memory: Building context-aware agents&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/machine-learning/building-smarter-ai-agents-agentcore-long-term-memory-deep-dive/" target="_blank" rel="noopener"&gt;Building smarter AI agents: AgentCore long-term memory deep dive&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/machine-learning/effectively-use-prompt-caching-on-amazon-bedrock/" target="_blank" rel="noopener"&gt;Effectively use prompt caching on Amazon Bedrock&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;&lt;a href="https://aws.amazon.com/blogs/machine-learning/securely-launch-and-scale-your-agents-and-tools-on-amazon-bedrock-agentcore-runtime/" target="_blank" rel="noopener"&gt;Securely launch and scale your agents and tools on Amazon Bedrock AgentCore runtime&lt;/a&gt;&lt;/li&gt;
+&lt;/ul&gt;
+&lt;p style="clear: both"&gt;&lt;/p&gt;
&lt;hr style="width: 100%"&gt;
&lt;h2&gt;About the authors&lt;/h2&gt;
&lt;footer&gt;
&lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
&lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
- &lt;p&gt;&lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21636-3.jpg" alt="Alex Thewsey" width="100" height="133"&gt;&lt;/p&gt;
+ &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-4.jpg" alt="Thiago Verney" width="100" height="133"&gt;
&lt;/div&gt;
- &lt;h3 class="lb-h4"&gt;Alex Thewsey&lt;/h3&gt;
- &lt;p&gt;Alex is an AI Specialist Solutions Architect at AWS, based in Singapore. He focuses on how open source technologies and open weight models can help customers around the world to build innovative AI solutions and tackle AI governance challenges.&lt;/p&gt;
+ &lt;h3 class="lb-h4"&gt;Thiago Verney&lt;/h3&gt;
+ &lt;p style="overflow: hidden"&gt;Thiago is a Front-End Engineer on the One MHS team at Amazon, specializing in AI-powered interfaces using React, TypeScript, and modern federated microfrontend architecture. He builds user-focused interfaces for operations-leader in FC, drawing on prior work on th…
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-5.jpg" alt="Sathya Balakrishnan" width="100" height="133"&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Sathya Balakrishnan&lt;/h3&gt;
+ &lt;p style="overflow: hidden"&gt;Sathya is a Pr. Cloud Architect in the Professional Services team at Amazon Web Services (AWS), specializing in data and machine learning (ML) solutions. He works with US federal financial clients. He is passionate about building pragmatic solutions to solve custo…
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/ML-21277-6.jpg" alt="Akarsha Sehwag" width="100" height="133"&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Akarsha Sehwag&lt;/h3&gt;
+ &lt;p style="overflow: hidden"&gt;Akarsha is a Sr.&amp;nbsp;Gen AI Data Scientist for Amazon Bedrock AgentCore GTM team. With over 7 years of expertise in AI/ML, she has built production-ready enterprise solutions across diverse customer segments in Generative AI, Deep Learning and Computer Vision…
&lt;/div&gt;
&lt;/footer&gt;</content:encoded>
- <enclosure length="353017803" type="video/mp4" url="https://d2908q01vomqb2.cloudfront.net/artifacts/DBSBlogs/ML-22047/Strix+Demo+Video+Censored.mp4"/>
-
+
</item>
<item>
- <title>Supercharge regulated workloads with Claude Code and Amazon Bedrock</title>
- <link>https://aws.amazon.com/blogs/machine-learning/supercharge-regulated-workloads-with-claude-code-and-amazon-bedrock/</link>
+ <title>Responsible AI governance: How AWS positions customers to align with ISO/IEC 42005:2025</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/responsible-ai-governance-how-aws-positions-customers-to-align-with-iso-iec-420052025/</link>
- <dc:creator><![CDATA[Bradley Wyman]]></dc:creator>
- <pubDate>Mon, 05 Oct 2026 17:25:20 +0000</pubDate>
- <category><![CDATA[Amazon Bedrock]]></category>
- <category><![CDATA[Intermediate (200)]]></category>
- <category><![CDATA[Technical How-to]]></category>
- <guid isPermaLink="false">0cbe16ca98bef589de7fcf6ba5ec623095bdb76d</guid>
-
- <description>Anthropic Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in the AWS GovCloud (US) Regions. Learn how to use them with Claude Code, Anthropic's agentic coding tool, for compliance-aligned, AI-assisted development on regulated and ITAR workloads.</description>
- <content:encoded>&lt;p&gt;&lt;em&gt;Please note that the following post is intended for informational purposes only. The approach detailed below may not be suitable for all organizations or compliance programs. It is important to evaluate this potential solution against the compliance requ…
-&lt;p&gt;The availability of Anthropic Claude Opus 5.5 and Claude Sonnet 5.5 in the AWS GovCloud (US) Regions introduces an on-ramp for AI-assisted development for workloads with regulatory or compliance requirements, including International Traffic in Arms Regulations (ITAR). Claude Sonnet 5 holds …
-&lt;p&gt;In this post, we explore how to use these models on &lt;a href="https://aws.amazon.com/bedrock" target="_blank" rel="noopener"&gt;Amazon Bedrock&lt;/a&gt; in AWS GovCloud (US) with &lt;a href="https://code.claude.com" target="_blank" rel="noopener"&gt;Claude Code&lt;/a&gt;, Anthropic’s agen…
-&lt;h2 id="amazon-bedrock-in-aws-govcloud-us"&gt;Amazon Bedrock in AWS GovCloud (US)&lt;/h2&gt;
-&lt;p&gt;&lt;a href="https://aws.amazon.com/govcloud-us/" target="_blank" rel="noopener"&gt;AWS GovCloud (US) Regions&lt;/a&gt; are designed specifically for US customers with elevated &lt;a href="https://docs.aws.amazon.com/govcloud-us/latest/UserGuide/govcloud-compliance.html" target="_blank" rel=…
-&lt;p&gt;&lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html" target="_blank" rel="noopener"&gt;Built-in data protection&lt;/a&gt; where customer content isn’t stored, logged, or used to train AWS models or shared with third parties.&lt;/p&gt;
-&lt;p&gt;FedRAMP Class D (formerly High) certification and DoD Cloud Service Provider (CSP) SRG IL4/IL5 authorization pathways, supporting government agencies’ compliance requirements. See the &lt;a href="https://aws.amazon.com/compliance/services-in-scope/FedRAMP/amazon-bedrock-models/" target="_bl…
-&lt;p&gt;Integration with existing security controls and compliance frameworks available in AWS GovCloud (US), maintaining the same high security standards as other AWS Regions while providing additional authorization pathways.&lt;/p&gt;

Diff display stops at 400 lines. The line counts above are from the whole diff. 61 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.