llm-catalog-archive

Change

fc524df

fc524dfa6a6ff0ba6b579e4510ea82fd069d5283 · commit on GitHub

aws-blog-feed: changed (671290 bytes, HTTP 200)

raw/aws-blog-feed/response.xml modified

Lines added
+4,169
Lines removed
-5,437
Stored bytes at this commit
671,290
Timestamp
observed
Raw artifact at this commit
raw/aws-blog-feed/response.xml
Recorded headers
observed_at2026-09-22T04:57:17.651Z
origin_datenull
status200
final URLhttps://aws.amazon.com/blogs/machine-learning/feed/
etagnull
last-modifiedTue, 22 Sep 2026 00:16:58 GMT
dateTue, 22 Sep 2026 04:57:17 GMT
agenull
cache-controlnull
cf-cache-statusnull
content-encodingnull
content-lengthnull
@@@ -5,7 +5,7 @@
<atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/>
<link>https://aws.amazon.com/blogs/machine-learning/</link>
<description>Official Machine Learning Blog of Amazon Web Services</description>
- <lastBuildDate>Fri, 18 Sep 2026 21:17:31 +0000</lastBuildDate>
+ <lastBuildDate>Mon, 21 Sep 2026 18:30:34 +0000</lastBuildDate>
<language>en-US</language>
<sy:updatePeriod>
hourly </sy:updatePeriod>
@@@ -13,6 +13,961 @@
1 </sy:updateFrequency>
<item>
+ <title>xAI’s Grok 4.6 is now available in Amazon Bedrock</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/xais-grok-4-6-is-now-available-in-amazon-bedrock/</link>
+
+ <dc:creator><![CDATA[Suheel Farooq]]></dc:creator>
+ <pubDate>Mon, 21 Sep 2026 18:30:34 +0000</pubDate>
+ <category><![CDATA[Amazon Bedrock]]></category>
+ <category><![CDATA[Amazon Machine Learning]]></category>
+ <category><![CDATA[Announcements]]></category>
+ <category><![CDATA[Artificial Intelligence]]></category>
+ <category><![CDATA[Foundation models]]></category>
+ <category><![CDATA[Generative AI]]></category>
+ <category><![CDATA[Intermediate (200)]]></category>
+ <category><![CDATA[Launch]]></category>
+ <guid isPermaLink="false">56dfeac3bfef33ff7b19f174cf3e455627c66990</guid>
+
+ <description>xAI's Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with Converse API and cross-
+ <content:encoded>&lt;p&gt;Today, we are announcing that xAI’s Grok 4.6 is available in &lt;a href="https://aws.amazon.com/bedrock/" target="_blank" rel="noopener"&gt;Amazon Bedrock&lt;/a&gt;, adding a frontier model built for long-running agents, coding, and knowledge work to the Bedrock m
+&lt;p&gt;This is xAI’s second model in Amazon Bedrock. When &lt;a href="https://aws.amazon.com/blogs/machine-learning/introducing-grok-on-amazon-bedrock/" target="_blank" rel="noopener"&gt;Grok 4.3 became generally available&lt;/a&gt;, xAI joined Amazon Bedrock as a model provider and the model was
+&lt;p&gt;This post covers what xAI says Grok 4.6 is designed for, how it is packaged on Amazon Bedrock, and how to send your first request.&lt;/p&gt;
+&lt;h2 id="what-grok-4.6-is-built-for"&gt;What Grok 4.6 is built for&lt;/h2&gt;
+&lt;p&gt;The capability and training details in this section come from xAI’s launch announcement, &lt;a href="https://x.ai/news/grok-4-6" target="_blank" rel="noopener"&gt;Introducing Grok 4.6&lt;/a&gt;.&lt;/p&gt;
+&lt;p&gt;Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. xAI describes the model as staying with complex tasks across many steps, whether that is researching a topic, analyzing information, working across a code base, or turn
+&lt;p&gt;On training, xAI reports a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. It then used Grok 4.5 to regenerate the supervised fine-
+&lt;p&gt;Two behaviors xAI calls out are worth noting for anyone building agents. On longer trajectories, the model began showing more self-testing and verification, checking its own work before moving on. It also produces stronger first passes on visual and interactive projects, establishing the st
+&lt;p&gt;On safety, xAI states that Grok 4.6’s safeguards have been improved and calibrated in line with the model’s capabilities, backed by what it describes as its widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, plus post-deployment and third-party testing.
+&lt;h2 id="reported-benchmark-results"&gt;Reported benchmark results&lt;/h2&gt;
+&lt;p&gt;xAI reports that Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. These are the figures it published for Grok 4.6 High at launch on August 12, 2026:&lt;/p&gt;
+&lt;table style="height: 490px" border="1px" width="856" cellpadding="10px"&gt;
+ &lt;tbody&gt;
+ &lt;tr&gt;
+ &lt;td&gt;&lt;strong&gt;Evaluation&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Grok 4.6 High&lt;/strong&gt;&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;AA Intelligence Index&lt;/td&gt;
+ &lt;td&gt;61&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;GDPVal-AA v2&lt;/td&gt;
+ &lt;td&gt;1753&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;CursorBench v3.2&lt;/td&gt;
+ &lt;td&gt;69.9%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;DeepSWE v1.1&lt;/td&gt;
+ &lt;td&gt;65.9%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;FrontierCode v1.1 (Extended)&lt;/td&gt;
+ &lt;td&gt;61.3%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;APEX-Agents&lt;/td&gt;
+ &lt;td&gt;57.5%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;Terminal-Bench v3.0&lt;/td&gt;
+ &lt;td&gt;26%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;APEX-SWE&lt;/td&gt;
+ &lt;td&gt;56.4%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;AA-Briefcase&lt;/td&gt;
+ &lt;td&gt;1577&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;Harvey LAB (Vals)&lt;/td&gt;
+ &lt;td&gt;15.8%&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;/tbody&gt;
+&lt;/table&gt;
+&lt;p&gt;Source: xAI, according to &lt;a class="uri" href="https://x.ai/news/grok-4-6" target="_blank" rel="noopener"&gt;https://x.ai/news/grok-4-6&lt;/a&gt;.&lt;/p&gt;
+&lt;p&gt;Several of those evaluations come from Artificial Analysis, so it helps to know what they measure. According to &lt;a href="https://artificialanalysis.ai/models" target="_blank" rel="noopener"&gt;Artificial Analysis&lt;/a&gt;, the Artificial Analysis Intelligence Index v4.1.1 is a composite
+&lt;p&gt;Artificial Analysis also tracks cost and latency alongside intelligence. Its cost-per-task metric is a weighted average cost per Intelligence Index task, derived from input, cache hit, cache write, reasoning, and answer token prices, which is a useful lens if you are sizing a reasoning-heav
+&lt;h2 id="what-grok-4.6-adds-on-bedrock"&gt;What Grok 4.6 adds on Bedrock&lt;/h2&gt;
+&lt;p&gt;Several Bedrock capabilities are new for this model rather than carried over from the earlier Grok launch.&lt;/p&gt;
+&lt;p&gt;&lt;strong&gt;The &lt;code&gt;bedrock-runtime&lt;/code&gt; endpoint.&lt;/strong&gt; Grok 4.6 is served on bedrock-runtime in addition to bedrock-mantle, so you can reach it with the AWS SDKs and the standard Bedrock control surface rather than only an OpenAI-compatible client.&lt;/p&gt;
+&lt;p&gt;&lt;strong&gt;The Converse API, including streaming.&lt;/strong&gt; Both &lt;code&gt;converse&lt;/code&gt; and &lt;code&gt;converse_stream&lt;/code&gt; are available. This is the practical payoff of runtime support: one message shape across models, and streaming through the usual Converse e
+&lt;p&gt;&lt;strong&gt;An &lt;code&gt;xhigh&lt;/code&gt; reasoning effort level.&lt;/strong&gt; Effort runs low, medium, high, xhigh, extending the range at the top end for problems where a deeper pass is worth the tokens. On Converse, set it through &lt;code&gt;additionalModelRequestFields={"reason
+&lt;p&gt;&lt;strong&gt;Cross-Region inference.&lt;/strong&gt; On bedrock-runtime you route through one of two inference profiles rather than pinning to a single Region. us.xai.grok-4.6 keeps traffic within the US geography when you have data residency requirements, and global.xai.grok-4.6 routes wor
+&lt;p&gt;&lt;strong&gt;Amazon Bedrock Guardrails.&lt;/strong&gt; Grok 4.6 now supports Guardrails on bedrock-runtime across its APIs, giving you content filters, denied topics, personally identifiable information (PII) redaction, and word policies. You attach a guardrail by ID and version on the req
+&lt;p&gt;&lt;strong&gt;Invocation logging.&lt;/strong&gt; With model invocation logging enabled, Grok 4.6 calls are captured as complete Amazon CloudWatch records: request body, response body, token counts including reasoning tokens, and the inference profile used. Useful for auditing agent runs whe
+&lt;p&gt;&lt;strong&gt;Prompt caching.&lt;/strong&gt; Cached input is billed at roughly a quarter of the standard input rate, which matters for agents that resend a large system prompt or document on every turn. Caching applies to a repeated prefix, so keep stable content at the front of the request
+&lt;p&gt;Tool calling, structured output, image input, response streaming, and encrypted reasoning content are available as well, but those date from the Grok 4.3 launch and are covered in that post.&lt;/p&gt;
+&lt;h2 id="how-grok-4.6-is-packaged-on-amazon-bedrock"&gt;How Grok 4.6 is packaged on Amazon Bedrock&lt;/h2&gt;
+&lt;p&gt;Grok 4.6 accepts text and image input and returns text. Audio, speech, video, and embedding modalities are not supported, and it does not generate images. The model is reachable through two endpoints, and the model ID differs depending on which one you use:&lt;/p&gt;
+&lt;table border="1px" width="100%" cellpadding="10px"&gt;
+ &lt;tbody&gt;
+ &lt;tr&gt;
+ &lt;td&gt;&lt;strong&gt;Endpoint&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Model ID&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Base URL&lt;/strong&gt;&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;bedrock-mantle&lt;/td&gt;
+ &lt;td&gt;&lt;code&gt;xai.grok-4.6&lt;/code&gt;&lt;/td&gt;
+ &lt;td&gt;https://bedrock-mantle.{region}.api.aws/openai/v1&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;bedrock-runtime&lt;/td&gt;
+ &lt;td&gt;&lt;code&gt;us.xai.grok-4.6&lt;/code&gt; (Geo) or &lt;code&gt;global.xai.grok-4.6&lt;/code&gt; (Global)&lt;/td&gt;
+ &lt;td&gt;https://bedrock-runtime.{region}.amazonaws.com/openai/v1&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;/tbody&gt;
+&lt;/table&gt;
+&lt;p&gt;&lt;a href="images/image1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21784-1.png" alt="Architecture diagram of Grok 4.6 access paths on Amazon Bedrock through the bedrock-mantle and bedrock
+&lt;p&gt;On the API side, Grok 4.6 supports the Responses API, the Chat Completions API, and the Converse API. The Invoke API is not supported.&lt;/p&gt;
+&lt;p&gt;Feature support differs by endpoint, which is the detail most likely to shape your integration choice:&lt;/p&gt;
+&lt;p&gt;On bedrock-mantle, supported features include client-side tool calling, reasoning, structured outputs, prompt caching, response streaming, projects, and abuse detection.&lt;/p&gt;
+&lt;p&gt;On bedrock-runtime, supported features include reasoning, prompt caching, response streaming, invocation logs, and projects (default project only). Structured outputs, server-side tool use, intelligent prompt routing, count tokens, and application inference profiles are not supported on tha
+&lt;p&gt;Tool calling works on both endpoints. The model returns a structured function request, your code executes it, and you pass the result back. On bedrock-runtime you can drive that loop through Converse’s &lt;code&gt;toolConfig&lt;/code&gt; or the OpenAI-compatible &lt;code&gt;tools&lt;/code&g
+&lt;p&gt;If your application depends on JSON Schema structured output, that points you at bedrock-mantle. If you want the Converse API or invocation logging, that points you at bedrock-runtime.&lt;/p&gt;
+&lt;h3 id="regions-and-inference-options"&gt;Regions and inference options&lt;/h3&gt;
+&lt;p&gt;Availability differs by endpoint. On bedrock-mantle, Grok 4.6 is available for in-Region inference in US West (Oregon) (us-west-2) . On bedrock-runtime, in-Region inference is not offered. Instead, you invoke the model through cross-Region inference profiles. Geo cross-Region inference is a
+&lt;p&gt;This is a change in shape from the Grok 4.3 launch, where, as noted in the &lt;a href="https://aws.amazon.com/blogs/machine-learning/introducing-grok-on-amazon-bedrock/" target="_blank" rel="noopener"&gt;Grok 4.3 post&lt;/a&gt;, the model used in-Region inference only and Geo and Global cro
+&lt;h3 id="service-tier-and-pricing"&gt;Service tier and pricing&lt;/h3&gt;
+&lt;p&gt;Grok 4.6 supports three service tiers. Standard is pay-per-token with no commitment, selected by setting &lt;code&gt;"service_tier": "default"&lt;/code&gt; or omitting the field. Priority delivers faster, prioritized processing for a premium (&lt;code&gt;"service_tier": "priority"&lt;/code&
+&lt;p&gt;The other two tiers are priced as multipliers on those Standard rates: Priority at 1.75x, a 75 percent premium, and Flex at 0.5x, a 50 percent discount. So the same workload that costs $2.20 per million input tokens on Standard in-Region runs $3.85 on Priority and $1.10 on Flex, which makes
+&lt;p&gt;For reference, xAI lists Grok 4.6 pricing starting at $2 per million input tokens and $6 per million output tokens, with a fast variant at twice the price. Always confirm current rates on the &lt;a href="https://aws.amazon.com/bedrock/pricing/" target="_blank" rel="noopener"&gt;Amazon Bedro
+&lt;h2 id="send-your-first-request"&gt;Send your first request&lt;/h2&gt;
+&lt;p&gt;Before your first call, confirm the model is available to you in the Bedrock console for the Region you plan to use. Grok 4.6 is served through inference profiles rather than on-demand throughput on the bare model ID, which is why requests name &lt;code&gt;us.xai.grok-4.6&lt;/code&gt; or &l
+&lt;p&gt;Grok 4.6 uses OpenAI-compatible APIs, so the OpenAI SDK works against either endpoint after you set the base URL. Install the SDK, and boto3 if you plan to use the Converse API:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-bash"&gt;pip install openai
+pip install boto3&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;Generate a long-term Amazon Bedrock API key from the Amazon Bedrock console for exploration, then set your environment. For bedrock-mantle:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-bash"&gt;export OPENAI_API_KEY="&amp;lt;provide your Bedrock API key&amp;gt;"
+export OPENAI_BASE_URL="https://bedrock-mantle.us-west-2.api.aws/openai/v1"&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;For bedrock-runtime:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-bash"&gt;export OPENAI_API_KEY="&amp;lt;provide your Bedrock API key&amp;gt;"
+export OPENAI_BASE_URL="https://bedrock-runtime.us-east-1.amazonaws.com/openai/v1"&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;A first request on bedrock-mantle with the Chat Completions API:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;from openai import OpenAI
+
+client = OpenAI()
+
+response = client.chat.completions.create(
+ model="xai.grok-4.6",
+ messages=[
+ {"role": "user", "content": "Can you explain the features of Amazon Bedrock?"}
+ ],
+)
+print(response)&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;On bedrock-runtime the difference is the model name: you pass a cross-Region inference profile instead of the bare model ID. This example also switches to the Responses API to show that shape:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;from openai import OpenAI
+
+client = OpenAI()
+
+response = client.responses.create(
+ model="us.xai.grok-4.6",
+ input="Can you explain the features of Amazon Bedrock?",
+)
+print(response)&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;And through the Converse API with boto3. Because reasoning is active, the first content block carries the reasoning and the answer sits in a later block, so search the blocks for the text rather than indexing &lt;code&gt;content[0]&lt;/code&gt;:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;import boto3
+
+client = boto3.client("bedrock-runtime", region_name="us-east-1")
+
+response = client.converse(
+ modelId="us.xai.grok-4.6",
+ messages=[
+ {"role": "user", "content": [{"text": "Can you explain the features of Amazon Bedrock?"}]}
+ ],
+ inferenceConfig={"maxTokens": 2048},
+)
+
+blocks = response["output"]["message"]["content"]
+text = next(b["text"] for b in blocks if "text" in b)
+print(text)&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;On Converse you set the effort level through additionalModelRequestFields rather than a reasoning parameter:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;response = client.converse(
+ modelId="us.xai.grok-4.6",
+ messages=[{"role": "user", "content": [{"text": "What is 17*23? Number only."}]}],
+ inferenceConfig={"maxTokens": 3000},
+ additionalModelRequestFields={"reasoning_effort": "xhigh"},
+)&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;Three operational notes. First, on bedrock-runtime, Grok 4.6 is not available for in-Region inference, so requests must name &lt;code&gt;us.xai.grok-4.6&lt;/code&gt; or &lt;code&gt;global.xai.grok-4.6&lt;/code&gt;.&lt;/p&gt;
+&lt;p&gt;Second, &lt;code&gt;bedrock:InvokeModel&lt;/code&gt; is evaluated against three resources: your account’s default project, the inference profile you name, and the underlying foundation model. The foundation model ARN is wildcarded across Regions because cross-Region profiles route outside t
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-json"&gt;{
+ "Version": "2012-10-17",
+ "Statement": [
+ {
+ "Effect": "Allow",
+ "Action": "bedrock:InvokeModel",
+ "Resource": [
+ "arn:aws:bedrock:{region}:{account-id}:project/default",
+ "arn:aws:bedrock:{region}:{account-id}:inference-profile/us.xai.grok-4.6",
+ "arn:aws:bedrock:*::foundation-model/xai.grok-4.6"
+ ]
+ },
+ {
+ "Effect": "Allow",
+ "Action": "bedrock:CallWithBearerToken",
+ "Resource": "*"
+ }
+ ]
+}&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;List every inference profile you plan to call. Profiles are scoped individually, so a policy naming &lt;code&gt;us.xai.grok-4.6&lt;/code&gt; does not cover &lt;code&gt;global.xai.grok-4.6&lt;/code&gt;.&lt;/p&gt;
+&lt;p&gt;Third, the two authentication mechanisms cover different code paths. An Amazon Bedrock API key in &lt;code&gt;OPENAI_API_KEY&lt;/code&gt; travels as a bearer token and authenticates the OpenAI-compatible calls on both endpoints. The boto3 Converse examples sign with SigV4 instead, drawing o
+&lt;p&gt;Treat a long-term API key as an exploration-only credential. For production, the Grok 4.3 launch post recommends short-term bearer tokens generated from your IAM credentials with the &lt;code&gt;aws-bedrock-token-generator&lt;/code&gt; package, because they expire automatically and keep acc
+&lt;h2 id="working-with-reasoning-effort"&gt;Working with reasoning effort&lt;/h2&gt;
+&lt;p&gt;Reasoning is active on Grok 4.6 by default, and you configure how much of it the model spends through the reasoning parameter with low (the default), medium, high, or xhigh. The xhigh level is new relative to what the Grok 4.3 launch post documented, where the levels were none, low, medium,
+&lt;p&gt;Reasoning content is encrypted. You can have it returned by passing &lt;code&gt;include: ["reasoning.encrypted_content"]&lt;/code&gt; on a Responses API request, then send that content back on subsequent turns to give the model its own prior reasoning as context in a multi-turn conversation
+&lt;p&gt;Encrypted reasoning is a Responses API feature, so this example uses the OpenAI client rather than the boto3 client from the Converse examples above:&lt;/p&gt;
+&lt;div class="hide-language"&gt;
+ &lt;pre&gt;&lt;code class="language-python"&gt;from openai import OpenAI
+
+client = OpenAI() # OPENAI_BASE_URL points at the bedrock-runtime endpoint
+
+response = client.responses.create(
+ model="us.xai.grok-4.6",
+ reasoning={"effort": "high"},
+ include=["reasoning.encrypted_content"],
+ input="Explain quantum entanglement simply.",
+)
+print(response.output_text)&lt;/code&gt;&lt;/pre&gt;
+&lt;/div&gt;
+&lt;p&gt;Because reasoning is by default and effort is per request, effort level is a real cost and latency control. Run short extraction and classification calls at low, and reserve high or xhigh for planning steps and long agent trajectories where an early mistake compounds. Benchmarking effort le
+&lt;h2 id="get-started"&gt;Get started&lt;/h2&gt;
+&lt;p&gt;Grok 4.6 on Amazon Bedrock gives you a model xAI built for long-running agents and ambitious interactive work, with a 500K token context window, four reasoning effort levels, image input, prompt caching, and a choice between the OpenAI-compatible bedrock-mantle endpoint and the bedrock-runt
+&lt;p&gt;To start building, review the &lt;a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-6.html" target="_blank" rel="noopener"&gt;Grok 4.6 model card&lt;/a&gt; for the current Region list, feature matrix, and parameter details, and check the &lt;a href="https://
+&lt;h2 id="sources"&gt;Sources&lt;/h2&gt;
+&lt;ul&gt;
+ &lt;li&gt;Amazon Bedrock Grok 4.6 model card: &lt;a class="uri" href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-6.html" target="_blank" rel="noopener"&gt;https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-xai-grok-4-6.html&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;xAI models in Amazon Bedrock: &lt;a class="uri" href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-xai.html" target="_blank" rel="noopener"&gt;https://docs.aws.amazon.com/bedrock/latest/userguide/model-cards-xai.html&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;xAI, Introducing Grok 4.6: &lt;a class="uri" href="https://x.ai/news/grok-4-6" target="_blank" rel="noopener"&gt;https://x.ai/news/grok-4-6&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;Artificial Analysis, model comparison and benchmark methodology: &lt;a class="uri" href="https://artificialanalysis.ai/models" target="_blank" rel="noopener"&gt;https://artificialanalysis.ai/models&lt;/a&gt;&lt;/li&gt;
+ &lt;li&gt;AWS, Introducing Grok on Amazon Bedrock (Grok 4.3): &lt;a class="uri" href="https://aws.amazon.com/blogs/machine-learning/introducing-grok-on-amazon-bedrock/" target="_blank" rel="noopener"&gt;https://aws.amazon.com/blogs/machine-learning/introducing-grok-on-amazon-bedrock/&lt;/a&gt;&lt;/
+&lt;/ul&gt;
+&lt;hr style="width: 100%"&gt;
+&lt;h2&gt;About the authors&lt;/h2&gt;
+&lt;footer&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;p&gt;&lt;img loading="lazy" class="alignnone size-full wp-image-139858" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/21/Screenshot-2026-09-21-at-10.54.29 AM.png" alt="" width="100" height="117"&gt;&lt;/p&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Suheel Farooq&lt;/h3&gt;
+ &lt;p&gt;Suheel is a Principal Solutions Architect at AWS, specializing in artificial intelligence, machine learning, and generative AI. He helps Foundation Model Provider customers design, build, modernize, and scale their AI/ML and generative AI workloads on AWS. His experience spans the AWS AI/
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;p&gt;&lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21784-3.jpg" alt="Ikenna Izugbokwe" width="100" height="133"&gt;&lt;/p&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Ikenna Izugbokwe&lt;/h3&gt;
+ &lt;p&gt;Ikenna is a Principal Solutions Architect at AWS specializing in networking, containers, and AI infrastructure. He guides model providers through scaling their training and inference systems while enabling rapid deployment of evolving frontier models on AWS. His work increasingly spans ag
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;p&gt;&lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21784-4.jpg" alt="Fabio Branco" width="100" height="133"&gt;&lt;/p&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Fabio Branco&lt;/h3&gt;
+ &lt;p&gt;Fabio is a Senior Customer Solutions Manager at Amazon Web Services (AWS) and strategic advisor guiding foundational model providers in their go-to-market journey. Prior to AWS, he held Product Management, Engineering, Consulting, and Technology Delivery roles across multiple Fortune 500
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;p&gt;&lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21784-5.jpg" alt="Saurabh Trikande" width="100" height="133"&gt;&lt;/p&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Saurabh Trikande&lt;/h3&gt;
+ &lt;p&gt;Saurabh is a Senior Product Manager for Amazon Bedrock and Amazon SageMaker Inference. He is passionate about working with customers and partners, motivated by the goal of democratizing AI. He focuses on core challenges related to deploying complex AI applications, inference with multi-te
+ &lt;/div&gt;
+ &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
+ &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
+ &lt;p&gt;&lt;img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/17/ML-21784-6.jpg" alt="Anirban Gupta" width="100" height="133"&gt;&lt;/p&gt;
+ &lt;/div&gt;
+ &lt;h3 class="lb-h4"&gt;Anirban Gupta&lt;/h3&gt;
+ &lt;p&gt;Anirban is a Principal Engineer at AWS based in Seattle, USA, where he focuses on the design of secure, high-scale model-serving infrastructure for Amazon Bedrock. He has driven the technical work behind several foundation-model launches on the platform. Prior to joining Amazon Bedrock, h
+ &lt;/div&gt;
+&lt;/footer&gt;</content:encoded>
+
+
+
+ </item>
+ <item>
+ <title>How BMW Group detects cost anomalies across 14,000 cloud accounts</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/how-bmw-group-detects-cost-anomalies-across-14000-cloud-accounts/</link>
+
+ <dc:creator><![CDATA[Tareq Haschemi]]></dc:creator>
+ <pubDate>Mon, 21 Sep 2026 16:36:10 +0000</pubDate>
+ <category><![CDATA[Advanced (300)]]></category>
+ <category><![CDATA[Customer Solutions]]></category>
+ <guid isPermaLink="false">b2786889523b42040f67f3b605d5e4d72111c12e</guid>
+
+ <description>BMW Group operates CLEA, a FinOps platform monitoring more than 14,000 cloud accounts. This post shows how BMW added automated daily cost anomaly detection, moving from reactive dashboards to proactive alerts using Prophet forecasting, AWS Step Functions, and a serverless pipeline
+ <content:encoded>&lt;p&gt;&lt;em&gt;This post is co-written with Philipp Karg from BMW Group and Christopher Masurek from Data Reply.&lt;/em&gt;&lt;/p&gt;
+&lt;p&gt;Cost anomalies are hard to spot when you run 14,000 cloud accounts. BMW Group operates Cloud Efficiency Analytics (CLEA), an in-house FinOps system built on AWS with Reply that monitors more than 14,000 cloud accounts across BMW Group’s cloud estate. CLEA began as a set of dashboards in &lt
+&lt;p&gt;To close that gap, CLEA now runs anomaly detection every day and sends email to account owners when spending departs from its expected pattern.&lt;/p&gt;
+&lt;p&gt;This post walks through the forecasting baseline, the filtering logic that decides which deviations are worth an alert, the alert engine, and the serverless architecture that processes every account daily for about $50 per month in compute.&lt;/p&gt;
+&lt;h2 id="what-clea-forecasts"&gt;What CLEA forecasts&lt;/h2&gt;
+&lt;p&gt;CLEA ingests billing data daily from &lt;a href="https://docs.aws.amazon.com/cur/latest/userguide/what-is-cur.html" target="_blank" rel="noopener"&gt;AWS Cost and Usage Reports&lt;/a&gt; (CUR), the primary source, along with the equivalent billing exports from the other providers in BMW Gro
+&lt;p&gt;A single AWS account running &lt;a href="https://aws.amazon.com/ec2/" target="_blank" rel="noopener"&gt;Amazon Elastic Compute Cloud (Amazon EC2)&lt;/a&gt;, &lt;a href="https://aws.amazon.com/s3/" target="_blank" rel="noopener"&gt;Amazon Simple Storage Service (Amazon S3)&lt;/a&gt;, &lt;a h
+&lt;p&gt;We use three terms throughout the rest of this post. &lt;em&gt;Expected spend&lt;/em&gt; is the predicted cost for one account-service pair on one day, based on 365 days of history. &lt;em&gt;Actual spend&lt;/em&gt; is the cost recorded in the billing data for that same account, service, an
+&lt;h2 id="forecasting-and-detection-pipeline"&gt;Forecasting and detection pipeline&lt;/h2&gt;
+&lt;p&gt;CLEA builds its cost baselines with Prophet, the open source forecasting library from Meta. We chose it for its simplicity and its steady performance on cost time series. The model trains on 365 days of daily cost history for each account-service pair, with additive seasonality.&lt;/p&gt;
+&lt;p&gt;&lt;a href="https://aws.amazon.com/step-functions/" target="_blank" rel="noopener"&gt;AWS Step Functions&lt;/a&gt; orchestrates the daily run. A preparation AWS Lambda function discovers the active accounts and writes the list to Amazon S3 as JSON. A Distributed Map then fans the work out a
+&lt;p&gt;The forecast output has two uses: a 12-month rolling forecast, and per-day predicted values that become the expected cost baseline for anomaly detection.&lt;/p&gt;
+&lt;p&gt;We treat forecasting as a pluggable module. The interfaces are the input format (daily cost per account-service) and the output format (per-day predicted values with confidence intervals), so the forecasting engine can be swapped out without touching the detection and alerting layers that a
+&lt;h2 id="from-forecast-to-actionable-alert"&gt;From forecast to actionable alert&lt;/h2&gt;
+&lt;p&gt;A forecast on its own is not an alert. Getting from one to the other takes three things: a baseline that each account-service pair can be measured against, a daily comparison that flags the days falling outside it, and a set of filters that decide which of those days are worth an owner’s at
+&lt;h3 id="setting-the-baseline"&gt;Setting the baseline&lt;/h3&gt;
+&lt;p&gt;A fixed rule, such as alerting whenever daily spend passes a set dollar amount, does not hold up at this scale. Accounts grow, adopt new services, and ramp workloads on purpose, and a fixed rule reads all of that as anomalous. Set the threshold high enough to stay quiet for the largest acco
+&lt;p&gt;The trade-off is that a model that adapts to a trend will eventually absorb one. A sustained step up in spend gets flagged for the first few days and then settles in as the new expected level as the training window catches up. Detection of this kind is strongest on spikes.&lt;/p&gt;
+&lt;p&gt;With a forecast in place, detection becomes a daily comparison. For every account-service pair, CLEA calculates the impact: actual spend minus expected spend. Where actual cost falls outside the confidence interval Prophet produced, CLEA flags the day as a potential anomaly. Because the int
+&lt;h3 id="filtering-down-to-what-matters"&gt;Filtering down to what matters&lt;/h3&gt;
+&lt;p&gt;Some services never enter the model. Before a detection runs, CLEA excludes low-spend services (averaging below $0.10 over the last 3 days), services with fewer than 10 days of history, and other specific line items and charge types that are not relevant to the forecast.&lt;/p&gt;
+&lt;p&gt;CLEA also applies a deviation threshold: a flagged day must deviate by at least 40% from expected spend to stay in scope. From there, the remaining detections pass through two groups of filters. Generic thresholds apply to every account: an anomaly has to clear both the deviation threshold
+&lt;p&gt;&lt;strong&gt;Account-cluster filtering.&lt;/strong&gt; A 900% jump sounds alarming until you look at the absolute numbers: An account that normally spends $0.10 on a service and then spends $1.00 has spiked, but the absolute overspend is negligible. CLEA sorts accounts into four clusters b
+&lt;table border="1px" width="100%" cellpadding="10px"&gt;
+ &lt;tbody&gt;
+ &lt;tr&gt;
+ &lt;td&gt;&lt;strong&gt;Cluster&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Trailing 3-month average spend&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Minimum impact to alert&lt;/strong&gt;&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;1&lt;/td&gt;
+ &lt;td&gt;Less than $100k&lt;/td&gt;
+ &lt;td&gt;More than $300&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;2&lt;/td&gt;
+ &lt;td&gt;$100k to $250k&lt;/td&gt;
+ &lt;td&gt;More than $500&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;3&lt;/td&gt;
+ &lt;td&gt;$250k to $500k&lt;/td&gt;
+ &lt;td&gt;More than $750&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;4&lt;/td&gt;
+ &lt;td&gt;More than $500k&lt;/td&gt;
+ &lt;td&gt;More than $1,000&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;/tbody&gt;
+&lt;/table&gt;
+&lt;p&gt;&lt;strong&gt;Service-specific thresholds.&lt;/strong&gt; Based on operational experience, certain services produce cost spikes as part of their normal usage pattern. &lt;a href="https://aws.amazon.com/glue/" target="_blank" rel="noopener"&gt;AWS Glue&lt;/a&gt;, &lt;a href="https://aws.amaz
+&lt;p&gt;&lt;strong&gt;Account-specific overrides.&lt;/strong&gt; Accounts on a reduced-sensitivity list must exceed three times the standard thresholds before an alert fires. This covers teams with known volatile workloads who asked for fewer notifications.&lt;/p&gt;
+&lt;p&gt;Figure 1 shows how each layer narrows the set: a wide band of Prophet anomalies on the left, and on the right the few that survive both the generic thresholds and the case-specific overrides.&lt;/p&gt;
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21627-1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21627-1.png" alt="A funnel narrowing left
+ &lt;p class="wp-caption-text"&gt;Figure 1: Each filtering layer reduces false positives&lt;/p&gt;
+&lt;/div&gt;
+&lt;h2 id="calibration-and-what-automation-cannot-decide"&gt;Calibration, and what automation cannot decide&lt;/h2&gt;
+&lt;p&gt;CLEA can see operations, usage types, and the resulting costs. It cannot see intent. Only the account owner knows whether a cost increase was planned, such as a new workload rollout or a migration. That boundary between detection and judgment is permanent, so we tune thresholds against user
+&lt;p&gt;One more piece of bookkeeping matters at this scale. CLEA merges consecutive flagged days into date ranges, and the grouping logic reads every historical model snapshot instead of only the latest run. Without that, ranges fragment whenever Prophet reclassifies an individual day between exec
+&lt;h2 id="alert-engine-and-delivery"&gt;Alert engine and delivery&lt;/h2&gt;
+&lt;p&gt;Alerting runs as a separate process once detection finishes. The engine queries the day’s active anomalies and deduplicates them by matching the detected date against the current date, so each anomaly produces a single alert. Anomalies that began within the last four days generate an alert.
+&lt;p&gt;Every alert email carries the context an owner needs to act:&lt;/p&gt;
+&lt;ul&gt;
+ &lt;li&gt;The account ID and name, the account owners, and the department hierarchy, from BMW Group metadata sources.&lt;/li&gt;
+ &lt;li&gt;The affected service.&lt;/li&gt;
+ &lt;li&gt;The anomaly date range and duration.&lt;/li&gt;
+ &lt;li&gt;Expected spend against actual spend.&lt;/li&gt;
+ &lt;li&gt;The absolute impact and the percentage deviation.&lt;/li&gt;
+ &lt;li&gt;The accumulated impact across every concurrent anomaly on that account.&lt;/li&gt;
+&lt;/ul&gt;
+&lt;p&gt;An Excel attachment holds the full table.&lt;/p&gt;
+&lt;p&gt;Figure 2 shows an alert for an example account. The email names the account and the recipient’s role. It states that costs exceeded expected spending by $1,775.13 and notes that anomalies are detected from spending spikes and may include false positives. Under Recommended actions it asks th
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21627-2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21627-2.png" alt="An anomaly alert email
+ &lt;p class="wp-caption-text"&gt;Figure 2: An anomaly alert as an account owner receives it&lt;/p&gt;
+&lt;/div&gt;
+&lt;h2 id="self-service-root-cause-analysis"&gt;Self-service root cause analysis&lt;/h2&gt;
+&lt;p&gt;When an alert lands, the owner can investigate without involving the platform team. CLEA provides an anomaly dashboard in Amazon Quick Sight that lists the detected anomalies for each account, using the same fields as the alert email. Owners can widen the filter to include anomalies that we
+&lt;p&gt;When you select an anomaly, CLEA opens a drill-down chart. Two bar charts break the account’s actual daily spend down by operation and by usage type, which is usually enough to confirm the spike and place it in time. A detail table lists the usage types driving the cost, such as &lt;code&gt
+&lt;p&gt;Figure 3 shows the drill-down for an account with a confirmed spike. The left chart plots daily usage cost by operation from early May to mid-June 2026. &lt;code&gt;RunInstances&lt;/code&gt; dominates, and a single day reaches about $2,600 against a baseline near $900. The right chart plots
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21627-3.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ML-21627-3.png" alt="Two bar charts of daily

Diff display stops at 400 lines. The line counts above are from the whole diff. 61 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.