Change
cd61676
cd61676bc5da4e208032c75da88e689a84caee41 · commit on GitHub
aws-blog-feed: changed (679716 bytes, HTTP 200)
raw/aws-blog-feed/response.xml modified
- Source
- aws-blog-feed
- Lines added
- +4,700
- Lines removed
- -4,205
- Stored bytes at this commit
- 679,716
- Timestamp
- observed
- Raw artifact at this commit
- raw/aws-blog-feed/response.xml
Recorded headers
| observed_at | 2026-09-10T04:43:13.706Z |
|---|---|
| origin_date | null |
| status | 200 |
| final URL | https://aws.amazon.com/blogs/machine-learning/feed/ |
| etag | null |
| last-modified | Wed, 09 Sep 2026 22:46:40 GMT |
| date | Thu, 10 Sep 2026 04:43:13 GMT |
| age | null |
| cache-control | null |
| cf-cache-status | null |
| content-encoding | null |
| content-length | null |
@
@@ -5,7 +5,7 @@ <atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/> <link>https://aws.amazon.com/blogs/machine-learning/</link> <description>Official Machine Learning Blog of Amazon Web Services</description>-
<lastBuildDate>Tue, 08 Sep 2026 22:06:58 +0000</lastBuildDate>+
<lastBuildDate>Wed, 09 Sep 2026 22:26:29 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod>@
@@ -13,64 +13,577 @@ 1 </sy:updateFrequency> <item>-
<title>Take on your most ambitious work with GPT-6 Astra on Amazon Bedrock</title>-
<link>https://aws.amazon.com/blogs/machine-learning/take-on-your-most-ambitious-work-with-gpt-6-astra-on-amazon-bedrock/</link>+
<title>Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM</title>+
<link>https://aws.amazon.com/blogs/machine-learning/deploying-qwen3-8-2-4t-a95b-on-amazon-sagemaker-hyperpod-with-vllm/</link> -
<dc:creator><![CDATA[Tanvi Girinath]]></dc:creator>-
<pubDate>Tue, 08 Sep 2026 22:06:58 +0000</pubDate>-
<category><![CDATA[Amazon Bedrock]]></category>-
<category><![CDATA[Announcements]]></category>-
<category><![CDATA[Intermediate (200)]]></category>-
<guid isPermaLink="false">68d48019de714857964056351fc4d79e51c3ab9f</guid>+
<dc:creator><![CDATA[Dmitry Soldatkin]]></dc:creator>+
<pubDate>Wed, 09 Sep 2026 22:26:29 +0000</pubDate>+
<category><![CDATA[Advanced (300)]]></category>+
<category><![CDATA[Amazon SageMaker HyperPod]]></category>+
<category><![CDATA[Technical How-to]]></category>+
<guid isPermaLink="false">6973e0078540c7d789cc946cb3c7e1fd791a8c2f</guid>-
<description>GPT-6 Astra from OpenAI is now generally available on Amazon Bedrock. It brings deeper reasoning and sharper judgment to your most demanding tasks, running on the Amazon Bedrock inference engine built for high performance, security, and scale.</description>-
<content:encoded><p><em>GPT-6 Astra from OpenAI brings greater depth and judgment to your most demanding tasks and runs on the Amazon Bedrock inference engine built for high performance, security, and scale.</em></p> -
<p>Organizations are already running AI agents that write code, analyze data, and automate complex workflows at production scale on Amazon Bedrock. GPT-6 Astra raises the potential of what those agents can deliver. It applies deeper reasoning and sharper judgment to complex business decisions,…-
<p>Today, GPT-6 Astra, the latest and most capable OpenAI model, is generally available on <a href="https://aws.amazon.com/bedrock/" target="_blank" rel="noopener">Amazon Bedrock</a>. You can call the <a href="https://aws.amazon.com/bedrock/openai/" target="_blank" rel="noopener…-
<h2 id="greater-depth-for-complex-decisions">Greater depth for complex decisions</h2> -
<p>GPT-6 Astra brings greater depth to work that requires you to reconcile competing inputs, trace dependencies, and determine what to prioritize. When performing financial analysis, it can help you reconcile conflicting data sources and identify discrepancies that could change a recommendatio…-
<p>For workflows that reuse the same context across requests, such as recurring document review, codebase analysis, or agents grounded in company standards, GPT-6 Astra supports both implicit and explicit prompt caching. With explicit caching, you can set cache breakpoints to control which con…-
<h2 id="layered-security-and-governance-for-production-ai">Layered security and governance for production AI</h2> -
<p>Model-level safeguards work alongside the security and governance controls of Amazon Bedrock. OpenAI evaluated GPT-6 Astra through its <a href="https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf" target="_blank" rel="noopener">Preparedness Fr…-
<p>Amazon Bedrock protects your inference data and governs access at every model invocation. Zero-operator access is enforced at the chip, so even AWS operators can’t access your prompts and completions during inference. Data is encrypted in transit and at rest. Access is governed by your AWS …-
<p>Your inference data isn’t used for model training, and using GPT-6 Astra doesn’t require you to opt into sharing your data with OpenAI. For <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/abuse-detection.html" target="_blank" rel="noopener">automated abuse detection</…-
<h2 id="build-code-and-work-with-astra">Build, code, and work with Astra</h2> -
<p>You can use&nbsp;GPT-6 Astra for inference, knowledge work, and software development.</p> -
<h3 id="for-platform-and-application-teams">Integrate Astra into your applications</h3> -
<p>You can integrate GPT-6 Astra directly into your applications using supported Amazon Bedrock APIs. Use it to power autonomous agents that handle complex, multistep workflows, build internal tools that analyze and synthesize documents at scale, or create customer-facing applications that req…-
<h3 id="for-business-and-operations-teams">Turn complex tasks into finished deliverables</h3> -
<p>ChatGPT Work is a productivity agent for turning complex business tasks into finished deliverables. With GPT-6 Astra, it can gather information across applications and files, use the web, and produce spreadsheets, slides, documents, and sites. You can control which applications and websites…-
<p>ChatGPT Work is available through the ChatGPT desktop app for Mac and Windows. New enterprise plugins introduced alongside this release extend browser-use capabilities across business intelligence tools, Workday, Navan, and Avalara for tasks across data analytics, operations, and finance. T…-
<h3 id="for-development-teams">Build, test, and ship code faster</h3> -
<p>Codex is a software engineering agent that works with local files, repositories, terminals, developer tools, and development environments to write features, fix bugs, run tests, and open pull requests. When you configure Codex to use GPT-6 Astra on Amazon Bedrock, it applies Astra’s reasoni…-
<p>You can access Codex through the ChatGPT desktop app, CLI, VS Code, JetBrains IDEs, and Xcode. For AWS development, the&nbsp;<a href="https://docs.aws.amazon.com/agent-toolkit/latest/userguide/quick-start.html" target="_blank" rel="noopener">Agent Toolkit for AWS</a>&nbs…-
<h2 id="get-started">Get started</h2> -
<p>You can get started in the <a href="https://us-east-1.console.aws.amazon.com/bedrock/home?region=us-east-1#/" target="_blank" rel="noopener">Amazon Bedrock console</a> or programmatically through supported Amazon Bedrock APIs. For information about supported <a href="https://…-
<p><em>Interested in how Amazon Bedrock can support your team? <a href="https://pages.awscloud.com/Amazon-Bedrock-Contact-Us.html" target="_blank" rel="noopener">Connect with us</a> to start the conversation.</em></p> +
<description>Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP specu…+
<content:encoded><p>On August 12, 2026, Alibaba’s Qwen team released <a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B" target="_blank" rel="noopener">Qwen3.8-2.4T-A95B</a>. This is the first time a Qwen-Max-class model has been made available as open weights. With 2…+
<p>Open weights models give you full control. Data stays within your infrastructure, inference behavior can be customized, and there are no per-token API fees at scale. The trade-off is operational: hosting a 2.4T-parameter model requires purpose-built GPU infrastructure and an optimized servi…+
<p>In this post we show how to deploy Qwen3.8-2.4T-A95B on <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod.html" target="_blank" rel="noopener">Amazon SageMaker HyperPod</a> using <a href="https://docs.vllm.ai/" target="_blank" rel="noopener">vLLM&…+
<p>This is the second post in our series on deploying open trillion-parameter models on Amazon SageMaker HyperPod. For the first post covering Kimi K3, see <a href="https://aws.amazon.com/blogs/machine-learning/deploying-kimi-k3-on-amazon-sagemaker-hyperpod-and-amazon-eks/" target="_blank" …+
<h2 id="qwen3.8-2.4t-a95b-at-a-glance">Qwen3.8-2.4T-A95B at a glance</h2> +
<p>Qwen3.8-2.4T-A95B (the open-weight release of Qwen3.8-Max) is the largest and most capable model in the Qwen family. The following is a summary of the key architectural details relevant to deployment.</p> +
<h3 id="architecture">Architecture</h3> +
<table border="1px" width="100%" cellpadding="10px"> +
<tbody> +
<tr> +
<td><strong>Attribute</strong></td> +
<td><strong>Value</strong></td> +
</tr> +
<tr> +
<td>Total parameters</td> +
<td>2.4 T</td> +
</tr> +
<tr> +
<td>Activated parameters per token</td> +
<td>95 B</td> +
</tr> +
<tr> +
<td>Architecture</td> +
<td>Fine-grained Mixture of Experts (MoE)</td> +
</tr> +
<tr> +
<td>Expert count</td> +
<td>512 routed + 1 shared (10 routed experts activated per token)</td> +
</tr> +
<tr> +
<td>Layers</td> +
<td>92</td> +
</tr> +
<tr> +
<td>Layer layout</td> +
<td>3 × (Gated DeltaNet → MoE) → 1 × (Gated Attention → MoE), repeated</td> +
</tr> +
<tr> +
<td>Context window</td> +
<td>262,144 tokens native. Extensible to 1,010,000</td> +
</tr> +
<tr> +
<td>Max output length</td> +
<td>128K tokens</td> +
</tr> +
<tr> +
<td>Multi-Token Prediction</td> +
<td>Native MTP draft heads (enables speculative decoding without a separate model)</td> +
</tr> +
</tbody> +
</table> +
<p>The hybrid attention design is key to efficient long-context inference. <strong>Gated DeltaNet</strong> layers (69 of 92) use linear attention with a bounded recurrent state, replacing the growing KV-cache with a fixed-size memory. <strong>Gated Attention</strong> la…+
<p>The fine-grained MoE distributes capacity across 512 small experts rather than a few large ones, improving routing efficiency and specialization. Only approximately 95B parameters are active per forward pass, so serving costs track activated parameters, not the full 2.4T.</p> +
<h3 id="capabilities-and-reasoning-control">Capabilities and reasoning control</h3> +
<p>Qwen3.8 is designed for agentic execution: multi-step coding, autonomous tool use, long-horizon planning, and complex research workflows. It includes built-in reasoning controls through the <code>reasoning_effort</code> parameter (<code>low</code>, <code>medium…+
<h3 id="model-weights-and-quantization">Model weights and quantization</h3> +
<p>The open weights are published on <a href="https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B" target="_blank" rel="noopener">Hugging Face</a> in the standard Transformers format. Community quantizations include MXFP4 and NVFP4 (W4A4), which compress the model to approximately 1.2 TB…+
<h3 id="benchmark-highlights">Benchmark highlights</h3> +
<p>According to the vendor’s benchmarking results, Qwen3.8-2.4T-A95 shows particular strength in research workflows (PaperBench 93.0), instruction following (IFBench 82.8), and terminal-based coding (86.6). It performs comparably with leading frontier models across most categories, with remain…+
<h2 id="why-amazon-sagemaker-hyperpod-for-large-moe-inference">Why Amazon SageMaker HyperPod for large MoE inference</h2> +
<p>Deploying a 2.4T-parameter model is not only a GPU problem. It requires orchestration that handles model download, container scheduling, health monitoring, autoscaling, and node failures without manual intervention. <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyper…+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-21725-1.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-21725-1.png" alt="Architecture diagram of…+
<p class="wp-caption-text">Figure 1: High-level architecture of Amazon SageMaker HyperPod</p>+
</div> +
<p><strong>EKS-orchestrated clusters.</strong> HyperPod clusters use Amazon Elastic Kubernetes Service (Amazon EKS) as the control plane. You get the full Kubernetes landscape (<code>kubectl</code>, Helm charts, custom resource definitions), while AWS manages the underl…+
<p><strong>Inference Operator.</strong> The HyperPod Inference Operator (installed automatically or as an EKS Add-on) provides a single custom resource definition (CRD), <code>InferenceEndpointConfig</code>, that declaratively specifies your model, container image, GPU …+
<ul> +
<li>Model weight download (from Hugging Face Hub, Amazon Simple Storage Service (Amazon S3), or Amazon FSx).</li> +
<li>Container scheduling and GPU allocation.</li> +
<li>Health checks and readiness gates.</li> +
<li>Rolling updates and endpoint lifecycle management.</li> +
<li>Autoscaling through KEDA with Amazon CloudWatch or Prometheus metrics.</li> +
</ul> +
<p><strong>Reserved capacity with Flexible Training Plans.</strong> The <code>ml.p6-b300.48xlarge</code> instance type requires reserved capacity. Flexible Training Plans provide committed GPU reservations that can be allocated directly to your HyperPod cluster. There’s…+
<p><strong>Resilience.</strong> HyperPod continuously monitors node health and automatically replaces degraded nodes. For sustained inference workloads running 24/7, this alleviates the operational overhead of manually detecting and recovering from hardware failures.</p> +
<p><strong>Additional inference features</strong> (Inference Operator v3.x):</p> +
<ul> +
<li><a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-model-deployment-dpd.html" target="_blank" rel="noopener">Disaggregated Prefill and Decode</a> (DPD) – separates prefill and decode onto distinct GPU pools for predictable per-token latency under concu…+
<li><a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-model-deployment-data-capture.html" target="_blank" rel="noopener">Inference data capture</a> – log inputs/outputs at the endpoint, load balancer, or pod level.</li> +
<li>Local NVMe model deployment – load weights from node-local storage to reduce cold-start latency.</li> +
<li><a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-model-deployment-custom-certs.html" target="_blank" rel="noopener">Amazon Route 53 DNS management</a> – automatic custom domain records for your endpoints.</li> +
</ul> +
<p>In short: You write a YAML manifest describing <em>what</em> to deploy. HyperPod handles <em>how</em> to run it reliably at scale.</p> +
<h2 id="infrastructure-sizing-matching-hardware-to-the-model">Infrastructure sizing: Matching hardware to the model</h2> +
<h3 id="the-p6-b300-instance">The p6-b300 instance</h3> +
<p>The <code>ml.p6-b300.48xlarge</code> provides the compute density required for single-node serving of Qwen3.8:</p> +
<table border="1px" width="100%" cellpadding="10px"> +
<tbody> +
<tr> +
<td><strong>Resource</strong></td> +
<td><strong>Specification</strong></td> +
</tr> +
<tr> +
<td>GPUs</td> +
<td>8× NVIDIA B300 (Blackwell Ultra)</td> +
</tr> +
<tr> +
<td>GPU memory</td> +
<td>288 GB HBM3e per GPU (<strong>2.1 TB total</strong>)</td> +
</tr> +
<tr> +
<td>GPU memory bandwidth</td> +
<td>8 TB/s per GPU</td> +
</tr> +
<tr> +
<td>GPU interconnect</td> +
<td>NVLink + NVSwitch, 14.4 TB/s bisection bandwidth</td> +
</tr> +
<tr> +
<td>FP4 compute</td> +
<td>~15 PFLOPS per GPU (120 PFLOPS total)</td> +
</tr> +
<tr> +
<td>vCPUs</td> +
<td>192 (Intel Xeon Emerald Rapids)</td> +
</tr> +
<tr> +
<td>System memory</td> +
<td>4,096 GiB</td> +
</tr> +
<tr> +
<td>Networking</td> +
<td>6,400 Gbps EFA</td> +
</tr> +
<tr> +
<td>Local storage</td> +
<td>3.8 TB NVMe SSD</td> +
</tr> +
</tbody> +
</table> +
<h3 id="why-nvfp4-quantization">Why NVFP4 quantization</h3> +
<p>At BF16 precision, Qwen3.8’s 2.4T parameters require approximately 4.8 TB of memory for weights alone, exceeding a single 8-GPU node. NVFP4 (W4A4) quantization compresses weights to approximately 4 bits per parameter, bringing the total weight footprint to approximately 1.2 TB. This fits co…+
<h3 id="memory-budget">Memory budget</h3> +
<p>A rough breakdown for a single p6-b300 node:</p> +
<table border="1px" width="100%" cellpadding="10px"> +
<tbody> +
<tr> +
<td><strong>Component</strong></td> +
<td><strong>Estimated Size</strong></td> +
<td><strong>Notes</strong></td> +
</tr> +
<tr> +
<td>Model weights (NVFP4)</td> +
<td>~1.2 TB</td> +
<td>2.4T params × 4 bits</td> +
</tr> +
<tr> +
<td>KV-cache (full attention layers)</td> +
<td>Variable</td> +
<td>23 layers × KV heads × context length</td> +
</tr> +
<tr> +
<td>Recurrent state (DeltaNet layers)</td> +
<td>Fixed ~50–100 GB</td> +
<td>69 layers × bounded state (does not grow with context)</td> +
</tr> +
<tr> +
<td>Activations + overhead</td> +
<td>~100–200 GB</td> +
<td>Tensor-parallel buffers, framework overhead</td> +
</tr> +
<tr> +
<td><strong>Available headroom</strong></td> +
<td><strong>~500–700 GB</strong></td> +
<td>For batching and longer contexts</td> +
</tr> +
</tbody> +
</table> +
<p>The hybrid attention architecture is a key advantage here: the 69 DeltaNet layers maintain a fixed-size recurrent state regardless of context length, unlike traditional models where KV-cache grows linearly with every layer. Only the 23 full-attention layers contribute to context-dependent m…+
<h3 id="throughput-expectations">Throughput expectations</h3> +
<p>Reference numbers from NVIDIA’s Day-0 benchmarks on GB300 NVL72 (FP8, 72 GPUs): &gt;4K tokens/sec/GPU, &gt;350 tokens/sec/user. A single 8-GPU p6-b300 node with NVFP4 will deliver proportionally lower aggregate throughput but remains well-suited for production inference workloads wi…+
<h3 id="capacity-procurement">Capacity procurement</h3> +
<p>The <code>ml.p6-b300.48xlarge</code> instance type isn’t available on-demand. You must procure capacity through a <strong>Flexible Training Plan</strong>, a committed reservation of GPU availability for your HyperPod cluster. Set the target Availability Zone to match…+
<h2 id="vllm-configuration-deep-dive">vLLM configuration deep dive</h2> +
<p>This section details the vLLM serving parameters for Qwen3.8 on a single p6-b300 node. The configuration is informed by the <a href="https://recipes.vllm.ai/Qwen/Qwen3.8-2.4T-A95B?hardware=b300&amp;variant=nvfp4&amp;features=tool_calling,reasoning,spec_decoding" target="_blank" r…+
<h3 id="base-serving-command">Base serving command</h3> +
<p>The full <code>vllm serve</code> invocation:</p> +
<div class="hide-language"> +
<pre><code class="language-bash">vllm serve Inferact/Qwen3.8-2.4T-A95B-NVFP4 \+
--tensor-parallel-size 8 \+
--quantization nvfp4 \+
--load-format fastsafetensors \+
--trust-remote-code \+
--enable-prefix-caching \+
--moe-backend auto \+
--reasoning-parser qwen3 \+
--enable-auto-tool-choice \+
--tool-call-parser qwen3 \+
--speculative-config '{"method":"mtp","num_speculative_tokens":1}' \+
--served-model-name Qwen3.8</code></pre> +
</div> +
<p>Key flags explained:</p> +
<ul> +
<li><code>--tensor-parallel-size 8</code> – shards the model across all 8 B300 GPUs.</li> +
<li><code>--quantization nvfp4</code> – activates NVIDIA FP4 (W4A4) quantization so the 2.4T model fits in 2.1 TB of GPU memory.</li> +
<li><code>--load-format fastsafetensors</code> – uses accelerated weight deserialization for faster cold-start.</li> +
<li><code>--trust-remote-code</code> – required for Qwen3.8’s custom modeling code on Hugging Face.</li> +
<li><code>--enable-prefix-caching</code> – reuses computed KV-cache across requests that share prompt prefixes. Critical for multi-turn agentic conversations where the system prompt and conversation history repeat.</li> +
<li><code>--moe-backend auto</code> – lets vLLM select the optimal MoE dispatch kernel for the hardware.</li> +
</ul> +
<h3 id="reasoning-thinking-mode">Reasoning (thinking mode)</h3> +
<p>The <code>--reasoning-parser qwen3</code> flag extracts reasoning content from the model’s <code>&lt;think&gt;...&lt;/think&gt;</code> output blocks. Key behaviors:</p> +
<ul> +
<li>Qwen3.8 reasoning is <strong>enabled by default</strong> – no extra flag needed on the model side.</li> +
<li>The API response separates <code>reasoning_content</code> (the thinking trace) from <code>content</code> (the final answer).</li> +
<li>To disable thinking per-request, pass <code>extra_body={"chat_template_kwargs": {"enable_thinking": False}}</code> in the client call.</li> +
<li>Structured output (<code>guided_json</code>, <code>guided_regex</code>) works alongside reasoning – the structured output engine constrains only the <code>content</code> field.</li> +
</ul> +
<h3 id="tool-calling-function-calling">Tool calling (function calling)</h3> +
<p>The <code>--enable-auto-tool-choice</code> and <code>--tool-call-parser qwen3</code> flags enable OpenAI-compatible function calling:</p> +
<ul> +
<li>Supports <code>tool_choice</code> values: <code>auto</code>, <code>required</code>, <code>none</code>, and named functions.</li> +
<li>Tool calls are parsed from the <code>content</code> field only — the <code>reasoning_content</code> is not parsed for function calls. This means the model can reason about <em>which</em> tool to call, then emit the structured call separately.</li>…+
<li>When <code>tool_choice="auto"</code> and <code>strict: true</code> is set on a tool definition, vLLM enforces schema-constrained decoding for tool arguments, facilitating valid JSON output.</li> +
</ul> +
<h3 id="speculative-decoding-native-mtp">Speculative decoding (native MTP)</h3> +
<p>The <code>--speculative-config '{"method":"mtp","num_speculative_tokens":1}'</code> flag enables Multi-Token Prediction using Qwen3.8’s built-in draft heads:</p> +
<ul> +
<li>Qwen3.8 was trained with MTP – lightweight draft heads are bundled in the model weights. No separate draft model download or configuration is required.</li> +
<li>The draft head predicts the next N tokens in parallel, then verifies them in a single forward pass. Accepted tokens skip individual decode steps, increasing throughput.</li> +
<li><code>num_speculative_tokens: 1</code> is the safe starting point. Increase to 2–3 for throughput-sensitive workloads once you’ve validated that the acceptance rate remains high (monitor through vLLM’s <code>/metrics</code> endpoint).</li> +
<li>MTP adds minimal latency overhead on the draft step because the heads reuse the model’s existing hidden states.</li> +
</ul> +
<h2 id="deployment-walkthrough-on-sagemaker-hyperpod">Deployment walkthrough on SageMaker HyperPod</h2> +
<p>The complete deployment manifests and scripts used in this post are available in our <a href="https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/SageMakerHyperpod/Qwen3.8-2.4T-A95B" target="_blank" rel="noopener">GitHub repository</a>.</p> +
<h3 id="prerequisites">Prerequisites</h3> +
<p>Before deploying the model, you need a running HyperPod cluster with p6-b300 capacity:</p> +
<ol type="1"> +
<li><strong>Create a HyperPod cluster with EKS orchestration.</strong> In the Amazon SageMaker AI console, navigate to <strong>HyperPod Clusters</strong>, then choose <strong>Create</strong>. Choose <strong>Orchestrated by Amazon EKS</strong> an…+
<li><strong>Provision a Flexible Training Plan.</strong> Under the instance group configuration, select <strong>Training plan</strong> as the capacity source. Create or attach a plan covering <code>ml.p6-b300.48xlarge</code> with the instance count and dura…+
<li><strong>Add a p6-b300 worker group.</strong> Add an instance group with <code>ml.p6-b300.48xlarge</code> and at least 1 instance. Wait for the cluster to reach <strong>Active</strong> state with healthy GPU nodes.</li> +
<li><strong>Verify access.</strong> Confirm you can reach the cluster: +
<div class="hide-language"> +
<pre><code class="language-bash">kubectl get nodes+
# Expect node(s) with nvidia.com/gpu: 8 capacity</code></pre> +
</div> </li> +
</ol> +
<h3 id="inferenceendpointconfig-manifest">InferenceEndpointConfig manifest</h3> +
<p>Apply the following <code>InferenceEndpointConfig</code> to deploy Qwen3.8 with the vLLM configuration:</p> +
<div class="hide-language"> +
<pre><code class="language-yaml">apiVersion: inference.sagemaker.aws.amazon.com/v1+
kind: InferenceEndpointConfig+
metadata:+
name: qwen38+
spec:+
modelName: qwen38+
instanceType: ml.p6-b300.48xlarge+
invocationEndpoint: v1/chat/completions+
replicas: 1+
modelSourceConfig:+
huggingFaceModel:+
modelId: Inferact/Qwen3.8-2.4T-A95B-NVFP4+
modelSourceType: huggingface+
worker:+
image: vllm/vllm-openai:qwen38+
modelInvocationPort:+
containerPort: 8000+
name: http+
modelVolumeMount:+
mountPath: /opt/ml/model+
name: model-weights+
resources:+
limits:+
nvidia.com/gpu: 8+
requests:+
nvidia.com/gpu: 8+
args:+
- "--model"+
- "/opt/ml/model"+
- "--serving-model-name"+
- "Qwen3.8"+
- "--linear-backend"+
- "flashinfer_cutedsl"+
- "--trust-remote-code"+
- "--enable-prefix-caching"+
- "--enable-auto-tool-choice"+
- "--tool-call-parser"+
- "qwen3_coder"+
- "--reasoning-parser"+
- "qwen3"+
- "--served-model-name"+
- "Qwen3.8"+
- "--tensor-parallel-size"+
- "8"+
environmentVariables:+
- name: "VLLM_ENGINE_READY_TIMEOUT_S"+
value: "1800"</code></pre> +
</div> +
<p>This manifest is also available in the <a href="https://github.com/aws-samples/sagemaker-genai-hosting-examples/tree/main/SageMakerHyperpod/Qwen3.8-2.4T-A95B" target="_blank" rel="noopener">GitHub repository</a>.</p> +
<h3 id="applying-and-monitoring">Applying and monitoring</h3> +
<p>Apply the manifest:</p> +
<div class="hide-language"> +
<pre><code class="language-bash">kubectl apply -f qwen.yaml</code></pre> +
</div> +
<p>Monitor the deployment progress:</p> +
<div class="hide-language"> +
<pre><code class="language-bash"># Watch the InferenceEndpointConfig status+
kubectl get inferenceendpointconfig qwen38 -w+
+
# Check pod status (model download and container startup)+
kubectl get pods -l model-name=qwen38+
+
# View vLLM startup logs+
kubectl logs -f &lt;pod-name&gt; --tail=100</code></pre> +
</div> +
<p>The deployment proceeds through these stages: <strong>model download</strong> (approximately 1.2 TB from Hugging Face, time depends on network bandwidth) then <strong>weight loading</strong> (fastsafetensors deserialization to GPU memory) then <strong>health ch…+
<p>After the endpoint shows <code>Ready</code>, you can send requests to the service:</p> +
<div class="hide-language"> +
<pre><code class="language-bash"># Get the service endpoint+
kubectl get svc -l model-name=qwen38+
+
# Quick health check+
curl http://&lt;service-endpoint&gt;:8000/health</code></pre> +
</div> +
<h2 id="inference-in-action-calling-the-endpoint">Inference in action: Calling the endpoint</h2> +
<p>After the endpoint is ready, it exposes an OpenAI-compatible API. You can use the standard OpenAI Python SDK, <code>curl</code>, or another HTTP client (note that in this deployment example the endpoint isn’t exposed to the public internet).</p> +
<h3 id="basic-chat-completion-with-reasoning">Basic chat completion (with reasoning)</h3> +
<div class="hide-language"> +
<pre><code class="language-python">from openai import OpenAI+
+
client = OpenAI(+
base_url="http://&lt;service-endpoint&gt;:8000/v1",+
api_key="unused", # vLLM does not require auth by default+
)+
+
response = client.chat.completions.create(+
model="Qwen3.8",+
messages=[{"role": "user", "content": "Explain the trade-offs of MoE vs dense models for inference."}],+
temperature=0.6,+
top_p=0.95,+
)+
+
# Reasoning trace (the model's thinking)+
print("Thinking:", response.choices[0].message.reasoning_content)+
+
# Final answer+
print("Answer:", response.choices[0].message.content)</code></pre> +
</div> +
<p>To control reasoning depth per request, pass <code>reasoning_effort</code>:</p> +
<div class="hide-language"> +
<pre><code class="language-python">response = client.chat.completions.create(+
model="Qwen3.8",+
messages=[{"role": "user", "content": "What is 2+2?"}],Diff display stops at 400 lines. The line counts above are from the whole diff. 37 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.