Change
5f66eab
5f66eab52d2587f7a904c9532b6142049ab88e32 · commit on GitHub
aws-blog-feed: changed (549833 bytes, HTTP 200)
raw/aws-blog-feed/response.xml modified
- Source
- aws-blog-feed
- Lines added
- +548
- Lines removed
- -611
- Stored bytes at this commit
- 549,833
- Timestamp
- observed
- Raw artifact at this commit
- raw/aws-blog-feed/response.xml
Recorded headers
| observed_at | 2026-10-09T06:08:54.672Z |
|---|---|
| origin_date | null |
| status | 200 |
| final URL | https://aws.amazon.com/blogs/machine-learning/feed/ |
| etag | null |
| last-modified | Thu, 08 Oct 2026 21:16:19 GMT |
| date | Fri, 09 Oct 2026 06:08:54 GMT |
| age | null |
| cache-control | null |
| cf-cache-status | null |
| content-encoding | null |
| content-length | null |
@
@@ -5,7 +5,7 @@ <atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/> <link>https://aws.amazon.com/blogs/machine-learning/</link> <description>Official Machine Learning Blog of Amazon Web Services</description>-
<lastBuildDate>Wed, 07 Oct 2026 23:32:02 +0000</lastBuildDate>+
<lastBuildDate>Thu, 08 Oct 2026 18:33:29 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod>@
@@ -13,6 +13,552 @@ 1 </sy:updateFrequency> <item>+
<title>Pay-per-inference for AI agents: How BlockRun and Incarna use Amazon Bedrock AgentCore payments</title>+
<link>https://aws.amazon.com/blogs/machine-learning/pay-per-inference-for-ai-agents-how-blockrun-and-incarna-use-amazon-bedrock-agentcore-payments/</link>+
+
<dc:creator><![CDATA[Peter Jiang]]></dc:creator>+
<pubDate>Thu, 08 Oct 2026 18:33:29 +0000</pubDate>+
<category><![CDATA[Amazon Bedrock AgentCore]]></category>+
<category><![CDATA[Customer Solutions]]></category>+
<category><![CDATA[Foundational (100)]]></category>+
<guid isPermaLink="false">0d337b36c3f07d0b91dbfe91ce290d78c3de5f56</guid>+
+
<description>Amazon Bedrock AgentCore payments gives AI agents a managed way to pay for services on demand, with spending limits enforced by the infrastructure. See how Incarna's agents pay BlockRun for model inference one request at a time over x402, cutting the work of adding x402 payment sup…+
<content:encoded><p>When an AI agent runs, it often needs to buy something to finish a task: a model inference, an API response, access web content, or a call to another agent. These purchases are small and frequent, sometimes a fraction of a cent each, and they happen inside the age…+
<p><a href="https://aws.amazon.com/blogs/machine-learning/agents-that-transact-introducing-amazon-bedrock-agentcore-payments-built-with-coinbase-and-stripe/" target="_blank" rel="noopener">Amazon Bedrock AgentCore payments</a> removes that burden. It gives agents a managed way to p…+
<h2 id="the-challenge-paying-for-inference-by-the-request">The challenge: Paying for inference by the request</h2>+
<p>Paying per inference is a high-frequency, low-value pattern. An agent might make hundreds of small purchases in a single session, each worth a fraction of a cent. Card networks weren’t built for sub-cent payments. Building your own rails means solving several hard problems at once. You must…+
<h2 id="what-agentcore-payments-provides">What AgentCore payments provides</h2>+
<p>Amazon Bedrock AgentCore is a platform to build, connect, and optimize agents at scale, with any framework or model. AgentCore payments is a managed capability of Amazon Bedrock AgentCore that builders can use to add payments to their agents in a few lines of code. It handles the payment pr…+
<ul>+
<li><strong>Managed wallets.</strong> Incarna provisions each agent’s wallet using the Coinbase CDP connector. The customer owns it and grants Incarna a delegated authorization to use the wallet.</li>+
<li><strong>Native protocol handling.</strong> When a paid endpoint answers with HTTP 402 (“Payment Required”), the agent uses AgentCore payments to make the payment over x402. It signs the transaction with the configured wallet and returns cryptographic proof to the merchant.<…+
<li><strong>Spending governance at the infrastructure layer.</strong> AgentCore payments enforces limits outside the model, so an agent can’t exceed them even if its prompt is manipulated.</li>+
<li><strong>Settlement you can audit.</strong> Payments settle in a stablecoin. Incarna uses USDC on the Base network, and each transaction is verifiable on-chain.</li>+
</ul>+
<p>The following diagram illustrates the end-to-end architecture. AgentCore runs the agent, BlockRun serves metered inference, and AgentCore payments connects to the customer’s wallet, enforces the spending limit, and signs each payment on behalf of the agent’s Incarna identity.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/21707-1.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/21707-1.png" alt="Architecture diagram: an agen…+
<p class="wp-caption-text">Figure 1: Pay-per-inference architecture. Agent request flows through AgentCore to BlockRun (HTTP 402), with AgentCore payments signing the x402 payment from the agent’s own wallet</p>+
</div>+
<p><em>Source: incarna.io/aws-blockrun-incarna-partnership</em></p>+
<h3 id="prerequisites">Prerequisites</h3>+
<p>The path Incarna followed is open to other teams, and AgentCore payments provisions the pieces for you. You can set them up through a guided conversation with the AgentCore payments skill in the Agent Toolkit for AWS (in Claude Code, Kiro, or Codex). You can also create each one yourself wi…+
<p>Start by storing your Coinbase CDP or Stripe Privy credentials as a payment credential provider, which keeps the secrets in AWS Secrets Manager instead of your code. Create a Payment Manager and connector to coordinate payments against those credentials, and set a default spending limit whi…+
<h2 id="the-flow-buying-and-selling-one-inference">The flow: Buying and selling one inference</h2>+
<p>In this integration, BlockRun is the seller. BlockRun serves metered model inference from a live catalog, and each call is quoted, paid, and settled on its own. The flow for a single inference is straightforward:</p>+
<ol type="1">+
<li>The agent needs a model call. It integrates with BlockRun, which handles provider selection and delivery. No per-provider subscription required.</li>+
<li>BlockRun answers with PaymentRequired challenge and a price for that specific call.</li>+
<li>It opens a payment session and calls <code>ProcessPayment</code>. AgentCore payments checks the quote against the spending limits set for the session and signs the authorization from the agent’s own wallet address.</li>+
<li>The seller verifies the payment signature.</li>+
<li>BlockRun serves the inference and records the charge, a small per-call amount.</li>+
</ol>+
<p>Because settlement happens per request, the agent pays only for what it uses, and a call the agent chooses not to make costs nothing.</p>+
<p>AgentCore payments supports two x402 payment schemes: <code>exact</code> and <code>upto</code>. The <code>exact</code> scheme is usually used when the price is known up front. The <code>upto</code> scheme suits resources with dynamic pricing. …+
<h2 id="keeping-spend-under-control">Keeping spend under control</h2>+
<p>The piece that makes builders comfortable letting an agent move real money is the payment session. A session sets a ceiling, the most the agent can spend, and AgentCore payments enforces that ceiling at the infrastructure layer. The agent’s own code and prompt can’t change it. Each session …+
<h2 id="results">Results</h2>+
<p>Incarna put the pay-per-inference flow into production on Base, with BlockRun serving the sell side and AgentCore payments governing every transaction.</p>+
<blockquote>+
<p><em>“The next generation of AI agents shouldn’t have to choose between better performance and sustainable economics. BlockRun’s open source router gives developers full control over their own model set, while continuously benefiting from BlockRun’s ongoing benchmark-driven routing im…+
<p>— Vicky Fu, Founder, BlockRun</p>+
</blockquote>+
<p>The Incarna team completed the full AgentCore payments integration in three days: one to build and two to test. That was roughly 200 lines of application code, against the two to three months originally scoped. Across the beta, agents have processed over 1,000 payments ranging from $0.001 t…+
<blockquote>+
<p><em>“AgentCore payments covered everything we needed for an agent to pay over x402: a wallet the customer owns, a funding and revocation flow, a spending limit the platform enforces, and signing that handles both versions of x402. We wrote none of it.”</em></p>+
<p>— Justin Zhou, Founder, Incarna</p>+
</blockquote>+
<h2 id="conclusion">Conclusion</h2>+
<p>AgentCore payments gives agents a governed, on-demand way to pay for the services they use. BlockRun and Incarna show how it comes together end to end: an agent that pays for inference, a provider that meters and settles it, and an identity that owns the transaction. If you’re building agen…+
<h2 id="getting-started">Getting started</h2>+
<p>Ready to add pay-per-inference to your agents? Here’s how to start:</p>+
<ol type="1">+
<li><strong>Set up AgentCore payments.</strong> Create a Payment Manager with your wallet connection and spending policies. Connect a Coinbase CDP wallet as your credential provider. See the <a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/payments.html" t…+
<li><strong>Open a payment session with your budget.</strong> Before the agent starts a task, open a session with a spending cap that fits the workload. The agent transacts within that ceiling and the infrastructure enforces it.</li>+
<li><strong>Call a paid endpoint.</strong> Point your agent at an x402-compatible service (like BlockRun). When the endpoint returns HTTP 402, call <code>ProcessPayment</code> with the payment details. AgentCore payments handles the signing and returns proof the agent …+
</ol>+
<p>To explore the code, see the <a href="https://github.com/awslabs/agentcore-samples/tree/main/01-features/08-agents-that-transact" target="_blank" rel="noopener">AgentCore payments samples on GitHub</a>.</p>+
<h2 id="learn-more">Learn more</h2>+
<ul>+
<li><a href="https://aws.amazon.com/blogs/machine-learning/amazon-bedrock-agentcore-payments-is-now-generally-available-enabling-agents-to-transact-safely-and-autonomously-at-scale/" target="_blank" rel="noopener">Amazon Bedrock AgentCore payments is now generally available: Enabling ag…+
<li><a href="https://aws.amazon.com/blogs/machine-learning/technical-deep-dive-agentcore-payments-and-innovation-in-agentic-commerce/" target="_blank" rel="noopener">Technical deep dive: AgentCore payments and innovation in agentic commerce</a></li>+
<li><a href="https://aws.amazon.com/blogs/machine-learning/enable-safe-agentic-payments-with-built-in-guardrails-using-amazon-bedrock-agentcore-payments/" target="_blank" rel="noopener">Enable safe agentic payments with built-in guardrails using AgentCore payments</a></li>+
</ul>+
<h2 id="about-blockrun-and-incarna">About BlockRun and Incarna</h2>+
<p><strong>BlockRun</strong> is a pay-as-you-go inference router that serves model inference over the x402 payment protocol. Agents reach a live catalog of models through a single metered endpoint. Each call is independently quoted, authorized, paid, and settled on Base (USDC). Blo…+
<p><strong>Incarna</strong>, built by SpreadX, gives AI agents a persistent identity that survives sessions, models, and runtimes. Each identity carries its own wallet, email address, social accounts, and an action history that stays attached to one ID across runs. When an agent pa…+
<p style="clear: both"></p>+
<hr style="width: 100%">+
<h2>About the authors</h2>+
<footer>+
<div class="blog-author-box" style="padding-top: 2.0em">+
<div class="blog-author-image" style="margin-right: 1.0em">+
<img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/01/ML-21707-2.jpg" alt="Peter Jiang" width="100" height="133">+
</div>+
<h3 class="lb-h4">Peter Jiang</h3>+
<p style="overflow: hidden">Peter is a Senior Software Developer at AWS, based in Seattle, WA. He is a core member of the engineering team behind the Amazon Bedrock AgentCore payments initiative, which enables AI agents with payments capability. With over 8 years of experience in the financi…+
</div>+
<div class="blog-author-box" style="padding-top: 2.0em">+
<div class="blog-author-image" style="margin-right: 1.0em">+
<img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/01/ML-21707-3.jpg" alt="Guy Bachar" width="100" height="133">+
</div>+
<h3 class="lb-h4">Guy Bachar</h3>+
<p style="overflow: hidden">Guy is a Senior Solutions Architect at AWS, working with fintech and capital markets firms on agentic AI, autonomous commerce, and cloud transformation. He focuses on building systems where AI agents act on behalf of customers across payments, customer experience,…+
</div>+
<div class="blog-author-box" style="padding-top: 2.0em">+
<div class="blog-author-image" style="margin-right: 1.0em">+
<img loading="lazy" class="alignnone size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/01/ML-21707-4.jpg" alt="Chethan Shriyan" width="100" height="133">+
</div>+
<h3 class="lb-h4">Chethan Shriyan</h3>+
<p style="overflow: hidden">Chethan is a Principal Product Manager, Technical at AWS, based in Seattle, WA. He brings nearly 13 years of experience in product and business management, including over 7 years at Amazon. He is passionate about building and delivering technology products that cr…+
</div>+
</footer></content:encoded>+
+
+
+
</item>+
<item>+
<title>Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod</title>+
<link>https://aws.amazon.com/blogs/machine-learning/share-gpu-clusters-across-teams-with-isolation-and-fairness-using-amazon-sagemaker-hyperpod/</link>+
+
<dc:creator><![CDATA[Giuseppe Angelo Porcelli]]></dc:creator>+
<pubDate>Thu, 08 Oct 2026 16:20:04 +0000</pubDate>+
<category><![CDATA[Amazon SageMaker HyperPod]]></category>+
<category><![CDATA[Best Practices]]></category>+
<category><![CDATA[Expert (400)]]></category>+
<guid isPermaLink="false">4b49d02253b76bfce1c012d6f6062cfcdca7f739</guid>+
+
<description>A reference architecture for securely sharing one Amazon SageMaker HyperPod EKS cluster across multiple teams, using AWS IAM Identity Center for authentication, per-team SageMaker Domains and Kubernetes namespaces for isolation, HyperPod Task Governance for fairness, and namespace-…+
<content:encoded><p>Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence. Consider a data science team training larg…+
<p>Amazon SageMaker HyperPod is a purpose-built AI service that simplifies the management of large-scale compute clusters for gen AI workloads. It provides resilient, optimized clusters orchestrated by Amazon Elastic Kubernetes Service (Amazon EKS) or Slurm, so organizations can run distribute…+
<p>In this post, we present a reference architecture for building a multi-tenant environment on Amazon SageMaker HyperPod with EKS. This architecture uses AWS IAM Identity Center for centralized authentication, per-team SageMaker AI domains for a tailored user experience, Kubernetes namespaces…+
<h2 id="architecture-overview">Architecture overview</h2>+
<p>The following diagram illustrates the high-level architecture of a multi-tenant HyperPod EKS deployment. In this example, two teams (Team A and Team B) share a single HyperPod EKS cluster, each operating within their own isolated namespace.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/multi-tenant-hp-eks.jpg" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/multi-tenant-hp-eks.jpg" alt="Layer…+
<p class="wp-caption-text">Figure 1: High-level multi-tenant architecture for two teams sharing one HyperPod EKS cluster</p>+
</div>+
<p>The architecture is structured as a layered flow from left to right, connecting user identity through authorization controls and into isolated workload namespaces on the cluster.</p>+
<h3 id="users-and-authentication">Users and authentication</h3>+
<p>On the far left, individual users from each team (User 1 from Team A, User 2 from Team B) interact with the system through two paths. Both paths authenticate through the AWS IAM Identity Center Portal, which federates with an external identity provider (such as Microsoft Entra ID) shown at …+
<p>The first path is through CLI access. Users authenticate with <code>aws sso login</code>, which redirects them to the Identity Center portal, and then obtain temporary credentials from their team’s permission set to submit tasks directly to the EKS cluster with <code>kubec…+
<p>The second path is through the Identity Center portal directly, where users select the SageMaker Studio application to sign in to their team-specific SageMaker AI domain.</p>+
<p>Each team has a corresponding permission set (TeamA permission set, TeamB permission set) that carries the AWS Identity and Access Management (IAM) policies needed for CLI workflows. Identity Center automatically provisions an IAM role for each permission set, shown in the diagram as <co…+
<h3 id="sagemaker-domains">SageMaker AI domains</h3>+
<p>From the Identity Center portal, users are routed to their team-specific SageMaker AI domain. Each domain (SageMaker AI domain Team A and SageMaker AI domain Team B) provides a dedicated Amazon SageMaker Studio GUI and is configured with a team-specific execution role (<code>TeamA-rol…+
<h3 id="eks-access-control">EKS access control</h3>+
<p>At the EKS boundary, access entries map IAM roles to Kubernetes permissions. The diagram shows access entries for <code>TeamA-role</code> and <code>TeamB-role</code> (the Studio execution roles), which authorize requests originating from the SageMaker Studio GUI. Acc…+
<h3 id="hyperpod-eks-cluster">HyperPod EKS cluster</h3>+
<p>The cluster itself is depicted with two cross-cutting platform layers at the top: HyperPod Observability (for monitoring and dashboards) and HyperPod Task Governance (for compute quota management and scheduling priorities). Below these layers, the cluster is partitioned into Namespace A (Te…+
<h3 id="storage">Storage</h3>+
<p>Beneath the cluster, the architecture includes two storage tiers. The first is a POSIX-compliant file system (Amazon FSx for Lustre or Amazon FSx for OpenZFS) organized into per-team shared directories (<code>/fsx/TeamA</code>, <code>/fsx/TeamB</code>) and per-user h…+
<p>This architecture isolates each team from authentication through authorization to workload execution, while sharing expensive GPU infrastructure efficiently.</p>+
<h2 id="authentication-and-access-control">Authentication and access control</h2>+
<p>The foundation of any multi-tenant system is robust authentication: verifying who users are before they interact with any resource. In this architecture, AWS IAM Identity Center serves as the centralized authentication layer, federating with an external identity provider to manage user iden…+
<h3 id="why-aws-iam-identity-center">Why AWS IAM Identity Center</h3>+
<p>AWS IAM Identity Center (successor to AWS Single Sign-On) provides a single place to manage workforce identities across AWS accounts and applications. For a multi-tenant HyperPod deployment, it offers several key capabilities:</p>+
<ul>+
<li><strong>Centralized identity management</strong> – Rather than maintaining separate user databases per AWS service, Identity Center provides a single source of truth for all user identities and their group memberships.</li>+
<li><strong>Federation with existing identity providers</strong> – Most enterprises already manage their workforce identities in systems like Microsoft Entra ID (formerly Azure AD), Okta, or Ping Identity. Identity Center integrates with these providers, so organizations can reuse…+
<li><strong>Native integration with SageMaker AI</strong> – SageMaker AI domains support Identity Center authentication, so users can sign in to SageMaker Studio through their corporate identity provider with single sign-on (SSO).</li>+
<li><strong>AWS account access</strong> – Identity Center can also grant users access to the underlying AWS account with specific permission sets, which support CLI workflows alongside the Studio GUI experience.</li>+
<li><strong>Required for Amazon Managed Grafana</strong> – Amazon Managed Grafana uses Identity Center as its authentication mechanism for workforce users, making it the natural choice when teams also need access to observability dashboards for monitoring their workloads.</li&g…+
</ul>+
<p>Learn more: <a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/what-is.html" target="_blank" rel="noopener">What is IAM Identity Center</a></p>+
<h3 id="configuring-identity-center-with-an-external-identity-provider">Configuring Identity Center with an external identity provider</h3>+
<p>In this reference architecture, we use Microsoft Entra ID as the external identity provider, though the same pattern applies to most standard Security Assertion Markup Language (SAML) 2.0 providers.</p>+
<p>The configuration involves:</p>+
<ol type="1">+
<li><strong>Group structure in the identity provider</strong> – In Entra ID, create groups that correspond to your organizational teams. In our example, we define three groups: <code>TeamA</code>, <code>TeamB</code>, and <code>Admin</code>. Each…+
<li><strong>SCIM provisioning</strong> – Enable SCIM (System for Cross-domain Identity Management) synchronization between Entra ID and AWS IAM Identity Center. SCIM provides automatic provisioning and de-provisioning of users and groups. When a new user is added to the <code&g…+
<li><strong>SAML-based authentication</strong> – Configure SAML 2.0 federation so that when users authenticate, they do so against Entra ID. Identity Center acts as the service provider, trusting the assertions from your Entra ID tenant.</li>+
</ol>+
<p>With this configuration, you manage team membership (which drives all downstream authorization decisions) in your existing corporate directory, and it propagates to AWS automatically.</p>+
<p>The following image shows an example of how organizational teams can be represented in Microsoft Entra ID, with dedicated groups for <code>TeamA</code>, <code>TeamB</code>, and <code>Admin</code>.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/entra_id_groups.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/entra_id_groups.png" alt="Microsoft Ent…+
<p class="wp-caption-text">Figure 2: Organizational teams represented as groups in Microsoft Entra ID</p>+
</div>+
<p>Then, the following image shows the corresponding groups in AWS IAM Identity Center, automatically provisioned from Entra ID through SCIM synchronization.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/iam_idc_groups.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/iam_idc_groups.png" alt="AWS IAM Identit…+
<p class="wp-caption-text">Figure 3: Corresponding groups in AWS IAM Identity Center, provisioned through SCIM</p>+
</div>+
<p>Learn more: <a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/manage-your-identity-source-idp.html" target="_blank" rel="noopener">Connect an external identity provider</a> · <a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/scim-profile-saml…+
<h2 id="authorization">Authorization</h2>+
<p>With authentication established, the next layer is authorization: controlling what actions each team can perform across AWS services and the Kubernetes cluster. Authorization in this architecture operates at two levels: IAM for service-level access, and Kubernetes RBAC for cluster-level acc…+
<h3 id="per-team-iam-roles">Per-team IAM roles</h3>+
<p>Each team requires a dedicated IAM role that encapsulates the AWS level permissions needed for their AI and machine learning (ML) workflows. These roles serve as the SageMaker AI domain execution role and define what AWS services the team can access.</p>+
<p>A typical team IAM role should include policies granting access to:</p>+
<ul>+
<li><strong>Amazon SageMaker AI</strong> – For managing HyperPod clusters, MLflow tracking servers, and other SageMaker AI resources through the SageMaker AI API.</li>+
<li><strong>Amazon S3</strong> – For reading training datasets and writing model artifacts, checkpoints, and logs. Scope these permissions to team-specific bucket prefixes.</li>+
<li><strong>Amazon CloudWatch</strong> – For viewing logs and metrics related to the team’s workloads.</li>+
<li><strong>Amazon EKS</strong> – Specifically, the <code>eks:AccessKubernetesApi</code> and <code>eks:MutateViaKubernetesApi</code> permissions, which the SageMaker Studio GUI needs to make Kubernetes API calls on behalf of the user (for example, listing S…+
</ul>+
<p>The trust policy on each IAM role must include <code>sagemaker.amazonaws.com</code> as a trusted principal, so SageMaker AI can assume the role on behalf of users when they operate through Studio. If you plan to reuse the same execution role as an EKS Pod Identity association fo…+
<p>Learn more: <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-roles.html" target="_blank" rel="noopener">How to use SageMaker AI execution roles</a></p>+
<h3 id="aws-account-access-through-identity-center">AWS account access through Identity Center</h3>+
<p>Beyond SageMaker Studio, teams often need direct AWS account access for CLI operations such as running <code>kubectl</code> commands, scripting workflows, or accessing resources programmatically. Identity Center permission sets provide this capability.</p>+
<p>For the <strong>Admin</strong> group, assign a permission set with administrative access as required by your company’s policies, granting the necessary account access for cluster management and administrative operations.</p>+
<p>For <strong>Team A</strong> and <strong>Team B</strong>, create permission sets with inline or managed policies that grant the permissions needed for CLI workflows directly. A typical team permission set includes permissions for <code>eks:AccessKubernetesApi<…+
<p>Users retrieve temporary credentials through the AWS Command Line Interface (AWS CLI) using <code>aws sso login</code>, which they can then use to configure <code>kubectl</code> for direct interaction with the EKS cluster.</p>+
<p>The following image shows the per-team permission sets in AWS IAM Identity Center, providing scoped AWS account access for CLI workflows such as running <code>kubectl</code> and <code>aws sso login</code> against the EKS cluster.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/iam_idc_permissionsets.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/iam_idc_permissionsets.png" alt=…+
<p class="wp-caption-text">Figure 4: Per-team permission sets in AWS IAM Identity Center for CLI workflows</p>+
</div>+
<p>Learn more: <a href="https://docs.aws.amazon.com/singlesignon/latest/userguide/permissionsetsconcept.html" target="_blank" rel="noopener">Manage AWS accounts with permission sets</a></p>+
<h3 id="configuring-the-aws-cli">Configuring the AWS CLI</h3>+
<p>Team members configure the AWS CLI to authenticate through Identity Center by running <code>aws configure sso</code>. This creates profiles in <code>~/.aws/config</code> that reference the appropriate Identity Center session and permission set. Each team member uses …+
<p>The resulting configuration defines a shared <code>sso-session</code> block for the Identity Center portal and one named profile per team, each pointing at that team’s permission set. Team members then run <code>aws sso login --profile &lt;team&gt;</code> to …+
<div class="hide-language">+
<pre><code class="language-ini">[sso-session my-sso]+
sso_start_url = https://d-xxxxxxxxxx.awsapps.com/start+
sso_region = us-west-2+
sso_registration_scopes = sso:account:access+
+
[default]+
sso_session = my-sso+
sso_account_id = 123456789012+
sso_role_name = OpsAdmin+
region = us-west-2+
+
[profile team-a]+
sso_session = my-sso+
sso_account_id = 123456789012+
sso_role_name = TeamA-permission-set+
region = us-west-2+
+
[profile team-b]+
sso_session = my-sso+
sso_account_id = 123456789012+
sso_role_name = TeamB-permission-set+
region = us-west-2</code></pre>+
</div>+
<p>Learn more: <a href="https://docs.aws.amazon.com/cli/latest/userguide/cli-configure-sso.html" target="_blank" rel="noopener">Configuring IAM Identity Center authentication with the AWS CLI</a></p>+
<h2 id="sagemaker-domains-1">SageMaker AI domains</h2>+
<p>SageMaker AI domains provide the workspace boundary for each team, offering a tailored user experience, pre-configured execution roles, and built-in integration with Identity Center authentication.</p>+
<h3 id="why-sagemaker-domains">Why SageMaker AI domains</h3>+
<p>Using one SageMaker AI domain per team is a well-established pattern for organizing multi-team environments. This approach offers several advantages:</p>+
<ul>+
<li><strong>Established multi-team pattern</strong> – AWS has documented this approach extensively for separating lines of business or teams with multiple domains, making it a proven and supported configuration.</li>+
<li><strong>Native Identity Center authentication</strong> – Each domain can be configured with Identity Center authentication, meaning users sign in once through their corporate identity provider and land directly in their team’s Studio environment.</li>+
<li><strong>Built-in team configuration</strong> – Domains already provide mechanisms to specify configurations for users and teams without requiring additional custom entities. For example, settings like team execution role can be specified at domain level, and overridden at user…+
<li><strong>Navigation customization</strong> – With domain settings, administrators can hide navigation items that are not relevant to the team’s workflows, presenting a focused interface tailored to HyperPod use cases.</li>+
</ul>+
<p>Learn more: <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sm-domain.html" target="_blank" rel="noopener">SageMaker AI domain entities and statuses</a> · <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/domain-multiple.html" target="_blank" rel="noopener">…+
<h3 id="setting-up-per-team-domains">Setting up per-team domains</h3>+
<p>Create one SageMaker AI domain per team with Identity Center authentication. In our example, we create <code>TeamA-domain</code> and <code>TeamB-domain</code>. Each domain is configured as follows:</p>+
<ol type="1">+
<li><strong>Default execution role</strong> – Set the domain’s default execution role to the team-specific IAM role created in the authorization step. All actions performed through Studio inherit the appropriate permissions as a result.</li>+
<li><strong>Identity Center group assignment</strong> – Add the corresponding Identity Center group (for example, the <code>TeamA</code> group) to the domain. This activates the SageMaker Studio application for all members of that group, granting them access to the Stu…+
<li><strong>Application assignment verification</strong> – After configuring group access, review the application assignment in Identity Center to confirm that the correct groups are mapped to the correct domains.</li>+
<li><strong>Navigation customization</strong> – Configure the default navigation settings for each domain to present only the relevant capabilities. For example, you might hide items not related to HyperPod workflows, providing a streamlined <em>HyperPod-focused</em> u…+
</ol>+
<p>The following image shows the SageMaker AI console with one domain per team (<code>TeamA-domain</code> and <code>TeamB-domain</code>), each providing an isolated workspace boundary.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/domains-list.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/domains-list.png" alt="SageMaker console l…+
<p class="wp-caption-text">Figure 5: One SageMaker Domain per team in the SageMaker console</p>+
</div>+
<p>Then the following image shows the details of <code>TeamA-domain</code>, including the assigned Identity Center groups.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/teama-domain.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/teama-domain.png" alt="TeamA-domain detail…+
<p class="wp-caption-text">Figure 6: TeamA-domain configuration with its assigned Identity Center groups</p>+
</div>+
<h2 id="hyperpod-eks-cluster-configuration">HyperPod EKS cluster configuration</h2>+
<p>The HyperPod EKS cluster is where workloads are executed. Multi-tenancy at the cluster level is achieved through Kubernetes namespaces for isolation and EKS access entries for authorization.</p>+
<h3 id="namespace-isolation">Namespace isolation</h3>+
<p>Create a dedicated Kubernetes namespace for each team, for example <code>hyperpod-ns-team-a</code> and <code>hyperpod-ns-team-b</code>. Namespaces provide a logical boundary within the cluster, isolating each team’s workloads (Spaces, training jobs, inference endpoin…+
<blockquote>+
<p><strong>Note: Namespaces are an isolation boundary, not a hard security boundary.</strong> This architecture targets <em>multi-team within a single organization</em>: teams that share a cluster under a common administrative domain and a baseline of mutual trust. It …+
<p>Namespaces, RBAC, and quotas prevent <em>accidental</em> interference (teams overwriting each other’s resources or exceeding their compute allocation) but aren’t a defense against a determined malicious tenant: namespaced pods share the same nodes and kernel, and cluster-scoped…+
<p>For untrusted tenants or strict regulatory isolation, use stronger boundaries such as separate clusters or accounts, dedicated node pools, and runtime sandboxing. For the multi-team scenario here, namespace isolation combined with RBAC, Task Governance quotas, and the POSIX identity contro…+
</blockquote>+
<p>Namespaces can be created manually with <code>kubectl create namespace</code> or provisioned automatically through HyperPod Task Governance, which manages namespaces as part of its quota and scheduling configuration.</p>+
<p>The following image shows the cluster namespaces (managed with HyperPod Task Governance), with one dedicated namespace per team (<code>hyperpod-ns-team-a</code> and <code>hyperpod-ns-team-b</code>) providing workload isolation.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/namespaces.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/namespaces.png" alt="Cluster namespaces mana…+
<p class="wp-caption-text">Figure 7: Dedicated Kubernetes namespace per team for workload isolation</p>+
</div>+
<p>Learn more: <a href="https://kubernetes.io/docs/concepts/overview/working-with-objects/namespaces/" target="_blank" rel="noopener">Kubernetes namespaces</a></p>+
<h3 id="network-isolation">Network isolation</h3>+
<p>Namespaces don’t restrict network traffic. By default, Kubernetes networking is flat: every pod can reach every other pod across all namespaces. As a result, a pod in <code>hyperpod-ns-team-a</code> can open a connection to a pod in <code>hyperpod-ns-team-b</code> un…+
<p>The recommended pattern is <em>default-deny</em> per namespace: start by denying all ingress (and optionally egress), then explicitly allow the traffic each team needs, typically intra-namespace communication plus required egress such as DNS, storage endpoints, and AWS APIs. The…+
<div class="hide-language">+
<pre><code class="language-yaml"># 1. Default-deny all ingress in the team's namespace.+
apiVersion: networking.k8s.io/v1+
kind: NetworkPolicy+
metadata:+
name: default-deny-ingress+
namespace: hyperpod-ns-team-a+
spec:+
podSelector: {} # applies to all pods in the namespace+
policyTypes:+
- Ingress+
---+
# 2. Allow ingress only from pods within the same namespace.+
apiVersion: networking.k8s.io/v1+
kind: NetworkPolicy+
metadata:+
name: allow-same-namespace+
namespace: hyperpod-ns-team-a+
spec:+
podSelector: {}+
policyTypes:+
- Ingress+
ingress:+
- from:+
- podSelector: {} # any pod in this namespace</code></pre>+
</div>+
<p><code>NetworkPolicy</code> enforcement depends on a Container Network Interface (CNI) that supports it. On EKS, you can enable network policy support in the Amazon Virtual Private Cloud (Amazon VPC) CNI.</p>+
<p>As with namespaces, <code>NetworkPolicies</code> reduce <em>accidental</em> cross-team reachability and shrink the scope, but they are not by themselves an adversarial security boundary on shared nodes. For stronger separation, consider dedicated node pools per team …+
<p>Learn more: <a href="https://kubernetes.io/docs/concepts/services-networking/network-policies/" target="_blank" rel="noopener">Kubernetes network policies</a> · <a href="https://docs.aws.amazon.com/eks/latest/userguide/cni-network-policy.html" target="_blank" rel="noopener"&g…+
<h3 id="eks-access-entries">EKS access entries</h3>+
<p>EKS access entries connect IAM principals to Kubernetes RBAC permissions. For each team, create two access entries:</p>+
<ul>+
<li><strong>Studio access entry</strong> – The IAM principal is the team’s SageMaker AI domain execution role. This entry is used when actions originate from the SageMaker Studio GUI.</li>+
<li><strong>CLI access entry</strong> – The IAM principal is the SSO-provisioned role created by Identity Center for the team’s permission set (following the pattern <code>AWSReservedSSO_&lt;permission-set-name&gt;_&lt;unique-id&gt;</code>). This entry …+
</ul>+
<p>Both entries are scoped to the team’s namespace with managed or custom Kubernetes policies. For example, both Team A entries grant permissions only within <code>hyperpod-ns-team-a</code>. The two entries can carry different RBAC policies if you want. For instance, the CLI entry …+
<p>With this scoping, whether access originates from Studio or from the CLI, users can only interact with resources in their own namespace. Attempting to list or modify resources in another team’s namespace results in a Kubernetes <code>Forbidden</code> error.</p>+
<p>For more advanced scenarios, you can use Kubernetes groups in the access entry to map users to custom ClusterRoles or Roles that provide fine-grained permissions beyond the standard managed policies.</p>+
<p>The following image shows an EKS access entry for Team B’s role, scoped down to the hyperpod-ns-team-b namespace, so its permissions apply only within Team B’s namespace.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/eks-access-entry-1.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/05/eks-access-entry-1.png" alt="EKS acc…+
<p class="wp-caption-text">Figure 8: EKS access entry scoped to Team B’s namespace</p>+
</div>+
<p>Learn more: <a href="https://docs.aws.amazon.com/eks/latest/userguide/access-entries.html" target="_blank" rel="noopener">Grant IAM users access to Kubernetes with EKS access entries</a></p>+
<h3 id="hyperpod-task-governance">HyperPod Task Governance</h3>+
<p>When Task Governance is enabled on the cluster, it provides an additional layer of resource management:</p>+
<ul>+
<li><strong>Compute quotas</strong> – Define how much GPU and CPU capacity each team can consume. This prevents a single team from monopolizing shared hardware during training runs.</li>+
<li><strong>Priorities</strong> – Assign scheduling priorities to each team or workload type, allowing critical production inference workloads to preempt experimental training jobs when resources are constrained.</li>+
<li><strong>Fair scheduling</strong> – With Task Governance, when multiple teams are competing for resources, allocation follows the configured policies rather than a first-come-first-served model.</li>+
</ul>+
<p>Configure Task Governance with appropriate quotas and priorities per team namespace, balancing between guaranteed minimum allocations and burst capacity for bursty workloads.</p>+
<p>The following image shows the Task Governance compute allocations for the two teams, with each team’s namespace assigned its own quota of cluster compute capacity.</p>+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/task-governance.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/10/02/task-governance.png" alt="HyperPod Task…+
<p class="wp-caption-text">Figure 9: Task Governance compute allocations per team namespace</p>+
</div>+
<p>Learn more: <a href="https://docs.aws.amazon.com/sagemaker/latest/dg/sagemaker-hyperpod-eks-operate-console-ui-governance.html" target="_blank" rel="noopener">SageMaker HyperPod task governance</a></p>+
<h2 id="storage-1">Storage</h2>+
<p>Storage is a key component of artificial intelligence and machine learning (AI/ML) environments shared across teams. Teams need high-performance file systems for training data, checkpoints, and model artifacts, while maintaining appropriate access boundaries between teams.</p>+
<h3 id="posix-compliant-file-systems">POSIX-compliant file systems</h3>+
<p>For workloads that require a shared, high-performance POSIX file system (common for distributed training where multiple nodes read the same dataset or write checkpoints), consider the following options:</p>+
<ul>+
<li><strong>Amazon FSx for Lustre</strong> – Provides high-throughput, low-latency parallel file system access, ideal for large-scale training workloads that need to read large datasets at high speed.</li>+
<li><strong>Amazon FSx for OpenZFS</strong> – Offers a general-purpose file system with strong POSIX semantics, snapshots, and compression. Well-suited for workloads that need traditional file system features alongside high performance.</li>+
<li><strong>Amazon Elastic File System (Amazon EFS)</strong> – Provides fully managed, elastic Network File System (NFS) storage. EFS also supports access points, which can simplify per-team directory isolation by mapping different mount points to different directories with enforc…+
</ul>+
<p>The storage layout typically follows this structure:</p>+
<ul>+
<li><strong>Per-team shared directories</strong> – Each team has a shared directory (for example, <code>/fsx/TeamA</code>, <code>/fsx/TeamB</code>) for datasets, models, and artifacts that all team members need to access.</li>+
<li><strong>Per-user home directories</strong> – Each user has a personal home directory (for example, <code>/home/User1</code>, <code>/home/User2</code>) for individual work, experiments, and notebooks.</li>+
</ul>+
<p>The POSIX permission model on these file systems relies on UIDs, GIDs, and supplemental groups to enforce access boundaries. These POSIX identities should then be propagated to the pod security context when a user launches a HyperPod Space or submits a training job, so that file system acce…+
<div class="hide-language">+
<pre><code class="language-python"># 1. Extract the caller's session identity from the admission request.+
# On EKS, requests from IAM-assumed roles (including IAM Identity Center)+
# surface the STS session name in userInfo.extra["sessionName"]. Its format+
# depends on how the session is created (e.g., an SSO short name, an email,+
# or a role-session-name); align your identity mapping with this value.+
def extract_username(admission_request):+
extra = admission_request["userInfo"]["extra"]+
...+
return extra["sessionName"][0]+
+
# 2. Look up the POSIX identity from a mapping table (e.g. DynamoDB).+
def lookup_posix_identity(username):+
item = posix_table.get_item(Key={"username": username})["Item"]+
...+
return {+
"uid": int(item["uid"]),+
"gid": int(item["gid"]),+
"supplementalGroups": [int(g) for g in item["supplementalGroups"]],+
}+
+
# 3. Patch the Pod security context with the resolved POSIX identity.+
def build_security_context_patch(pod, posix):+
...+
return [{+
"op": "add",+
"path": "/spec/securityContext",+
"value": {+
"runAsUser": posix["uid"],+
"runAsGroup": posix["gid"],+
"fsGroup": posix["gid"],+
"supplementalGroups": posix["supplementalGroups"],+
},+
}]</code></pre>+
</div>+
<p>Learn more: <a href="https://docs.aws.amazon.com/fsx/latest/LustreGuide/what-is.html" target="_blank" rel="noopener">FSx for Lustre</a> · <a href="https://docs.aws.amazon.com/fsx/latest/OpenZFSGuide/what-is-fsx.html" target="_blank" rel="noopener">FSx for OpenZFS</a>…+
<h3 id="amazon-s3-storage">Amazon S3 storage</h3>+
<p>For object storage, access to S3 buckets is governed by the team’s IAM execution role. You can create per-team buckets or use a shared bucket with per-team prefixes, relying on IAM policies to enforce isolation. Pods within the cluster need appropriate service accounts configured with IAM R…+
<p>Learn more: <a href="https://docs.aws.amazon.com/eks/latest/userguide/iam-roles-for-service-accounts.html" target="_blank" rel="noopener">IAM roles for service accounts (IRSA)</a> · <a href="https://docs.aws.amazon.com/eks/latest/userguide/pod-identities.html" target="_blank"…+
<h2 id="hyperpod-spaces">HyperPod Spaces</h2>+
<p>HyperPod Spaces provide interactive development environments (IDEs) running directly on cluster nodes. On a shared cluster, Spaces must be properly scoped to each team’s namespace and configured with appropriate resource templates.</p>Diff display stops at 400 lines. The line counts above are from the whole diff. 73 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.