llm-catalog-archive

Change

a595fbf

a595fbfd2ddc4e0340733523e8f2b9244ee4932f · commit on GitHub

aws-blog-feed: changed (734012 bytes, HTTP 200)

raw/aws-blog-feed/response.xml modified

Lines added
+2,971
Lines removed
-3,118
Stored bytes at this commit
734,012
Timestamp
observed
Raw artifact at this commit
raw/aws-blog-feed/response.xml
Recorded headers
observed_at2026-09-19T04:38:08.708Z
origin_datenull
status200
final URLhttps://aws.amazon.com/blogs/machine-learning/feed/
etagnull
last-modifiedSat, 19 Sep 2026 00:23:02 GMT
dateSat, 19 Sep 2026 04:38:08 GMT
agenull
cache-controlnull
cf-cache-statusnull
content-encodingnull
content-lengthnull
@@@ -5,7 +5,7 @@
<atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/>
<link>https://aws.amazon.com/blogs/machine-learning/</link>
<description>Official Machine Learning Blog of Amazon Web Services</description>
- <lastBuildDate>Thu, 17 Sep 2026 18:02:01 +0000</lastBuildDate>
+ <lastBuildDate>Fri, 18 Sep 2026 21:17:31 +0000</lastBuildDate>
<language>en-US</language>
<sy:updatePeriod>
hourly </sy:updatePeriod>
@@@ -13,597 +13,1065 @@
1 </sy:updateFrequency>
<item>
- <title>Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent</title>
- <link>https://aws.amazon.com/blogs/machine-learning/reduce-time-to-hire-for-quality-candidates-with-ai-powered-amazon-connect-talent/</link>
-
+ <title>Amazon SageMaker Inference: 2026 year-to-date launches in review</title>
+ <link>https://aws.amazon.com/blogs/machine-learning/amazon-sagemaker-inference-2026-year-to-date-launches-in-review/</link>
- <dc:creator><![CDATA[Ayesha Borker]]></dc:creator>
- <pubDate>Thu, 17 Sep 2026 17:55:20 +0000</pubDate>
- <category><![CDATA[Amazon Connect]]></category>
+ <dc:creator><![CDATA[Kareem Syed-Mohammed]]></dc:creator>
+ <pubDate>Fri, 18 Sep 2026 20:52:14 +0000</pubDate>
+ <category><![CDATA[Amazon SageMaker AI]]></category>
<category><![CDATA[Announcements]]></category>
- <guid isPermaLink="false">a6559b315a35bac095e4d57dabbff35e09f15496</guid>
-
- <description>Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while providing applicants w
- <content:encoded>&lt;p&gt;Hiring at scale in industries such as retail, logistics, hospitality, and others has its fair share of challenges. Recruiting teams are expected to fill hundreds of roles within tight timelines, often with limited capacity and with tools that weren’t designed to s
-&lt;p&gt;Today, we’re launching &lt;a href="https://aws.amazon.com/products/connect/talent/" target="_blank" rel="noopener"&gt;Amazon Connect Talent,&lt;/a&gt; an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, a
-&lt;p&gt;Recruiters configure the &lt;a href="https://docs.aws.amazon.com/talent/latest/userguide/evaluations.html" target="_blank" rel="noopener"&gt;evaluation criteria&lt;/a&gt;, assessments, and interview questions based on job requirements. &lt;a href="https://docs.aws.amazon.com/talent/latest/u
-&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-1.png" target="_blank" rel="noopener"&gt;&lt;img class="alignnone wp-image-139627 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b5
-&lt;p&gt;&lt;em&gt;Amazon Connect Talent lets recruiters configure AI-led interviews and candidate assessments in minutes, tailored to each role.&lt;/em&gt;&lt;/p&gt;
-&lt;p&gt;&lt;strong&gt; &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-2-1.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" class="alignnone wp-image-139629 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836c
-&lt;p&gt;&lt;em&gt;Candidates experience a consistent, structured AI-led interview — available whenever they’re ready, day or night.&lt;/em&gt;&lt;/p&gt;
-&lt;h2&gt;&lt;strong&gt;Informed by decades of Amazon’s hiring science&lt;/strong&gt;&lt;/h2&gt;
-&lt;p&gt;Amazon is one of the world’s largest employers, and Amazon Connect Talent puts decades of Amazon’s hiring science to work for organizations, with a highly configurable solution that adapts to specific hiring requirements. With consistent, evidence-based assessments applied to every candidat
-&lt;h3&gt;&lt;strong&gt;Reducing human preconceptions in hiring&lt;/strong&gt;&lt;/h3&gt;
-&lt;p&gt;Connect Talent’s AI focuses exclusively on measuring a candidate’s job-related competencies, testing abilities including problem-solving, logic, listening, and role-specific capabilities. All candidate data is anonymized during AI evaluation, removing factors that can introduce unconscious
-&lt;h3&gt;&lt;strong&gt;Transparency and consistent evaluation standards&lt;/strong&gt;&lt;/h3&gt;
-&lt;p&gt;Connect Talent communicates clearly to both recruiters and candidates about what data is collected, how it is used, and what is not collected. Candidates are informed of what to expect before they proceed to take the evaluation. Each competency is scored against a rubric that defines what a
-&lt;h3&gt;&lt;strong&gt;Human in the loop: recruiter control at every step&lt;/strong&gt;&lt;/h3&gt;
-&lt;p&gt;Recruiters maintain final decision authority over every hire. Connect Talent gives them scored candidate summaries with competency breakdowns, complete interview transcripts, comparative analytics, and clear reasoning behind every score, all in a single view instead of pieced together from
-&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-3.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" class="alignnone wp-image-139630 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f
-&lt;p&gt;&lt;em&gt;Amazon Connect Talent lets recruiters configure AI-led interviews and candidate assessments in minutes, tailored to each role.&lt;/em&gt;&lt;/p&gt;
-&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-4.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" class="alignnone wp-image-139631 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f
-&lt;p&gt;&lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-5.png" target="_blank" rel="noopener"&gt;&lt;img loading="lazy" class="alignnone wp-image-139632 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f
-&lt;p&gt;&lt;em&gt;Amazon Connect Talent lets recruiters review candidate responses and evaluation scores with details and full transparency&lt;/em&gt;&lt;/p&gt;
-&lt;h2&gt;&lt;strong&gt;Enterprise-grade security built on AWS infrastructure&lt;/strong&gt;&lt;/h2&gt;
-&lt;p&gt;Hiring data is sensitive. Candidate information deserves the same level of protection&amp;nbsp;you’d&amp;nbsp;expect for financial records or health information. Amazon Connect Talent delivers enterprise-grade security built on AWS infrastructure:&amp;nbsp;the same foundation trusted by ban
-&lt;p&gt;Built-in security controls help protect candidate data, with rigorous measures to meet your requirements. Access is controlled, auditable, and configurable. When your candidates share their information, they can trust&amp;nbsp;it’s&amp;nbsp;protected.&lt;/p&gt;
-&lt;ul&gt;
- &lt;li&gt;&lt;strong&gt;Integrity monitoring and fraud protection:&amp;nbsp;&lt;/strong&gt;Amazon Connect Talent uses&amp;nbsp;text-based&amp;nbsp;analysis to flag unnatural cadence, filler words, pauses, and response latency.&amp;nbsp;Human review is mandatory for every flag. No candidate is disqu
- &lt;li&gt;&lt;strong&gt;Audit trails and explainability:&amp;nbsp;&lt;/strong&gt;Every&amp;nbsp;candidate&amp;nbsp;interaction is logged with a complete audit trail and clear job-related evaluation reasoning. The system includes ongoing monitoring and tracking to help your organization&amp;nbsp;dem
-&lt;/ul&gt;
-&lt;h2&gt;&lt;strong&gt;Improved business outcomes&lt;/strong&gt;&lt;/h2&gt;
-&lt;p&gt;Amazon Connect Talent improves business outcomes by getting recruiters out of the workflow management cycle and empowering them with faster decision-making. Candidates get a faster and more flexible hiring experience. Organizations fill roles before lost revenue, rising costs, and competiti
-&lt;p&gt;Your business moves fast and now your hiring can too.&lt;/p&gt;
-&lt;p&gt;&lt;iframe loading="lazy" title="Amazon Connect Talent Introduction" width="500" height="281" src="https://www.youtube-nocookie.com/embed/0X1PH7ZwvRo?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" r
-&lt;p&gt;&lt;a href="https://aws.amazon.com/products/connect/talent/" target="_blank" rel="noopener noreferrer"&gt;Learn more about Amazon Connect Talent&lt;/a&gt; and discover how AI-powered hiring can transform your talent acquisition strategy.&lt;/p&gt;
-&lt;h2&gt;About the authors&lt;/h2&gt;
-&lt;footer&gt;
- &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
- &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
- &lt;img loading="lazy" class="alignnone wp-image-139633 size-thumbnail" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/Ayesha-100x133.jpg" alt="" width="100" height="133"&gt;
- &lt;/div&gt;
- &lt;h3 class="lb-h4"&gt;Ayesha Borker&lt;/h3&gt;
- &lt;p style="overflow: hidden"&gt;Ayesha is a Principal Solutions Architect, Applied AI at AWS. With over a decade of experience at the intersection of human and AI collaboration specializing in customer experience. She simplifies the complexity of AI, helping organizations stay focused on the out
- &lt;/div&gt;
- &lt;div class="blog-author-box" style="padding-top: 2.0em"&gt;
- &lt;div class="blog-author-image" style="margin-right: 1.0em"&gt;
- &lt;img loading="lazy" class="alignnone size-thumbnail wp-image-139635" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/Kate-100x133.jpg" alt="" width="100" height="133"&gt;
- &lt;/div&gt;
- &lt;h3 class="lb-h4"&gt;Kate Totaro&lt;/h3&gt;
- &lt;p style="overflow: hidden"&gt;Kate is the Principal Product Manager for Amazon Connect Talent, AWS’s AI-powered hiring service. She brings 20 years of experience building global businesses, including more than 15 years at Amazon and AWS across product, business, and operations. At Amazon, she
- &lt;/div&gt;
-&lt;/footer&gt;</content:encoded>
-
-
-
-
-
- </item>
- <item>
- <title>Selecting a vector store for Amazon Bedrock Knowledge Bases</title>
- <link>https://aws.amazon.com/blogs/machine-learning/selecting-a-vector-store-for-amazon-bedrock-knowledge-bases/</link>
-
- <dc:creator><![CDATA[Deepak Dalakoti]]></dc:creator>
- <pubDate>Thu, 17 Sep 2026 15:53:13 +0000</pubDate>
- <category><![CDATA[Amazon Bedrock Knowledge Bases]]></category>
- <category><![CDATA[Best Practices]]></category>
<category><![CDATA[Intermediate (200)]]></category>
- <guid isPermaLink="false">e129d8e2814c4c6f57fdc5feec60ab4d06f5e04c</guid>
+ <guid isPermaLink="false">2ee3ed6a4b11eb69260daaa48214c0e6564f2a28</guid>
- <description>Choosing the right vector store for your Amazon Bedrock Knowledge Bases RAG application affects performance and cost. This post compares Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors across three RAG use cases, with benchmarks and a practi
- <content:encoded>&lt;p&gt;When building a Retrieval Augmented Generation (RAG) solution with &lt;a href="https://aws.amazon.com/bedrock/knowledge-bases/" target="_blank" rel="noopener"&gt;Amazon Bedrock Knowledge Bases&lt;/a&gt;, selecting the right vector store impacts performance and cos
-&lt;p&gt;For broader guidance across all AWS vector solutions, see &lt;a href="https://aws.amazon.com/blogs/machine-learning/aws-vector-solutions-build-agentic-ai-where-your-data-lives/" target="_blank" rel="noopener"&gt;AWS vector solutions: Build agentic AI where your data lives&lt;/a&gt;. For the
-&lt;h2 id="how-vector-databases-fit-into-rag-solutions"&gt;How vector databases fit into RAG solutions&lt;/h2&gt;
-&lt;p&gt;A RAG architecture combines the capabilities of large language models (LLMs) with information retrieval systems to generate more accurate, up-to-date, and contextually relevant responses. It is based on the mathematical concept of a &lt;em&gt;vector&lt;/em&gt;, where the text is translated
-&lt;p&gt;When a user submits a query, it is converted into a vector embedding using an embedding model. The vector database, where document content has been pre-processed, chunked, and stored as vector embeddings, performs a similarity search to find the chunks whose embeddings are most similar to t
-&lt;p&gt;The vector database serves as the key bridge between raw information and contextual understanding. It transforms unstructured data into a searchable, semantically meaningful knowledge space that helps large language models deliver more precise and relevant responses. Vector databases achiev
-&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
- &lt;a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-19758-1.jpeg" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-19758-1.jpeg" alt="Diagram of a RAG arch
- &lt;p class="wp-caption-text"&gt;Figure 1: Retrieval Augmented Generation (RAG) architecture, where documents are chunked, embedded, and stored in a vector database during ingestion, and at query time the query is embedded, similar chunks are retrieved, and passed to the LLM as context for response
-&lt;/div&gt;
-&lt;h2 id="vector-store-backends-for-amazon-bedrock-knowledge-bases"&gt;Vector store backends for Amazon Bedrock Knowledge Bases&lt;/h2&gt;
-&lt;p&gt;Amazon Bedrock Knowledge Bases with a customer-managed (unmanaged) configuration supports three vector store backends. For the full AWS vector portfolio covering six services, see &lt;a href="https://aws.amazon.com/blogs/machine-learning/aws-vector-solutions-build-agentic-ai-where-your-data
-&lt;p&gt;&lt;a href="https://aws.amazon.com/opensearch-service/" target="_blank" rel="noopener"&gt;Amazon OpenSearch Service&lt;/a&gt; provides high-speed results from data held in memory. It supports high-dimensional vector embeddings with both managed cluster and serverless deployment options, and
-&lt;p&gt;&lt;a href="https://aws.amazon.com/rds/aurora/" target="_blank" rel="noopener"&gt;Amazon Aurora PostgreSQL with pgvector&lt;/a&gt; combines the high-performance relational database capabilities of Amazon Aurora with pgvector’s vector similarity search functionality. It supports multiple ind
-&lt;p&gt;&lt;a href="https://aws.amazon.com/s3/features/vectors/" target="_blank" rel="noopener"&gt;Amazon S3 Vectors&lt;/a&gt; is the AWS cloud object storage service with native vector support, designed for cost-effective storage and querying of vector embeddings at scale. It provides sub-second q
-&lt;p&gt;To understand how these options perform in practice, let’s examine three distinct RAG use cases, each with different latency, cost, and search requirements, and see which vector database is the best fit for each.&lt;/p&gt;
-&lt;h2 id="use-case-1-product-catalog-search"&gt;Use case 1: Product catalog search&lt;/h2&gt;
-&lt;p&gt;Ecommerce platforms face the challenge of helping customers find exactly what they’re looking for among thousands of products. An effective product search tool must understand natural language queries and scale to handle thousands of concurrent queries during peak shopping periods while mai
-&lt;h3 id="why-amazon-opensearch-is-the-best-fit-for-this-use-case"&gt;Why Amazon OpenSearch is the best fit for this use case&lt;/h3&gt;
-&lt;p&gt;Amazon OpenSearch Serverless is well suited for product catalog search because it supports combining semantic understanding with traditional keyword matching through hybrid search capabilities. When dealing with large product catalogs, performance matters. Amazon OpenSearch Serverless handl
-&lt;p&gt;What makes it particularly valuable for ecommerce is the built-in support for complex filtering and aggregations that power faceted navigation (think filtering by price, brand, or color). You can also choose from multiple distance metrics like cosine similarity or Euclidean distance to fine
-&lt;p&gt;Amazon OpenSearch Serverless Classic collections offer several optimization options to balance cost and search quality, as detailed in the following section.&lt;/p&gt;
-&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; Amazon Bedrock Knowledge Bases supports both Amazon OpenSearch Serverless and Managed Clusters. The following benchmarks were run on Serverless Classic collections. Amazon OpenSearch Serverless NextGen collections (generally available May 2026) aren’t yet
-&lt;h3 id="performance-analysis-and-optimizations-for-amazon-opensearch-serverless-vector-search"&gt;Performance analysis and optimizations for Amazon OpenSearch serverless vector search&lt;/h3&gt;
-&lt;p&gt;Amazon OpenSearch is highly configurable and provides several configuration options. Be careful when selecting these options because they can significantly affect the performance of the vector index. We consider some of these options targeted at optimizing cost and database size and quantit
-&lt;p&gt;Some of the common optimization options are:&lt;/p&gt;
-&lt;ol type="1"&gt;
- &lt;li&gt;&lt;strong&gt;Size of vector embeddings&lt;/strong&gt;: A larger vector can generally contain more semantic information about the embedded text. However, it also leads to higher memory consumption, which increases vector index size and cost. Modern embedding models like Amazon Titan Text
- &lt;li&gt;&lt;strong&gt;Data type of embeddings&lt;/strong&gt;: We can also reduce vector index size (and thus cost) by storing embeddings in lower precision data types, such as binary embeddings. This can significantly reduce the size of the vector index.&lt;/li&gt;
- &lt;li&gt;&lt;strong&gt;Disk optimized storage&lt;/strong&gt;: Amazon OpenSearch Serverless Classic collections offer disk-based vector search (on_disk mode) that applies 32× binary quantization internally while rescoring against full-precision vectors from disk. This preserves quality while reduci
-&lt;/ol&gt;
-&lt;p&gt;Depending on the indexing algorithm used, users may also configure HNSW parameters (ef_construction, m) to tune the trade-off between index build time, memory usage, and search accuracy (for practical guidance, see &lt;a href="https://opensearch.org/blog/a-practical-guide-to-selecting-hnsw-
-&lt;h3 id="dataset"&gt;Dataset&lt;/h3&gt;
-&lt;p&gt;We use the “Shopping Queries Data Set” (ESCI), a large dataset of difficult search queries provided by Amazon. The dataset contains 1,215,851 unique US products (title, description, bullets, and brand; approximately 1,140 characters median) and 97,345 judged queries. For each query, the dat
-&lt;p&gt;Query:&lt;/p&gt;
-&lt;p&gt;&lt;code&gt;self-seal envelopes without window&lt;/code&gt;&lt;/p&gt;
-&lt;p&gt;Relevant product title:&lt;/p&gt;
-&lt;p&gt;&lt;code&gt;BAZIC Security Self Seal Envelope 4 1/8" x 9 1/2" #10, No Window Tint Pattern Mailing Envelopes, Peel &amp;amp; Seal, Office Checks Invoices (30/Pack), 1-Pack&lt;/code&gt;&lt;/p&gt;
-&lt;p&gt;Irrelevant product title:&lt;/p&gt;
-&lt;p&gt;&lt;code&gt;ValBox 200 Count #8 Double Window Envelopes 3 5/8" x 8 11/16" Flip and Seal Double Window Security Check Envelopes- Security Tint Pattern Designed for Home Office Secure Mailing&lt;/code&gt;&lt;/p&gt;
-&lt;p&gt;We sample 5,000 queries (approximately 19 judged products per query, approximately 17 relevant) and index all 1,215,851 product descriptions for benchmarking. We measure retrieval quality (NDCG@10), latency (p50/p95/p99 at concurrency 1 and 10), and index size (ANN in-memory footprint).&lt;
-&lt;h3 id="vector-index-construction"&gt;Vector index construction&lt;/h3&gt;
-&lt;p&gt;We test all combinations of embedding dimension (1024, 512, 256) and data type (float, binary), totaling six configurations, plus 1024-float in on_disk mode at the default compression_level: 32x, compared against the 1024-float in-memory baseline (seven configurations total). All indexes us
+ <description>Amazon SageMaker AI shipped 13 inference launches in year-to-date across two deployment paths: fully managed endpoints and Amazon SageMaker HyperPod Inference. This post reviews each launch, from inference recommendations and capacity-aware instance pools to tiered KV caching and d
+ <content:encoded>&lt;p&gt;Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements are measured in tokens per second, cold starts can span multiple minutes as containers and weights transfer, GPU capacity is constrained, and traditional monit
+&lt;p&gt;Amazon SageMaker AI offers customers the ability to deploy AI models and consume them by the instance (instead of by the token), using two paths: managed endpoints for teams that want AWS to handle infrastructure and operations, and Amazon SageMaker HyperPod Inference for teams that need Ku
+&lt;h2 id="choose-the-deployment-that-fits-your-workload"&gt;Choose the deployment that fits your workload&lt;/h2&gt;
+&lt;p&gt;The table below compares the two deployment paths across seven dimensions.&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
&lt;tbody&gt;
&lt;tr&gt;
- &lt;td&gt;&lt;strong&gt;Configuration&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Embedding size&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Embedding type&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Dimension&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Endpoints&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;HyperPod&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;In memory&lt;/td&gt;
- &lt;td&gt;1024&lt;/td&gt;
- &lt;td&gt;float (baseline)&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Infrastructure&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Fully managed by AWS&lt;/td&gt;
+ &lt;td&gt;Managed Kubernetes stack&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;In memory&lt;/td&gt;
- &lt;td&gt;512&lt;/td&gt;
- &lt;td&gt;float&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Deploy target&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Console, SDK, CLI&lt;/td&gt;
+ &lt;td&gt;kubectl, Terraform, Console, CLI, SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;In memory&lt;/td&gt;
- &lt;td&gt;256&lt;/td&gt;
- &lt;td&gt;float&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Scaling&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Managed auto scaling with Amazon CloudWatch&lt;/td&gt;
+ &lt;td&gt;Auto scaling with Karpenter, KEDA, CloudWatch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;In memory&lt;/td&gt;
- &lt;td&gt;1024&lt;/td&gt;
- &lt;td&gt;binary&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Customization and Control&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Customizable at the container and model layers&lt;/td&gt;
+ &lt;td&gt;More customizability with Node level access, frameworks and AMI.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;In memory&lt;/td&gt;
- &lt;td&gt;512&lt;/td&gt;
- &lt;td&gt;binary&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;API protocol&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;OpenAI compatible with SageMaker endpoint&lt;/td&gt;
+ &lt;td&gt;HTTP, gRPC and custom load balancer capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;In memory&lt;/td&gt;
- &lt;td&gt;256&lt;/td&gt;
- &lt;td&gt;binary&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Best for&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Fast and fully managed deployment with minimal ops overhead&lt;/td&gt;
+ &lt;td&gt;Kubernetes-based, train-to-serve multi-cloud/hybrid-cloud deployments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;On disk (32×)&lt;/td&gt;
- &lt;td&gt;1024&lt;/td&gt;
- &lt;td&gt;float&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Corresponding Launches:&lt;/strong&gt;&lt;/td&gt;
+ &lt;/tr&gt;
+ &lt;tr&gt;
+ &lt;td&gt;&lt;strong&gt;Launches&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Inference recommendations, Capacity Aware Inference, OpenAI API, Container Caching, Observability, Async Inference Inline Payloads, Prefix-Aware Routing&lt;/td&gt;
+ &lt;td&gt;Simplified Operator, Tiered KV Cache, Data Capture, Performance Features, Disaggregated Prefill and Decode for HyperPod Inference, Model Caching&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
-&lt;h3 id="evaluation-methodology"&gt;Evaluation methodology&lt;/h3&gt;
-&lt;p&gt;For each configuration, we create a vector index in Amazon OpenSearch Serverless (Classic collection), ingest all 1,215,851 products, wait for merges to settle, then run an adaptive warm-up until latency stabilizes before measuring. We measure 1,000 queries × 3 repetitions at concurrency 1
-&lt;ol type="1"&gt;
- &lt;li&gt;Retrieval latency: Latency is measured as the time taken to retrieve relevant matches from the vector index as reported by the Amazon OpenSearch results. This doesn’t include the time to convert text to embeddings.&lt;/li&gt;
- &lt;li&gt;Retrieval performance: We use the &lt;a href="https://en.wikipedia.org/wiki/Discounted_cumulative_gain#Normalized_DCG" target="_blank" rel="noopener"&gt;Normalized Discounted Cumulative Gain (NDCG)&lt;/a&gt; metric to score the retrievals for each query. This metric measures the quality o
- &lt;li&gt;Index size: We report the ANN (Approximate Nearest Neighbor) index size, which is the in-memory structure that drives search compute cost and determines capacity requirements. This differs from total store size, which includes the _source JSON copy of each document and varies with documen
-&lt;/ol&gt;
-&lt;h3 id="results"&gt;Results&lt;/h3&gt;
-&lt;p&gt;Table 1: Semantic search (k-NN only) performance across seven Amazon OpenSearch Serverless configurations (1,215,851 indexed vectors, 5,000 queries, k=10). Latency is server-side at concurrency 1 unless noted. Deltas are paired bootstrap against the 1024-float in-memory baseline. See Table
-&lt;ol type="1"&gt;
- &lt;li&gt;Reducing dimensions doesn’t always reduce quality. On this dataset, 512-float was statistically indistinguishable from the 1024-float baseline (NDCG 0.3628 vs 0.3627, p = 0.87) at half the index size (2.79 vs 5.34 GiB) and lower latency (25 vs 31 ms p50). Dropping to 256 dimensions showed
- &lt;li&gt;Binarization offers large index size reductions, but the quality cost depends on the number of dimensions. At 1024 dimensions, binary embeddings reduced index size by 13.4× (0.40 vs 5.34 GiB) with a 5.2 percent NDCG loss and comparable latency (22 vs 31 ms p50). At 256 dimensions the qual
- &lt;li&gt;Disk mode (on_disk 32×) preserves quality at the cost of latency. At 1024 dimensions, disk mode achieved NDCG 0.3610 (−0.5 percent vs baseline) with the same 0.40 GiB index size as 1024-binary, but at approximately 3× higher latency (99 ms p50 vs 31 ms in-memory). Both 1024-binary and on_
-&lt;/ol&gt;
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="images/image1.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/18/ML-21809-1-1.png" alt="Two SageMaker AI inference deployment paths delivered in 2026: managed endpoints and HyperPod Inference" wid
+ &lt;p class="wp-caption-text"&gt;Figure 1: Two inference paths delivered in 2026&lt;/p&gt;
+&lt;/div&gt;
+&lt;h2 id="sagemaker-ai-endpoints-from-model-to-production-in-hours"&gt;SageMaker AI endpoints: From model to production in hours&lt;/h2&gt;
+&lt;p&gt;Managed SageMaker Inference endpoints are the faster path for teams that want AWS to handle GPU provisioning, scaling, and operational monitoring. You bring the model and define the performance target. SageMaker handles the rest. The seven launches year-to-date in 2026 below address deploym
+&lt;h3 id="inference-recommendations-and-benchmarking-april-2026"&gt;Inference recommendations and benchmarking (April 2026)&lt;/h3&gt;
+&lt;p&gt;Blog: &lt;a href="https://aws.amazon.com/blogs/machine-learning/amazon-sagemaker-ai-now-supports-optimized-generative-ai-inference-recommendations/" target="_blank" rel="noopener"&gt;Amazon SageMaker AI now supports optimized generative AI inference recommendations&lt;/a&gt;&lt;/p&gt;
+&lt;p&gt;Choosing the right instance type, serving container, and optimization settings for a generative AI model typically takes two to three weeks of manual benchmarking against 1000+ combinations, requiring expertise most teams do not have in-house. Inference recommendations automate this end-to-
+&lt;p&gt;Customers specify a model and performance goal (cost, latency, or throughput). SageMaker then runs a three-step process:&lt;/p&gt;
+&lt;div style="width: 810px" class="wp-caption alignnone"&gt;
+ &lt;a href="images/image2.png" target="_blank" rel="noopener"&gt;&lt;img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/18/ML-21809-2-1.png" alt="The three-step inference recommendations process: narrow, optimize, and benchmark" width="800"&gt;&lt;/a&gt;
+ &lt;p class="wp-caption-text"&gt;Figure 2: Inference recommendations 3-step process&lt;/p&gt;
+&lt;/div&gt;
+&lt;ul&gt;
+ &lt;li&gt;&lt;strong&gt;Narrow.&lt;/strong&gt; Filter the instance type space by analyzing model architecture, size, and memory requirements.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Optimize.&lt;/strong&gt; Apply goal-aligned techniques: EAGLE 3.0 speculative decoding for throughput, kernel tuning for latency, tensor parallelism based on model size.&lt;/li&gt;
+ &lt;li&gt;&lt;strong&gt;Benchmark.&lt;/strong&gt; Run NVIDIA AIPerf on real GPU infrastructure with statistically rigorous multi-run confidence reporting.&lt;/li&gt;
+&lt;/ul&gt;
+&lt;p&gt;The output is a SageMaker Model Package with deployment-ready configurations and validated metrics: time to first token (TTFT), inter-token latency (ITL), P50/P90/P99 latency percentiles, throughput, and cost projection. In a demonstrated example, throughput optimization on &lt;code&gt;GPT-
+&lt;h3 id="capacity-aware-instance-pools-may-2026"&gt;Capacity-aware instance pools (May 2026)&lt;/h3&gt;
+&lt;p&gt;Blog: &lt;a href="https://aws.amazon.com/blogs/machine-learning/capacity-aware-inference-automatic-instance-fallback-for-sagemaker-ai-endpoints/" target="_blank" rel="noopener"&gt;Capacity-aware inference: automatic instance fallback for SageMaker AI endpoints&lt;/a&gt;&lt;/p&gt;
+&lt;p&gt;When a SageMaker endpoint required a single instance type, a capacity shortage meant the endpoint failed before serving a single request. Instance pools address that single point of failure.&lt;/p&gt;
+&lt;p&gt;Customers define a prioritized list of up to five instance types. SageMaker automatically works through the list at endpoint creation, during scale-out, and during scale-in. At creation, SageMaker tries the first-choice type and falls back immediately if capacity is unavailable. During scal
+&lt;p&gt;Per-instance-type CloudWatch metric dimensions enable weighted scaling policies for heterogeneous fleets. Each pool entry can reference a separate optimized model configuration (tensor parallelism on high-memory instances, speculative decoding on mid-tier, quantization on smaller fallbacks)
+&lt;h3 id="openai-compatible-apis-may-2026"&gt;OpenAI-compatible APIs (May 2026)&lt;/h3&gt;
+&lt;p&gt;Blog: &lt;a href="https://aws.amazon.com/blogs/machine-learning/announcing-openai-compatible-api-support-for-amazon-sagemaker-ai-endpoints/" target="_blank" rel="noopener"&gt;Announcing OpenAI-compatible API support for Amazon SageMaker AI endpoints&lt;/a&gt;&lt;/p&gt;
+&lt;p&gt;Applications built on the OpenAI SDK, LangChain, or Strands Agents previously required custom client adapters and authentication rewrites to work with SageMaker-hosted models. That migration cost was a real barrier.&lt;/p&gt;
+&lt;p&gt;SageMaker endpoints now expose an /openai/v1 path supporting Chat Completions with streaming. Migration requires changing only the endpoint URL. SDK calls, streaming logic, and prompt formatting remain identical. Authentication uses bearer tokens generated from existing AWS credentials, val
+&lt;p&gt;Multi-model endpoints allow hosting multiple models, each callable through the same OpenAI SDK with independent resource allocation. For agentic workloads, AI agents can run entirely on customer-owned GPU infrastructure using the same OpenAI-compatible interface they were built on. Availabl
+&lt;h3 id="container-caching-june-2026"&gt;Container caching (June 2026)&lt;/h3&gt;
+&lt;p&gt;Blog: &lt;a href="https://aws.amazon.com/blogs/machine-learning/introducing-container-caching-in-amazon-sagemaker-ai-for-faster-model-scaling/" target="_blank" rel="noopener"&gt;Introducing container caching in Amazon SageMaker AI for faster model scaling&lt;/a&gt;&lt;/p&gt;
+&lt;p&gt;During inference auto scaling events, new instances responding to traffic spikes previously had to pull the full container image from Amazon Elastic Container Registry (Amazon ECR) before serving requests. For large serving containers exceeding 10 GB, that pull alone added several minutes o
+&lt;p&gt;Container caching pre-pulls images automatically, so new instances launch with the container already available locally. Zero configuration, no code changes, no container modifications. It activates automatically on supported accelerator instance types. With Qwen3-8B on ml.g6.2xlarge using t
+&lt;p&gt;Container caching is the third layer in a three-part scaling optimization suite:&lt;/p&gt;
&lt;table border="1px" width="100%" cellpadding="10px"&gt;
&lt;tbody&gt;
&lt;tr&gt;
- &lt;td&gt;&lt;strong&gt;Configuration&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;NDCG@10&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Δ vs baseline&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;p50 (ms)&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;p95 (ms)&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;p99 (ms)&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;Index size (ANN)&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;p50 @ conc 10&lt;/strong&gt;&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;1024 float&lt;/td&gt;
- &lt;td&gt;0.3627&lt;/td&gt;
- &lt;td&gt;baseline&lt;/td&gt;
- &lt;td&gt;31&lt;/td&gt;
- &lt;td&gt;44&lt;/td&gt;
- &lt;td&gt;52&lt;/td&gt;
- &lt;td&gt;5.34 GiB&lt;/td&gt;
- &lt;td&gt;161 ms&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;512 float&lt;/td&gt;
- &lt;td&gt;0.3628&lt;/td&gt;
- &lt;td&gt;+0.0002&lt;/td&gt;
- &lt;td&gt;25&lt;/td&gt;
- &lt;td&gt;37&lt;/td&gt;
- &lt;td&gt;62&lt;/td&gt;
- &lt;td&gt;2.79 GiB&lt;/td&gt;
- &lt;td&gt;149 ms&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;256 float&lt;/td&gt;
- &lt;td&gt;0.3468&lt;/td&gt;
- &lt;td&gt;−4.4%&lt;/td&gt;
- &lt;td&gt;22&lt;/td&gt;
- &lt;td&gt;35&lt;/td&gt;
- &lt;td&gt;47&lt;/td&gt;
- &lt;td&gt;1.51 GiB&lt;/td&gt;
- &lt;td&gt;85 ms&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;1024 binary&lt;/td&gt;
- &lt;td&gt;0.3438&lt;/td&gt;
- &lt;td&gt;−5.2%&lt;/td&gt;
- &lt;td&gt;22&lt;/td&gt;
- &lt;td&gt;37&lt;/td&gt;
- &lt;td&gt;55&lt;/td&gt;
- &lt;td&gt;0.40 GiB&lt;/td&gt;
- &lt;td&gt;70 ms&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Layer&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Optimization&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Impact&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;512 binary&lt;/td&gt;
- &lt;td&gt;0.3200&lt;/td&gt;
- &lt;td&gt;−11.8%&lt;/td&gt;
- &lt;td&gt;19&lt;/td&gt;
- &lt;td&gt;30&lt;/td&gt;
- &lt;td&gt;47&lt;/td&gt;
- &lt;td&gt;0.32 GiB&lt;/td&gt;
- &lt;td&gt;57 ms&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Detection&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Sub-minute CloudWatch metrics&lt;/td&gt;
+ &lt;td&gt;Triggers scale-up 6x faster than standard 1-minute metrics&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;256 binary&lt;/td&gt;
- &lt;td&gt;0.2600&lt;/td&gt;
- &lt;td&gt;−28.3%&lt;/td&gt;
- &lt;td&gt;17&lt;/td&gt;
- &lt;td&gt;24&lt;/td&gt;
- &lt;td&gt;32&lt;/td&gt;
- &lt;td&gt;0.28 GiB&lt;/td&gt;
- &lt;td&gt;52 ms&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;Existing instances&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Instance-store data caching&lt;/td&gt;
+ &lt;td&gt;Removes image pull and model download for instances already running&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
- &lt;td&gt;1024 float, on_disk 32×&lt;/td&gt;
- &lt;td&gt;0.3610&lt;/td&gt;
- &lt;td&gt;−0.5%&lt;/td&gt;
- &lt;td&gt;99&lt;/td&gt;
- &lt;td&gt;139&lt;/td&gt;
- &lt;td&gt;176&lt;/td&gt;
- &lt;td&gt;0.40 GiB&lt;/td&gt;
- &lt;td&gt;255 ms&lt;/td&gt;
+ &lt;td&gt;&lt;strong&gt;New instances&lt;/strong&gt;&lt;/td&gt;
+ &lt;td&gt;Container image caching&lt;/td&gt;
+ &lt;td&gt;Avoids image pull time; 51% startup latency reduction demonstrated&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
-&lt;h3 id="hybrid-search-results"&gt;Hybrid search results&lt;/h3&gt;
-&lt;p&gt;Table 2: Hybrid search comparison at 1024 dimensions (1,215,851 vectors, 5,000 queries, k=10). Hybrid uses normalization-processor with 0.7 semantic / 0.3 lexical weighting.&lt;/p&gt;
-&lt;table border="1px" width="100%" cellpadding="10px"&gt;
- &lt;tbody&gt;
- &lt;tr&gt;
- &lt;td&gt;&lt;strong&gt;Method&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;1024-float NDCG&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;float p50&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;1024-binary NDCG&lt;/strong&gt;&lt;/td&gt;
- &lt;td&gt;&lt;strong&gt;binary p50&lt;/strong&gt;&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;Keyword (BM25)&lt;/td&gt;
- &lt;td&gt;0.3141&lt;/td&gt;
- &lt;td&gt;11 ms&lt;/td&gt;
- &lt;td&gt;0.3177&lt;/td&gt;
- &lt;td&gt;17 ms&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;Semantic (k-NN)&lt;/td&gt;
- &lt;td&gt;0.3633&lt;/td&gt;
- &lt;td&gt;30 ms&lt;/td&gt;
- &lt;td&gt;0.3451&lt;/td&gt;
- &lt;td&gt;17 ms&lt;/td&gt;
- &lt;/tr&gt;
- &lt;tr&gt;
- &lt;td&gt;Hybrid&lt;/td&gt;
- &lt;td&gt;0.3850&lt;/td&gt;
- &lt;td&gt;38 ms&lt;/td&gt;
- &lt;td&gt;0.3658&lt;/td&gt;
- &lt;td&gt;35 ms&lt;/td&gt;
- &lt;/tr&gt;
- &lt;/tbody&gt;
-&lt;/table&gt;
-&lt;p&gt;Hybrid search provides a consistent quality lift over semantic-only search. On this dataset, hybrid improved NDCG by +6.0 percent over semantic-only for both float (0.3633 → 0.3850) and binary (0.3451 → 0.3658). In particular, 1024-binary with hybrid (0.3658) exceeded 1024-float with semant
-&lt;p&gt;Hybrid search adds latency. The fusion step roughly doubles sequential latency (30 → 38 ms for float, 17 → 35 ms for binary) and the gap widens under concurrency.&lt;/p&gt;
-&lt;p&gt;Note: This dataset (product search) favors keyword matching: BM25 alone reached NDCG 0.314, only 13 percent behind semantic. On natural-language RAG queries the absolute hybrid lift may differ. The fusion weight (0.7/0.3) was set for demonstration, not tuned.&lt;/p&gt;
-&lt;h2 id="use-case-2-deep-research-agent"&gt;Use case 2: Deep research agent&lt;/h2&gt;
-&lt;p&gt;Deep research agents represent a significant evolution beyond traditional RAG systems. While standard RAG performs a single retrieval-generation cycle, deep research agents tackle complex, multi-turn research tasks through dynamic reasoning, adaptive planning, and iterative information retr

Diff display stops at 400 lines. The line counts above are from the whole diff. 65 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.