Change
a8f7c92
a8f7c9249bb9732f4f437e9c4a45669a8c87f743 · commit on GitHub
aws-blog-feed: changed (780721 bytes, HTTP 200)
raw/aws-blog-feed/response.xml modified
- Source
- aws-blog-feed
- Lines added
- +5,812
- Lines removed
- -4,787
- Stored bytes at this commit
- 780,721
- Timestamp
- observed
- Raw artifact at this commit
- raw/aws-blog-feed/response.xml
Recorded headers
| observed_at | 2026-09-18T04:44:07.147Z |
|---|---|
| origin_date | null |
| status | 200 |
| final URL | https://aws.amazon.com/blogs/machine-learning/feed/ |
| etag | null |
| last-modified | Fri, 18 Sep 2026 01:28:37 GMT |
| date | Fri, 18 Sep 2026 04:44:07 GMT |
| age | null |
| cache-control | null |
| cf-cache-status | null |
| content-encoding | null |
| content-length | null |
@
@@ -5,7 +5,7 @@ <atom:link href="https://aws.amazon.com/blogs/machine-learning/feed/" rel="self" type="application/rss+xml"/> <link>https://aws.amazon.com/blogs/machine-learning/</link> <description>Official Machine Learning Blog of Amazon Web Services</description>-
<lastBuildDate>Wed, 16 Sep 2026 19:00:00 +0000</lastBuildDate>+
<lastBuildDate>Thu, 17 Sep 2026 18:02:01 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod>@
@@ -13,602 +13,873 @@ 1 </sy:updateFrequency> <item>-
<title>Improving HCLS AI reasoning with open-source agent skills</title>-
<link>https://aws.amazon.com/blogs/machine-learning/improving-hcls-ai-reasoning-with-open-source-agent-skills/</link>+
<title>Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent</title>+
<link>https://aws.amazon.com/blogs/machine-learning/reduce-time-to-hire-for-quality-candidates-with-ai-powered-amazon-connect-talent/</link>+
-
<dc:creator><![CDATA[Michael Hsieh]]></dc:creator>-
<pubDate>Wed, 16 Sep 2026 19:00:00 +0000</pubDate>-
<category><![CDATA[Amazon Bedrock]]></category>-
<category><![CDATA[Amazon Quick Suite]]></category>+
<dc:creator><![CDATA[Ayesha Borker]]></dc:creator>+
<pubDate>Thu, 17 Sep 2026 17:55:20 +0000</pubDate>+
<category><![CDATA[Amazon Connect]]></category> <category><![CDATA[Announcements]]></category>-
<category><![CDATA[Artificial Intelligence]]></category>-
<category><![CDATA[Generative AI]]></category>-
<category><![CDATA[Healthcare]]></category>+
<guid isPermaLink="false">a6559b315a35bac095e4d57dabbff35e09f15496</guid>+
+
<description>Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while providing applicants w…+
<content:encoded><p>Hiring at scale in industries such as retail, logistics, hospitality, and others has its fair share of challenges. Recruiting teams are expected to fill hundreds of roles within tight timelines, often with limited capacity and with tools that weren’t designed to s…+
<p>Today, we’re launching <a href="https://aws.amazon.com/products/connect/talent/" target="_blank" rel="noopener">Amazon Connect Talent,</a> an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, a…+
<p>Recruiters configure the <a href="https://docs.aws.amazon.com/talent/latest/userguide/evaluations.html" target="_blank" rel="noopener">evaluation criteria</a>, assessments, and interview questions based on job requirements. <a href="https://docs.aws.amazon.com/talent/latest/u…+
<p><a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-1.png" target="_blank" rel="noopener"><img class="alignnone wp-image-139627 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b5…+
<p><em>Amazon Connect Talent lets recruiters configure AI-led interviews and candidate assessments in minutes, tailored to each role.</em></p> +
<p><strong> <a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-2-1.png" target="_blank" rel="noopener"><img loading="lazy" class="alignnone wp-image-139629 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836c…+
<p><em>Candidates experience a consistent, structured AI-led interview — available whenever they’re ready, day or night.</em></p> +
<h2><strong>Informed by decades of Amazon’s hiring science</strong></h2> +
<p>Amazon is one of the world’s largest employers, and Amazon Connect Talent puts decades of Amazon’s hiring science to work for organizations, with a highly configurable solution that adapts to specific hiring requirements. With consistent, evidence-based assessments applied to every candidat…+
<h3><strong>Reducing human preconceptions in hiring</strong></h3> +
<p>Connect Talent’s AI focuses exclusively on measuring a candidate’s job-related competencies, testing abilities including problem-solving, logic, listening, and role-specific capabilities. All candidate data is anonymized during AI evaluation, removing factors that can introduce unconscious …+
<h3><strong>Transparency and consistent evaluation standards</strong></h3> +
<p>Connect Talent communicates clearly to both recruiters and candidates about what data is collected, how it is used, and what is not collected. Candidates are informed of what to expect before they proceed to take the evaluation. Each competency is scored against a rubric that defines what a…+
<h3><strong>Human in the loop: recruiter control at every step</strong></h3> +
<p>Recruiters maintain final decision authority over every hire. Connect Talent gives them scored candidate summaries with competency breakdowns, complete interview transcripts, comparative analytics, and clear reasoning behind every score, all in a single view instead of pieced together from …+
<p><a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-3.png" target="_blank" rel="noopener"><img loading="lazy" class="alignnone wp-image-139630 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f…+
<p><em>Amazon Connect Talent lets recruiters configure AI-led interviews and candidate assessments in minutes, tailored to each role.</em></p> +
<p><a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-4.png" target="_blank" rel="noopener"><img loading="lazy" class="alignnone wp-image-139631 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f…+
<p><a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/ConnectTalent-5.png" target="_blank" rel="noopener"><img loading="lazy" class="alignnone wp-image-139632 size-full" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f…+
<p><em>Amazon Connect Talent lets recruiters review candidate responses and evaluation scores with details and full transparency</em></p> +
<h2><strong>Enterprise-grade security built on AWS infrastructure</strong></h2> +
<p>Hiring data is sensitive. Candidate information deserves the same level of protection&nbsp;you’d&nbsp;expect for financial records or health information. Amazon Connect Talent delivers enterprise-grade security built on AWS infrastructure:&nbsp;the same foundation trusted by ban…+
<p>Built-in security controls help protect candidate data, with rigorous measures to meet your requirements. Access is controlled, auditable, and configurable. When your candidates share their information, they can trust&nbsp;it’s&nbsp;protected.</p> +
<ul> +
<li><strong>Integrity monitoring and fraud protection:&nbsp;</strong>Amazon Connect Talent uses&nbsp;text-based&nbsp;analysis to flag unnatural cadence, filler words, pauses, and response latency.&nbsp;Human review is mandatory for every flag. No candidate is disqu…+
<li><strong>Audit trails and explainability:&nbsp;</strong>Every&nbsp;candidate&nbsp;interaction is logged with a complete audit trail and clear job-related evaluation reasoning. The system includes ongoing monitoring and tracking to help your organization&nbsp;dem…+
</ul> +
<h2><strong>Improved business outcomes</strong></h2> +
<p>Amazon Connect Talent improves business outcomes by getting recruiters out of the workflow management cycle and empowering them with faster decision-making. Candidates get a faster and more flexible hiring experience. Organizations fill roles before lost revenue, rising costs, and competiti…+
<p>Your business moves fast and now your hiring can too.</p> +
<p><iframe loading="lazy" title="Amazon Connect Talent Introduction" width="500" height="281" src="https://www.youtube-nocookie.com/embed/0X1PH7ZwvRo?feature=oembed" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" r…+
<p><a href="https://aws.amazon.com/products/connect/talent/" target="_blank" rel="noopener noreferrer">Learn more about Amazon Connect Talent</a> and discover how AI-powered hiring can transform your talent acquisition strategy.</p> +
<h2>About the authors</h2> +
<footer> +
<div class="blog-author-box" style="padding-top: 2.0em"> +
<div class="blog-author-image" style="margin-right: 1.0em">+
<img loading="lazy" class="alignnone wp-image-139633 size-thumbnail" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/Ayesha-100x133.jpg" alt="" width="100" height="133">+
</div> +
<h3 class="lb-h4">Ayesha Borker</h3> +
<p style="overflow: hidden">Ayesha is a Principal Solutions Architect, Applied AI at AWS. With over a decade of experience at the intersection of human and AI collaboration specializing in customer experience. She simplifies the complexity of AI, helping organizations stay focused on the out…+
</div> +
<div class="blog-author-box" style="padding-top: 2.0em"> +
<div class="blog-author-image" style="margin-right: 1.0em">+
<img loading="lazy" class="alignnone size-thumbnail wp-image-139635" src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/16/Kate-100x133.jpg" alt="" width="100" height="133">+
</div> +
<h3 class="lb-h4">Kate Totaro</h3> +
<p style="overflow: hidden">Kate is the Principal Product Manager for Amazon Connect Talent, AWS’s AI-powered hiring service. She brings 20 years of experience building global businesses, including more than 15 years at Amazon and AWS across product, business, and operations. At Amazon, she …+
</div> +
</footer></content:encoded>+
+
+
+
+
+
</item>+
<item>+
<title>Selecting a vector store for Amazon Bedrock Knowledge Bases</title>+
<link>https://aws.amazon.com/blogs/machine-learning/selecting-a-vector-store-for-amazon-bedrock-knowledge-bases/</link>+
+
<dc:creator><![CDATA[Deepak Dalakoti]]></dc:creator>+
<pubDate>Thu, 17 Sep 2026 15:53:13 +0000</pubDate>+
<category><![CDATA[Amazon Bedrock Knowledge Bases]]></category>+
<category><![CDATA[Best Practices]]></category> <category><![CDATA[Intermediate (200)]]></category>-
<category><![CDATA[Kiro]]></category>-
<category><![CDATA[Life Sciences]]></category>-
<category><![CDATA[Open Source]]></category>-
<category><![CDATA[Strands Agents]]></category>-
<guid isPermaLink="false">fd565c39336c51430810e927ad66178c9158fbf3</guid>-
-
<description>AI agents on foundation models often misapply healthcare and life sciences decision frameworks, citing the right guideline but applying it incorrectly. This post shares 38 open-source agent skills across 11 HCLS domains that close this gap, with installation steps, three worked use…-
<content:encoded><p>AI agents built on foundation models (FMs) often misapply healthcare and life sciences (HCLS) decision frameworks, even when they’ve seen the guidelines in training and in the system prompt. Ask an agent to classify a TP53 missense variant using ACMG/AMP criteria.…-
<p>In this post, we share a collection of 38 open source agent skills spanning 11 HCLS domains that help close this methodology gap. We walk through installation and show how to use them across agentic AI services. We share our evaluation results to demonstrate measurable improvement across dr…-
<h2 id="solution-overview">Solution overview</h2> -
<p>Agent skills in the HCLS Agent Skills collection are structured markdown documents (<code>SKILL.md</code>) that encode domain decision procedures into a format AI agents can consume at inference time through progressive disclosure. Following the <a href="https://agentskills.i…-
<p>Skills in this repository are sorted as either reasoning or pipeline skills. Reasoning skills encode methodology and decision frameworks that guide how the agent thinks. For example, the <a href="https://github.com/aws-samples/sample-hcls-agent-skills/blob/main/skills/genomic-variant-int…-
<p>This dual taxonomy gives agents both the judgment to make correct decisions and the technical precision to execute them. Unlike Retrieval Augmented Generation (RAG), which retrieves limited passages from indexed documents to augment the response generation, skills encode the decision proced…-
<p>Three properties make skills distinct from other approaches to domain specialization. Skills are auditable, portable, and straightforward to maintain. Every decision criterion is human-readable in markdown format, not hidden in model weights. A skill works across over 20 services (<a hre…-
<p>Now that you understand what skills contain, let’s set them up.</p> -
<h2 id="prerequisites">Prerequisites</h2> -
<p>To follow along with the examples in this post, you need one of the supported services from AWS: <a href="https://kiro.dev/ide/" target="_blank" rel="noopener">Kiro</a> or <a href="https://kiro.dev/cli/" target="_blank" rel="noopener">Kiro CLI</a> for interactive ski…-
<p>Start by cloning the repository:</p> -
<div class="hide-language"> -
<pre><code class="language-bash">git clone https://github.com/awslabs/hcls-agent-skills.git-
cd hcls-agent-skills</code></pre> -
</div> -
<p>To install skills only without the agent configuration, use the universal <a href="https://github.com/vercel-labs/skills" target="_blank" rel="noopener">skills CLI</a>:</p> -
<div class="hide-language"> -
<pre><code class="language-bash">npx skills add awslabs/hcls-agent-skills</code></pre> -
</div> -
<p>For Kiro, the <code>install.sh</code> script installs both skills and a pre-configured agent that equips them. The agent handles skill routing automatically, so you don’t need to invoke individual skills by name. Run <code>./install.sh --target kiro</code>, then swit…-
<p>For the AWS Strands Agents SDK, load skills directly in your Python code:</p> -
<div class="hide-language"> -
<pre><code class="language-python">from strands import Agent-
from strands.skills import AgentSkills-
-
agent = Agent(-
model=model_id,-
skills=AgentSkills(skills="./skills/"),-
)</code></pre> -
</div> -
<p>For <a href="https://aws.amazon.com/bedrock/agentcore/" target="_blank" rel="noopener">AgentCore</a>, follow <a href="https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-skills.html" target="_blank" rel="noopener">Skills</a> to add agent skills to a…-
<p>For Amazon Quick Desktop, run <code>./install.sh --target quick-desktop</code> to see the full instructions for adding skills in the graphical interface. Alternatively, follow the instructions in <a href="https://docs.aws.amazon.com/quick/latest/userguide/skills-desktop.html"…-
<h2 id="solution-walkthrough">Solution walkthrough</h2> -
<p>With skills installed, we demonstrate three deployment patterns: the simplest single-agent approach in Quick Desktop, multi-agent orchestration in Kiro CLI that addresses context engineering challenges, and production deployment with Strands SDK on Amazon Bedrock AgentCore. We then show thr…-
<h3 id="agent-skills-in-action-with-quick-desktop">Agent skills in action with Quick Desktop</h3> -
<p>With skills installed, Quick Desktop’s agent gains structured HCLS domain reasoning without additional configuration. When you ask a domain question, the agent automatically activates relevant skills based on trigger patterns in your query. For example, asking “What is the RAF impact of cod…-
<div style="width: 640px;" class="wp-video">-
<video class="wp-video-shortcode" id="video-139153-1" width="640" height="360" preload="metadata" controls="controls">-
<source type="video/mp4" src="https://d2908q01vomqb2.cloudfront.net/artifacts/DBSBlogs/ML-21213/quick-desktop-risk-adjustment-withskill.mp4?_=1">-
</video>-
</div> -
<p>Amazon Quick Desktop chat with the risk-adjustment skill dynamically loaded to respond to a RAF coding question</p> -
<h3 id="multi-agent-architecture-with-kiro">Multi-agent architecture with Kiro</h3> -
<p>Loading all 38 skills into a single agent context consumes ~80K tokens. This is workable with large-context models, but it creates a context engineering challenge. The agent must select the right subset from 38 available skills on every query and irrelevant skill content competes for attent…-
<p>Kiro CLI’s multi-agent architecture solves both problems. A lightweight coordinator agent (no skills loaded) routes queries to eight domain specialists, each loading only its relevant skills (approximately 15K tokens per specialist). The coordinator handles intent classification while the s…-
<figure>-
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/11/ML-21213-2.png" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/11/ML-21213-2.png" alt="Table of the eight doma…-
<figcaption aria-hidden="true">-
Table of the eight domain specialist agents in the Kiro CLI multi-agent architecture and the skills assigned to each-
</figcaption>-
</figure> -
<p>The multi-agent configuration is defined in JSON agent files. Refer to the <a href="https://github.com/aws-samples/sample-hcls-agent-skills/blob/main/agents/multiagent/kiro/hcls-multiagent.json" target="_blank" rel="noopener">coordinator agent config</a> for the routing logic, a…-
<div style="width: 640px;" class="wp-video">-
<video class="wp-video-shortcode" id="video-139153-2" width="640" height="360" preload="metadata" controls="controls">-
<source type="video/mp4" src="https://d2908q01vomqb2.cloudfront.net/artifacts/DBSBlogs/ML-21213/kiro-multiagent-drug-repurposing.mp4?_=2">-
</video>-
</div> -
<p>Kiro CLI answering a drug repurposing question in a code base using multi-agent routing and dynamic skill activation</p> -
<h3 id="strands-sdk-integration">Strands SDK integration</h3> -
<p>The AWS Strands Agents SDK provides native skill loading for building custom HCLS agents:</p> -
<div class="hide-language"> -
<pre><code class="language-python">from strands import Agent-
from strands.skills import AgentSkills-
from strands.multiagent import MultiAgentOrchestrator-
-
# Define domain specialists with their skill sets-
genomics_agent = Agent(-
name="hcls-genomics",-
model=model_id,-
skills=AgentSkills(skills="./skills/genomics/"),-
)-
-
imaging_agent = Agent(-
name="hcls-imaging",-
model=model_id,-
skills=AgentSkills(skills="./skills/imaging/"),-
)-
-
# Coordinator routes to specialists-
coordinator = MultiAgentOrchestrator(-
agents=[genomics_agent, imaging_agent, ...],-
model=model_id,-
)-
-
response = coordinator("Classify NM_000546.6:c.743G&gt;A in TP53 using ACMG criteria")</code></pre> -
</div> -
<h3 id="deploying-to-amazon-bedrock-agentcore">Deploying to Amazon Bedrock AgentCore</h3> -
<p>After your skill-equipped agent works locally, you can move it to production. Amazon Bedrock AgentCore provides an alternative path to inject skills into hosted agents. In addition to embedding them in the Strands agent code, you can configure skills at the environment level so they’re avai…-
<p>With deployment covered, let’s look at what skill-equipped agents produce in practice. The following sample use cases are drawn from our evaluation prompt set.</p> -
<h3 id="use-case-1-evaluating-repurposing-candidates-for-rare-fibrotic-disease-in-drug-discovery">Use case 1: Evaluating repurposing candidates for rare fibrotic disease in drug discovery</h3> -
<p>A team at a biotech company investigating drug repurposing for idiopathic pulmonary fibrosis (IPF) wants to evaluate approved drugs that modulate TGF-β1 signaling through the receptor kinase TGFBR1 (ALK5). In practice, a researcher needs to query drug-gene interaction databases, rank candid…-
<p>Before adding skills, the agent provides a general literature review listing known TGFBR1 inhibitors without structured ranking criteria, evidence hierarchy, or translatability assessment framework. After equipping the agent with skills, the agent triggers <a href="https://github.com/aws…-
<ol type="1"> -
<li>The agent applies the DGIdb query framework, prioritizing interaction types (inhibitor &gt; modulator &gt; binder) and source databases (ChEMBL, DrugBank) over lower-confidence sources.</li> -
<li>It ranks candidates using a structured evidence hierarchy where direct target engagement outweighs pathway-level evidence, which in turn outweighs phenotypic association, with existing indication relevance applied as a modifier.</li> -
<li>It assesses mechanism-of-action overlap by mapping TGFBR1 inhibition to the key IPF pathological processes: fibroblast-to-myofibroblast transition, epithelial-mesenchymal transition, and extracellular matrix deposition.</li> -
<li>It evaluates clinical translatability using T0→T1 criteria, examining existing safety data from the original indication, therapeutic window compatibility, and concordance between available preclinical fibrosis models and human disease.</li> -
</ol> -
<p>The skill chain transforms a surface-level response into a structured regulatory-aware evaluation with quantified evidence rankings.</p> -
<h3 id="use-case-2-building-a-cms-hcc-risk-adjustment-pipeline-in-healthcare-claims-operations">Use case 2: Building a CMS-HCC risk adjustment pipeline in healthcare claims operations</h3> -
<p>A Medicare Advantage plan with 12,000 members needs to calculate Risk Adjustment Factor (RAF) scores from ICD-10 diagnosis claims data using CMS-HCC Model V28 coefficients. The pipeline must apply the <code>ICD-10-to-HCC</code> crosswalk, resolve disease hierarchies correctly, a…-
<p>Before adding skills, the agent produces a plausible but incomplete pipeline, often missing hierarchy resolution entirely, using outdated V24 coefficients, or applying hierarchies after summing (which inflates scores). After equipping the agent with skills, the agent triggers <a href="ht…-
<ol type="1"> -
<li>The agent generates correct SQL that joins diagnosis codes to the <code>ICD-10-to-HCC</code> crosswalk table with deduplication within the measurement year, making sure each HCC is counted only once per member.</li> -
<li>It implements V28 hierarchy resolution correctly, where HCC 18 (Diabetes with Chronic Complications) supersedes HCC 19 (Diabetes without Complications) and HCC 326 (CKD Stage 5) supersedes HCC 327 (CKD Stage 4), helping prevent double-counting at multiple specificity levels.</li> -
<li>It applies the correct demographic segmentation by categorizing members into community, institutional, or dual-eligible populations with age/sex adjustments before summing HCC coefficients.</li> -
<li>It proactively explains that skipping hierarchy resolution double-counts conditions at multiple specificity levels, systematically inflating RAF scores and creating audit liability under CMS RADV review.</li> -
</ol> -
<p>The skill supports producing audit-defensible RAF scores rather than inflated estimates that would trigger CMS RADV audit findings.</p> -
<h3 id="use-case-3-t1-weighted-mri-preprocessing-for-voxel-based-morphometry-in-medical-imaging-research">Use case 3: T1-weighted MRI preprocessing for voxel-based morphometry in medical imaging research</h3> -
<p>A neuroimaging study with 45 healthy adults needs a standard T1w preprocessing pipeline for voxel-based morphometry (VBM) analysis. Raw DICOM data has been converted to NIfTI. The pipeline must reorient, correct bias field, skull-strip, and register to MNI152 space in the correct order and …-
<p>Before adding skills, the agent suggests a reasonable pipeline but may order bias correction after skull stripping (which biases brain masks), use inappropriate thresholds, or omit failure mode detection strategies. After equipping the agent with skills, the agent triggers <a href="https…+
<guid isPermaLink="false">e129d8e2814c4c6f57fdc5feec60ab4d06f5e04c</guid>+
+
<description>Choosing the right vector store for your Amazon Bedrock Knowledge Bases RAG application affects performance and cost. This post compares Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors across three RAG use cases, with benchmarks and a practi…+
<content:encoded><p>When building a Retrieval Augmented Generation (RAG) solution with <a href="https://aws.amazon.com/bedrock/knowledge-bases/" target="_blank" rel="noopener">Amazon Bedrock Knowledge Bases</a>, selecting the right vector store impacts performance and cos…+
<p>For broader guidance across all AWS vector solutions, see <a href="https://aws.amazon.com/blogs/machine-learning/aws-vector-solutions-build-agentic-ai-where-your-data-lives/" target="_blank" rel="noopener">AWS vector solutions: Build agentic AI where your data lives</a>. For the…+
<h2 id="how-vector-databases-fit-into-rag-solutions">How vector databases fit into RAG solutions</h2> +
<p>A RAG architecture combines the capabilities of large language models (LLMs) with information retrieval systems to generate more accurate, up-to-date, and contextually relevant responses. It is based on the mathematical concept of a <em>vector</em>, where the text is translated …+
<p>When a user submits a query, it is converted into a vector embedding using an embedding model. The vector database, where document content has been pre-processed, chunked, and stored as vector embeddings, performs a similarity search to find the chunks whose embeddings are most similar to t…+
<p>The vector database serves as the key bridge between raw information and contextual understanding. It transforms unstructured data into a searchable, semantically meaningful knowledge space that helps large language models deliver more precise and relevant responses. Vector databases achiev…+
<div style="width: 810px" class="wp-caption alignnone">+
<a href="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-19758-1.jpeg" target="_blank" rel="noopener"><img src="https://d2908q01vomqb2.cloudfront.net/f1f836cb4ea6efb2a0b1b99f41ad8b103eff4b59/2026/09/09/ML-19758-1.jpeg" alt="Diagram of a RAG arch…+
<p class="wp-caption-text">Figure 1: Retrieval Augmented Generation (RAG) architecture, where documents are chunked, embedded, and stored in a vector database during ingestion, and at query time the query is embedded, similar chunks are retrieved, and passed to the LLM as context for response…+
</div> +
<h2 id="vector-store-backends-for-amazon-bedrock-knowledge-bases">Vector store backends for Amazon Bedrock Knowledge Bases</h2> +
<p>Amazon Bedrock Knowledge Bases with a customer-managed (unmanaged) configuration supports three vector store backends. For the full AWS vector portfolio covering six services, see <a href="https://aws.amazon.com/blogs/machine-learning/aws-vector-solutions-build-agentic-ai-where-your-data…+
<p><a href="https://aws.amazon.com/opensearch-service/" target="_blank" rel="noopener">Amazon OpenSearch Service</a> provides high-speed results from data held in memory. It supports high-dimensional vector embeddings with both managed cluster and serverless deployment options, and…+
<p><a href="https://aws.amazon.com/rds/aurora/" target="_blank" rel="noopener">Amazon Aurora PostgreSQL with pgvector</a> combines the high-performance relational database capabilities of Amazon Aurora with pgvector’s vector similarity search functionality. It supports multiple ind…+
<p><a href="https://aws.amazon.com/s3/features/vectors/" target="_blank" rel="noopener">Amazon S3 Vectors</a> is the AWS cloud object storage service with native vector support, designed for cost-effective storage and querying of vector embeddings at scale. It provides sub-second q…+
<p>To understand how these options perform in practice, let’s examine three distinct RAG use cases, each with different latency, cost, and search requirements, and see which vector database is the best fit for each.</p> +
<h2 id="use-case-1-product-catalog-search">Use case 1: Product catalog search</h2> +
<p>Ecommerce platforms face the challenge of helping customers find exactly what they’re looking for among thousands of products. An effective product search tool must understand natural language queries and scale to handle thousands of concurrent queries during peak shopping periods while mai…+
<h3 id="why-amazon-opensearch-is-the-best-fit-for-this-use-case">Why Amazon OpenSearch is the best fit for this use case</h3> +
<p>Amazon OpenSearch Serverless is well suited for product catalog search because it supports combining semantic understanding with traditional keyword matching through hybrid search capabilities. When dealing with large product catalogs, performance matters. Amazon OpenSearch Serverless handl…+
<p>What makes it particularly valuable for ecommerce is the built-in support for complex filtering and aggregations that power faceted navigation (think filtering by price, brand, or color). You can also choose from multiple distance metrics like cosine similarity or Euclidean distance to fine…+
<p>Amazon OpenSearch Serverless Classic collections offer several optimization options to balance cost and search quality, as detailed in the following section.</p> +
<p><strong>Note:</strong> Amazon Bedrock Knowledge Bases supports both Amazon OpenSearch Serverless and Managed Clusters. The following benchmarks were run on Serverless Classic collections. Amazon OpenSearch Serverless NextGen collections (generally available May 2026) aren’t yet …+
<h3 id="performance-analysis-and-optimizations-for-amazon-opensearch-serverless-vector-search">Performance analysis and optimizations for Amazon OpenSearch serverless vector search</h3> +
<p>Amazon OpenSearch is highly configurable and provides several configuration options. Be careful when selecting these options because they can significantly affect the performance of the vector index. We consider some of these options targeted at optimizing cost and database size and quantit…+
<p>Some of the common optimization options are:</p> <ol type="1"> -
<li>The agent specifies the correct processing order with justification: reorient to standard space, then bias field correction before skull stripping, then brain extraction with parameters tuned for healthy adults, and finally registration to the MNI152 template.</li> -
<li>It explains the critical ordering dependency. Intensity inhomogeneity at brain borders causes the skull-stripping algorithm to remove too much or too little tissue if bias correction hasn’t been applied first, particularly in temporal and frontal regions.</li> -
<li>It provides a complete bash script with error checking at each stage and quality control outputs for visual inspection of intermediate results.</li> -
<li>It documents failure modes at each step: incorrect orientation metadata, residual signal shading near surface coils, neck tissue inclusion when extraction thresholds are too permissive, and registration failure at ventricular boundaries in older subjects.</li> +
<li><strong>Size of vector embeddings</strong>: A larger vector can generally contain more semantic information about the embedded text. However, it also leads to higher memory consumption, which increases vector index size and cost. Modern embedding models like Amazon Titan Text …+
<li><strong>Data type of embeddings</strong>: We can also reduce vector index size (and thus cost) by storing embeddings in lower precision data types, such as binary embeddings. This can significantly reduce the size of the vector index.</li> +
<li><strong>Disk optimized storage</strong>: Amazon OpenSearch Serverless Classic collections offer disk-based vector search (on_disk mode) that applies 32× binary quantization internally while rescoring against full-precision vectors from disk. This preserves quality while reduci…</ol> -
<p>The skill catches the ordering dependency that would introduce systematic bias into the VBM analysis.</p> -
<p>These use cases illustrate how HCLS skills reshape agent behavior qualitatively to produce more domain-aligned responses, but there’s always a question of how much better it is for researchers and developers.</p> -
<h3 id="evaluation-results">Evaluation results</h3> -
<p>We conducted a pairwise evaluation to measure skill impact across 410 domain prompts (380 single-skill and 30 cross-skill) using two harness configurations. One of the two agent harnesses is Kiro CLI in which the Auto model is used to allow Kiro to select an optimal model for the task. The …-
<p>We employ five scoring dimensions for the large language model (LLM) judge to measure how skills impact the agent’s response. Scientific accuracy evaluates the correctness of facts, mechanisms, citations, and domain knowledge. Coherence assesses whether the response follows a logical struct…-
<p>The judge scores each dimension with 0–100 scale. However, LLM judges exhibit score compression, a phenomenon where the scores cluster in a certain range, making raw deltas (for example, +1.5) difficult to interpret. We therefore report two primary metrics. Firstly, win rate (WR), a percent…-
<p>Overall, skills win 69.5–85.9 percent of head-to-head comparisons in the two agent harness configurations. Skills improve critical thinking, actionability, and scientific accuracy in both configurations. The strongest signal is on critical thinking, confirming that skills’ primary contribut…+
<p>Depending on the indexing algorithm used, users may also configure HNSW parameters (ef_construction, m) to tune the trade-off between index build time, memory usage, and search accuracy (for practical guidance, see <a href="https://opensearch.org/blog/a-practical-guide-to-selecting-hnsw-…+
<h3 id="dataset">Dataset</h3> +
<p>We use the “Shopping Queries Data Set” (ESCI), a large dataset of difficult search queries provided by Amazon. The dataset contains 1,215,851 unique US products (title, description, bullets, and brand; approximately 1,140 characters median) and 97,345 judged queries. For each query, the dat…+
<p>Query:</p> +
<p><code>self-seal envelopes without window</code></p> +
<p>Relevant product title:</p> +
<p><code>BAZIC Security Self Seal Envelope 4 1/8" x 9 1/2" #10, No Window Tint Pattern Mailing Envelopes, Peel &amp; Seal, Office Checks Invoices (30/Pack), 1-Pack</code></p> +
<p>Irrelevant product title:</p> +
<p><code>ValBox 200 Count #8 Double Window Envelopes 3 5/8" x 8 11/16" Flip and Seal Double Window Security Check Envelopes- Security Tint Pattern Designed for Home Office Secure Mailing</code></p> +
<p>We sample 5,000 queries (approximately 19 judged products per query, approximately 17 relevant) and index all 1,215,851 product descriptions for benchmarking. We measure retrieval quality (NDCG@10), latency (p50/p95/p99 at concurrency 1 and 10), and index size (ANN in-memory footprint).<…+
<h3 id="vector-index-construction">Vector index construction</h3> +
<p>We test all combinations of embedding dimension (1024, 512, 256) and data type (float, binary), totaling six configurations, plus 1024-float in on_disk mode at the default compression_level: 32x, compared against the 1024-float in-memory baseline (seven configurations total). All indexes us…<table border="1px" width="100%" cellpadding="10px"> <tbody> <tr> -
<td><strong>Metric</strong></td> -
<td><strong>Kiro CLI</strong></td> -
<td><strong>Strands Agent</strong></td> +
<td><strong>Configuration</strong></td> +
<td><strong>Embedding size</strong></td> +
<td><strong>Embedding type</strong></td> </tr> <tr> -
<td>Prompts evaluated</td> -
<td>410</td> -
<td>410</td> +
<td>In memory</td> +
<td>1024</td> +
<td>float (baseline)</td> </tr> <tr> -
<td>Skills overall WR (d)</td> -
<td>69.5% (0.39)</td> -
<td>85.9% (0.97)</td> +
<td>In memory</td> +
<td>512</td> +
<td>float</td> </tr> <tr> -
<td>Critical thinking WR (d)</td> -
<td>78.0% (0.65)</td> -
<td>85.1% (1.03)</td> +
<td>In memory</td> +
<td>256</td> +
<td>float</td> </tr> <tr> -
<td>Scientific accuracy WR (d)</td> -
<td>69.3% (0.34)</td> -
<td>86.2% (0.85)</td> +
<td>In memory</td> +
<td>1024</td> +
<td>binary</td> </tr> <tr> -
<td>Actionability WR (d)</td> -
<td>68.0% (0.37)</td> -
<td>77.3% (0.56)</td> +
<td>In memory</td> +
<td>512</td> +
<td>binary</td> </tr> <tr> -
<td>Baseline-benefit correlation (r)</td> -
<td>-0.59</td> -
<td>-0.61</td> +
<td>In memory</td> +
<td>256</td> +
<td>binary</td> </tr> <tr> -
<td>Max variance reduction</td> -
<td>-61.9%</td> -
<td>-52.1%</td> +
<td>On disk (32×)</td> +
<td>1024</td> +
<td>float</td> </tr> </tbody> </table> -
<h4 id="effect-by-baseline-strength">Effect by baseline strength</h4> -
<p>There is strong evidence showing agent skills help the most when the base agent struggles the most. The Pearson correlation between baseline response quality and skill benefit is −0.59 in Kiro CLI and −0.61 in Strands agent. We categorize the prompts based on the baseline agent’s overall sc…-
<p>However, the strong tier finding is not absolute. Cross-domain reasoning skills achieve 80 percent win rate even at a strong baseline of 90.2, demonstrating that well-designed methodology frameworks add value across the quality spectrum when they teach a decision procedure the model wouldn’…-
<p>The overall scores for baseline and skilled agents by baseline strength are shown in the following tables.</p> -
<p>In Kiro CLI</p> +
<h3 id="evaluation-methodology">Evaluation methodology</h3> +
<p>For each configuration, we create a vector index in Amazon OpenSearch Serverless (Classic collection), ingest all 1,215,851 products, wait for merges to settle, then run an adaptive warm-up until latency stabilizes before measuring. We measure 1,000 queries × 3 repetitions at concurrency 1 …+
<ol type="1"> +
<li>Retrieval latency: Latency is measured as the time taken to retrieve relevant matches from the vector index as reported by the Amazon OpenSearch results. This doesn’t include the time to convert text to embeddings.</li> +
<li>Retrieval performance: We use the <a href="https://en.wikipedia.org/wiki/Discounted_cumulative_gain#Normalized_DCG" target="_blank" rel="noopener">Normalized Discounted Cumulative Gain (NDCG)</a> metric to score the retrievals for each query. This metric measures the quality o…+
<li>Index size: We report the ANN (Approximate Nearest Neighbor) index size, which is the in-memory structure that drives search compute cost and determines capacity requirements. This differs from total store size, which includes the _source JSON copy of each document and varies with documen…+
</ol> +
<h3 id="results">Results</h3> +
<p>Table 1: Semantic search (k-NN only) performance across seven Amazon OpenSearch Serverless configurations (1,215,851 indexed vectors, 5,000 queries, k=10). Latency is server-side at concurrency 1 unless noted. Deltas are paired bootstrap against the 1024-float in-memory baseline. See Table …+
<ol type="1"> +
<li>Reducing dimensions doesn’t always reduce quality. On this dataset, 512-float was statistically indistinguishable from the 1024-float baseline (NDCG 0.3628 vs 0.3627, p = 0.87) at half the index size (2.79 vs 5.34 GiB) and lower latency (25 vs 31 ms p50). Dropping to 256 dimensions showed…+
<li>Binarization offers large index size reductions, but the quality cost depends on the number of dimensions. At 1024 dimensions, binary embeddings reduced index size by 13.4× (0.40 vs 5.34 GiB) with a 5.2 percent NDCG loss and comparable latency (22 vs 31 ms p50). At 256 dimensions the qual…+
<li>Disk mode (on_disk 32×) preserves quality at the cost of latency. At 1024 dimensions, disk mode achieved NDCG 0.3610 (−0.5 percent vs baseline) with the same 0.40 GiB index size as 1024-binary, but at approximately 3× higher latency (99 ms p50 vs 31 ms in-memory). Both 1024-binary and on_…+
</ol> <table border="1px" width="100%" cellpadding="10px"> <tbody> <tr> -
<td><strong>Baseline Tier</strong></td> -
<td><strong>N</strong></td> -
<td><strong>Baseline (mean±sd)</strong></td> -
<td><strong>Skills (mean±sd)</strong></td> -
<td><strong>Delta</strong></td> -
<td><strong>Win Rate</strong></td> +
<td><strong>Configuration</strong></td> +
<td><strong>NDCG@10</strong></td> +
<td><strong>Δ vs baseline</strong></td> +
<td><strong>p50 (ms)</strong></td> +
<td><strong>p95 (ms)</strong></td> +
<td><strong>p99 (ms)</strong></td> +
<td><strong>Index size (ANN)</strong></td> +
<td><strong>p50 @ conc 10</strong></td> +
</tr> +
<tr> +
<td>1024 float</td> +
<td>0.3627</td> +
<td>baseline</td> +
<td>31</td> +
<td>44</td> +
<td>52</td> +
<td>5.34 GiB</td> +
<td>161 ms</td> +
</tr> +
<tr> +
<td>512 float</td> +
<td>0.3628</td> +
<td>+0.0002</td> +
<td>25</td> +
<td>37</td> +
<td>62</td> +
<td>2.79 GiB</td> +
<td>149 ms</td> +
</tr> +
<tr> +
<td>256 float</td> Diff display stops at 400 lines. The line counts above are from the whole diff. 80 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.