llm-catalog-archive

Change

a766469

a7664698f6f046e3e25dc100c87259b89a17dc7b · commit on GitHub

pytorch-blog-feed: changed (250073 bytes, HTTP 200)

raw/pytorch-blog-feed/response.xml modified

Lines added
+314
Lines removed
-356
Stored bytes at this commit
250,073
Timestamp
origin
Raw artifact at this commit
raw/pytorch-blog-feed/response.xml
Recorded headers
observed_at2026-10-01T05:54:04.027Z
origin_date2026-10-01T05:03:02.000Z
status200
final URLhttps://pytorch.org/blog/feed/
etag"b69b3761a12fc970358ce6972d4ac014"
last-modifiedWed, 30 Sep 2026 16:45:20 GMT
dateThu, 01 Oct 2026 05:54:03 GMT
age3061
cache-controlpublic, max-age=60, s-maxage=43200, stale-while-revalidate=86400, stale-if-error=604800
cf-cache-statusnull
content-encodingnull
content-length250073
@@@ -12,7 +12,7 @@
<atom:link href="https://pytorch.org/blog/feed/" rel="self" type="application/rss+xml" />
<link>https://pytorch.org</link>
<description></description>
- <lastBuildDate>Thu, 24 Sep 2026 21:06:09 +0000</lastBuildDate>
+ <lastBuildDate>Wed, 30 Sep 2026 14:39:49 +0000</lastBuildDate>
<language>en-US</language>
<sy:updatePeriod>
hourly </sy:updatePeriod>
@@@ -28,6 +28,317 @@
<height>32</height>
</image>
<item>
+ <title>A Ray-Focused Guide to PyTorch Conference North America</title>
+ <link>https://pytorch.org/blog/a-ray-focused-guide-to-pytorch-conference-north-america/</link>
+
+ <dc:creator><![CDATA[PyTorch Foundation]]></dc:creator>
+ <pubDate>Wed, 30 Sep 2026 16:45:20 +0000</pubDate>
+ <category><![CDATA[Blog]]></category>
+ <guid isPermaLink="false">https://pytorch.org/?p=171171</guid>
+
+ <description><![CDATA[TL;DR In only a few weeks, PyTorch Conference North America 2026 will begin in San Jose, bringing together the open source AI community to share ideas and collaborate. If you’ve...]]></description>
+ <content:encoded><![CDATA[<h3><span style="font-weight: 400;">TL;DR</span></h3>
+<p><span style="font-weight: 400;">In only a few weeks, PyTorch Conference North America 2026 will begin in San Jose, bringing together the open source AI community to share ideas and collaborate. If you’ve heard of Ray but never had a reason to dig in, this is the year as more teams seek to scale A…
+<p><span style="font-weight: 400;">Ray is a distributed compute framework that helps teams scale any AI framework, library or model. It provides a consistent surface for developers to go from curating petabytes of multimodal data, to training a model across multiple nodes and then deploying it in pr…
+<h2><span style="font-weight: 400;">The open source AI compute stack</span></h2>
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1298071"><b>Keynote: Evolving Ray and Kubernetes Together for the AI Era</b><b><br />
+</b></a> <i><span style="font-weight: 400;">Ion Stoica, Anyscale / Databricks / UC Berkeley · Tue 9:35 AM · Grand Ballroom</span></i></p>
+<p><span style="font-weight: 400;">Ray co-creator Ion Stoica will deliver a keynote on where Ray and Kubernetes are headed together in the AI era.</span></p>
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1256376"><b>Crossing the Divide: Co-Evolving Kubernetes and Ray for the AI Era</b><b><br />
+</b></a> <i><span style="font-weight: 400;">Jago Macleod, Google, Ion Stoica, Anyscale / Databricks / UC Berkeley · Tue 2:50 PM · Room 210BF</span></i></p>
+<p><span style="font-weight: 400;">The AI workload orchestration landscape is currently split across two massive, parallel universes: the CNCF (the bedrock of cloud native and modern infrastructure) and the PyTorch Foundation (the epicenter of AI/ML innovation). While these foundations operate indep…
+<h2><span style="font-weight: 400;">Production-scale data curation and model training</span></h2>
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1257598"><b>Ray All the Way Down: A Heterogeneous, Elastic PyTorch Training Stack at LinkedIn</b><b><br />
+</b></a> <i><span style="font-weight: 400;">Tommy Li, Tao Huang, LinkedIn · Tue 3:25 PM · Room 210AE</span></i></p>
+<p><span style="font-weight: 400;">If you want a concrete blueprint for scaling PyTorch past a single machine, this is the session to attend. LinkedIn&#8217;s team walks through a three-layer Ray-based training stack: a Ray-actor-based data loading library that streams Avro, Parquet, and Iceberg; a …
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1241540"><b>PyTorch-Native Feature Transformation and Training Framework for Uber Eats Recommendation</b></a><br />
+<i><span style="font-weight: 400;">Peng Zhang, Ke Chen, Xandra Zhu, Uber ·</span></i> <i><span style="font-weight: 400;">Oct 21 · 4:55–5:20 PM · Room LL20CD</span></i></p>
+<p><span style="font-weight: 400;">Uber&#8217;s talk on migrating the Eats recommendation stack from TensorFlow/Horovod to PyTorch leans on Ray Data as a core piece of the redesign, replacing legacy Spark-based preprocessing with distributed feature-stat computation and zero-copy batch transformatio…
+<h2><span style="font-weight: 400;">Scaling the LLM Serving Layer with Ray</span></h2>
+<p><span style="font-weight: 400;">Ray isn&#8217;t just powering dedicated training sessions this year, it&#8217;s also showing up as critical infrastructure inside talks with a different primary focus:</span></p>
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1228462"><b>Keeping GPUs Busy: High-Speed Storage for PyTorch via fsspec</b><b><br />
+</b></a> <i><span style="font-weight: 400;">Ankita Luthra, Trinadh Kotturu, Google · Tue 3:40 PM · Room LL21DEF</span></i></p>
+<p><span style="font-weight: 400;">A bottleneck has shifted from compute to storage: GPUs sitting idle while waiting on legacy REST-based data access. The proposed fix, Rapid Storage, brings a high-throughput gRPC-based protocol to PyTorch via fsspec. According to the team, the payoff extends across…
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1256403"><b>From PyTorch to Production: Serving a Physics-Constrained Generative Model with ONNX, Ray, and vLLM</b></a><b><br />
+</b><i><span style="font-weight: 400;">Arun Sharma, University of Minnesota · Tue 2:50–3:15 PM · Room 210AE</span></i></p>
+<p><span style="font-weight: 400;">A deep engineering narrative following a physics-constrained generative model &#8211; PC-RF, a conditional rectified-flow model for climate downscaling &#8211; from training through to a served, production stack. The talk covers exporting to ONNX two different ways…
+<h2><span style="font-weight: 400;">Post-training and RL with Ray</span></h2>
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1257963"><b>Building a Post-Training Platform for Teams That Don&#8217;t Own the Training Loop</b></a><br />
+<i><span style="font-weight: 400;">Gaurav Arora, Shunyao Li, Eric Wang, Pinterest · Weds 3:25–3:50 PM · Room LL20CD</span></i></p>
+<p><span style="font-weight: 400;">As post-training expanded across teams at Pinterest (SFT, DPO, GRPO on vision-language models), each built its own stack: different frameworks, data formats, and distributed strategies. Teams spent weeks integrating OSS frameworks with internal infra before trainin…
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1345867"><b>SkyRL: Democratizing Scalable RL Training</b></a><br />
+<span style="font-weight: 400;">Sumanth Hedge, Eric Tang, Anyscale </span><i><span style="font-weight: 400;">· Oct 20 4:55 PM-5:20 PM · LL20A (Lower Level)</span></i></p>
+<p><span style="font-weight: 400;">Agents have taken center stage in 2026, with ever-growing interest from companies in training custom agents with reinforcement learning. As agents shift to longer horizon, multi-turn interactions, the systems challenges and requirements on underlying training infra…
+<ul>
+<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">How SkyRL provides scalable fully async RL training on 350 billion+ parameter MoE models with Megatron and vLLM</span></li>
+<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">SkyRL&#8217;s multi-tenant Tinker Engine, which allows researchers to efficiently use their own hardware for RL training while using Tinker&#8217;s flexible training APIs to iterate on recipes</span></li>
+<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">SkyRL&#8217;s redesign towards HTTP-based APIs for scalable inference, and our contributions of native RL APIs to vLLM</span></li>
+<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Community recipes built on top of SkyRL including large MoE training on long-horizon tasks, custom recursive language models, and more.</span></li>
+</ul>
+<h2><span style="font-weight: 400;">Demo Theater</span></h2>
+<p><a href="https://events.linuxfoundation.org/pytorch-conference-north-america/program/schedule/?id=1345871"><b>Evolving Ray Core for Post-Training at Scale</b></a><br />
+<span style="font-weight: 400;">Mengjin Yan, (Ray Core), Josh Lee, Anyscale </span><i><span style="font-weight: 400;">· </span></i><span style="font-weight: 400;">Oct 21 3:55 PM-4:05 PM </span><i><span style="font-weight: 400;">·</span></i><span style="font-weight: 400;"> Community Expo (Concourse L…
+<p><span style="font-weight: 400;">Before a post-training job trains a single step, its trainers and inference engines must be scheduled onto the cluster, and every step after that, fresh weights must move between them.  Ray Core evolves directly from the workloads it powers, shaped by the post-trai…
+<h2><span style="font-weight: 400;">Meet the Developers</span></h2>
+<p><i><span style="font-weight: 400;">Tues ·  3:50 PM-4:25 PM · Community Expo</span></i></p>
+<p><span style="font-weight: 400;">Join Ray experts for an interactive “Meet the Developers” session for a chance to learn in a small group setting and ask your most pressing questions.</span></p>
+<h2><span style="font-weight: 400;">Learn more about Ray at PyTorch Conference North America</span></h2>
+<p><span style="font-weight: 400;">Taken together, these sessions tell a coherent story: Ray has moved from a scaling library you reach for to foundational infrastructure spanning orchestration, training, data loading, storage, and observability. If you&#8217;re building AI infrastructure on PyTorch…
+<p><a href="https://hubs.ly/Q04tDx8f0"><span style="font-weight: 400;">View the complete schedule here</span></a></p>
+<p><a href="https://hubs.ly/Q04tDw_W0"><span style="font-weight: 400;">Register for PyTorch Conference North America today</span></a></p>
+<p>&nbsp;</p>
+]]></content:encoded>
+
+
+
+ </item>
+ <item>
+ <title>From Upstream Changes to Downstream Confidence: Inside Torch Spyre&#8217;s Integration with PyTorch CRCR</title>
+ <link>https://pytorch.org/blog/from-upstream-changes-to-downstream-confidence-inside-torch-spyres-integration-with-pytorch-crcr/</link>
+
+ <dc:creator><![CDATA[Mehant Kammakomati (IBM), Jewel K M (Red Hat), Anubhav Jana (IBM), Padmanabha Venkatagiri Seshadri (IBM)]]></dc:creator>
+ <pubDate>Wed, 30 Sep 2026 13:40:56 +0000</pubDate>
+ <category><![CDATA[Blog]]></category>
+ <guid isPermaLink="false">https://pytorch.org/?p=171223</guid>
+
+ <description><![CDATA[TL;DR PyTorch&#8217;s Cross-Repository CI Relay (CRCR) gives out-of-tree accelerators a clean, scalable way to plug into upstream CI &#8211; and it leaves each backend free to decide which of PyTorch&#8217;s...]]></description>
+ <content:encoded><![CDATA[<h2>TL;DR</h2>
+<p>PyTorch&#8217;s Cross-Repository CI Relay (CRCR) gives out-of-tree accelerators a clean, scalable way to plug into upstream CI &#8211; and it leaves each backend free to decide which of PyTorch&#8217;s tens of thousands of tests actually matter for its hardware, and how to keep that answer curren…
+<h2>Introduction</h2>
+<p>PyTorch&#8217;s <a href="https://pytorch.org/blog/introducing-cross-repository-ci-relay-scalable-ci-for-pytorchs-out-of-tree-backends/">Cross-Repository CI Relay</a> (CRCR) closes the coordination gap between PyTorch and out-of-tree (OOT) accelerator repositories by providing a standardized way t…
+<p>The interesting engineering starts one step later, in the decisions the relay leaves to each backend: what are the test candidates, which dispatches deserve a build, which of PyTorch&#8217;s tens of thousands of tests are meaningful on your hardware, how to adapt tests written for CUDA without fo…
+<h2>Three Evolving Candidates under Test: OOT Accelerator PyTorch Backend, PyTorch Core, Test Suite</h2>
+<p><!-- IMAGE PLACEHOLDER: Figure 1 - Testing Surface in PyTorch and Torch Spyre (re-upload via WP media library) --></p>
+<p><em><img fetchpriority="high" decoding="async" class="alignnone wp-image-171228 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/image-1-1.png" alt="" width="1886" height="828" srcset="https://pytorch.org/wp-content/uploads/2026/09/image-1-1.png 1886w, https://pytorch.org/wp-content…
+<p><em>Figure 1: Testing Surface in PyTorch and Torch Spyre</em></p>
+<h3>Challenge 1: Deciding what to test: three moving targets, and which combination to run</h3>
+<p>Often, OOT accelerators have to deal with three evolving candidates to test after receiving a dispatch from <a href="https://pytorch.org/blog/introducing-cross-repository-ci-relay-scalable-ci-for-pytorchs-out-of-tree-backends/">PyTorch CRCR</a>. First, an OOT accelerator PyTorch backend codebase,…
+<p>Knowing the three candidates leads directly to the next decision: which combination of them to actually run. When a dispatch arrives, the OOT accelerator backend may have progressed since the last one, PyTorch core has new commits, and the test suite itself may have changed. The primary combinati…
+<p>Optionally, this can be expanded to include other useful combinations, such as testing a range of OOT accelerator backend versions against a moving PyTorch core and test suite. This can help uncover forward- and backward-compatibility issues for each OOT accelerator backend version. Further, each…
+<h2>Taming the PyTorch Test Suite</h2>
+<h3>Challenge 2: Identifying the candidate tests for your OOT accelerator backend</h3>
+<p><!-- IMAGE PLACEHOLDER: Figure 2 - Agentic Workflow for Test Selection (re-upload via WP media library) --></p>
+<p><em><img decoding="async" class="alignnone wp-image-171231 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/image-2-1.png" alt="" width="1468" height="520" srcset="https://pytorch.org/wp-content/uploads/2026/09/image-2-1.png 1468w, https://pytorch.org/wp-content/uploads/2026/09/imag…
+<p><em>Figure 2: Agentic Workflow for Test Selection</em></p>
+<p>PyTorch has tens of thousands of tests. Manually identifying which ones are relevant to your backend doesn&#8217;t scale and keeping that selection current as both PyTorch and your backend evolve is worse. Hence, we built an agentic pipeline to do it.</p>
+<p>The pipeline runs in four stages (refer Figure 2). The first two stages build context; the last two make decisions: first from code, then from execution.</p>
+<ol>
+<li>High-level selection. A first agent narrows the search space from the whole test tree to the folders and top-level files worth analysing, using two inputs: which PyTorch extension points the backend actually hooks into, and plain-language preferences about scope. A backend that does not hook <co…
+<li>Repository memory. A Repository Memory Generator turns that shortlist into a queryable index which includes symbols, files, LLM-written summaries and per-test embeddings. This lets the selection agent in the next phase to reason over the whole candidate set at once instead of being limited to wh…
+<li>Low-level selection. The second agent then reads the backend&#8217;s own codebase, its documentation, and metadata (e.g., supported operators), then queries the repository exact tests to run and buckets them into <code>mandatory_success</code> or <code>skip</code>. Every decision is recorded wit…
+<li>Refinement from real execution. The selected tests run on real hardware. The agent takes a second pass using execution logs that leads to catching runtime failures and numerical differences that static analysis misses. The output is a corrected, re-bucketed test set with updated reasoning.</li>
+</ol>
+<p>This pipeline currently produces a per-file config for each upstream test file we run, covering thousands of individually named test cases, and a merge run exercises them on real hardware. Coverage grows automatically enabling a new op in the backend to admit every test that was waiting on it, wi…
+<p>Auditability matters: Every bucketing choice is written back into the config as a comment explaining why. When someone asks &#8220;why is this test skipped?&#8221;, the answer is in the file, not in a model&#8217;s memory. The agent proposes; the config is the reviewable artifact of record.</p>
+<p>Furthermore, the agentic pipeline can be customized for specific use cases. Two such use cases are discussed below.</p>
+<h3>Use Case A: Enabling a New Op</h3>
+<p><em>When a new op is enabled in the backend, which tests should I turn on?</em></p>
+<p>This involves preparing a test-to-op mapping, where, for each test, we record the operators it exercises and build a reverse index that maps each operator to the complete list of tests that exercise it. The Repository Memory Generator captures this metadata during a one-time execution pass:</p>
+<ol>
+<li>Eager path: captured via TorchDispatchMode</li>
+<li>Compile path: captured via TORCH_LOGS</li>
+</ol>
+<p>However, the tests captured through this process cannot be added directly to the test suite. Although a test may exercise a desired operator, it may also exercise other PyTorch features that are either not relevant to the OOT accelerator or are not currently supported. Therefore, the operator inf…
+<p>Result: Enabling one op surfaces all relevant tests without any manual curation.</p>
+<h3>Use Case B: Upgrading PyTorch Versions</h3>
+<p><em>When upgrading to a new PyTorch version, how can test selection be updated without reprocessing everything?</em></p>
+<p>Across PyTorch versions, the test suite can evolve as new tests are added and existing tests are modified or removed. These changes can be incorporated into the repository memory to reflect the updated test suite. The low-level selection agent can then be run only on the delta that is the newly a…
+<p>Result: Version upgrades become config diffs, not full reprocessing. For example, upgrading from PyTorch 2.13 to 2.14 meant evaluating ~4000 changed tests instead of tens of thousands.</p>
+<h3>Challenge 3: Adapting the candidate tests</h3>
+<p>PyTorch tests are often parameterized across dtypes, shapes, and operators. Skipping an entire test because one dtype isn&#8217;t supported is too coarse, in such cases, only that dtype is excluded while the rest run. Deeper adaptations (e.g., removing CUDA-specific hardcoding) traditionally requ…
+<p>The framework provides three controls:</p>
+<ol>
+<li>Parameter-level knobs: Subselect dtypes, exclude specific shapes, add coverage for a dtype of interest.</li>
+<li>Outcome bucketing: Assign each test to mandatory_success (must pass), xfail (expected to fail), xfail_strict (expected to fail, and an unexpected pass is treated as failure), or skip.</li>
+<li>Capability-driven inclusion: Declare supported ops/dtypes globally; enabling a new op automatically admits all tests waiting on it.</li>
+</ol>
+<p>Concretely, this is what the declarative interface looks like in practice. A file entry sets a default policy for anything not explicitly listed, then lists tests in buckets, and an edits: block adapts individual cases without touching upstream source:</p>
+<pre><code class="language-yaml">test_suite_config:
+ labels: [trunk]
+ files:
+ - path: ${TORCH_ROOT}/test/test_view_ops.py
+ unlisted_test_mode: skip # default policy for unlisted tests
+ tests:
+ # Basic tensor metadata manipulation which is fully supported on Spyre.
+ - names:
+ - TestOldViewOps::test_broadcast_tensors
+ - TestViewOps::test_contiguous_self
+ mode: mandatory_success
+ # bool/int64 fail only because test setup uses aten::random_.from.
+ - names:
+ - TestOldViewOps::test_broadcast_to
+ mode: mandatory_success
+ edits:
+ dtypes:
+ exclude:
+ - name: bool
+ - name: int64
+ global:
+ supported_dtypes: [{name: bfloat16}, {name: float16}, {name: float32}]
+ supported_ops:
+ - name: _scaled_mm
+ dtypes: [{name: bfloat16}]
+</code></pre>
+<p>Given these additional user-facing controls, the agentic workflow can be extended to generate the corresponding configuration files that can be consumed by the OOT test reuse framework. The low-level agent along with test repository memory can then capture not only test outcomes (e.g., <code>mand…
+<p>Four properties of the schema do most of the work:</p>
+<ol>
+<li>A declared default. <code>unlisted_test_mode</code> sets the outcome for anything unlisted, so a test added upstream tomorrow has a defined result instead of breaking CI the day it lands.</li>
+<li>Four outcome buckets. <code>mandatory_success</code>, <code>xfail</code>, <code>xfail_strict</code> and <code>skip</code>. The strict variant earns its place: an unexpected pass surfaces as a signal to promote the test, not silence.</li>
+<li>Per-case adaptation. <code>edits:</code> applies to individual dtypes, ops and modules, so an unsupported dtype costs one dtype rather than the whole test.</li>
+<li>Capability declared once. <code>global.supported_ops</code> and <code>supported_dtypes</code> are declared per file, so enabling an operator in the backend is a one-line change that admits every test waiting on it.</li>
+</ol>
+<p>How it works: The framework patches upstream&#8217;s <code>@ops</code>, <code>@modules</code>, and <code>@dtypes</code> decorators at collection time, emitting pytest marks (<code>op__</code>, <code>dtype__</code>). This means the same config drives both adaptation and selection: <code>-m op__add…
+<p>Key guarantee: The upstream test tree stays pristine. No patches against PyTorch&#8217;s tests. A version bump becomes a config delta, not a merge conflict. The framework targets the generic privateuse1 device, so nothing is hardware-specific.</p>
+<h2>CRCR Integration into Existing Workflows</h2>
+<p><!-- IMAGE PLACEHOLDER: Figure 3 - Workflow with CRCR Integration for Torch Spyre (re-upload via WP media library) --></p>
+<p><em><img decoding="async" class="alignnone wp-image-171232 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/image-3-1-scaled.png" alt="" width="2560" height="1147" srcset="https://pytorch.org/wp-content/uploads/2026/09/image-3-1-scaled.png 2560w, https://pytorch.org/wp-content/uploa…
+<p><em>Figure 3: Workflow with CRCR Integration for Torch Spyre</em></p>
+<h3>Challenge 4: Consuming the dispatch: gate on the payload, resolve the SHA yourself</h3>
+<table>
+<thead>
+<tr>
+<th>Dispatch Field</th>
+<th>Purpose</th>
+</tr>
+</thead>
+<tbody>
+<tr>
+<td>SHA</td>
+<td>Identifies the exact upstream commit to validate.</td>
+</tr>
+<tr>
+<td>PR Number</td>
+<td>Associates downstream test results with a specific upstream PR.</td>
+</tr>
+<tr>
+<td>Action</td>
+<td>Specifies whether the PR was opened, updated (synchronize), or closed/merged.</td>
+</tr>
+<tr>
+<td>Base Branch (<code>pull_request.base.ref</code>)</td>
+<td>Identifies the target branch and enables filtering for main and release branches such as <code>release/{version}</code>.</td>
+</tr>
+<tr>
+<td>PR Label (<code>pull_request.labels</code>)</td>
+<td>Labels used with the pull request. Can be used to identify merged PRs through label &#8220;Merged&#8221;.</td>
+</tr>
+</tbody>
+</table>
+<p><em>Table 1: Key Dispatch Fields</em></p>
+<p>Integrating CRCR starts with consuming the dispatch from PyTorch. The dispatch payload contains key pieces of information as shown in the Table 1. Dispatches can be filtered by consuming these pieces of information for various use cases. Filtering dispatches for pull requests merged to main requi…
+<pre><code class="language-bash">HAS_MERGED=$(echo "$PAYLOAD_JSON" | jq -r '.payload.pull_request.labels // [] | map(.name) | index("Merged") // ""')
+</code></pre>
+<p>Edge case: A race condition can cause the label to be applied after the dispatch, or a manual merge may skip the label entirely. Defensive handling (e.g., polling or a fallback heuristic) may be needed for critical workflows.</p>
+<p>Apart from this dispatch-driven use case, there are two other use cases that may be of interest to OOT accelerators: testing against PyTorch nightly builds and testing against PyTorch releases.</p>
+<h3>Use Case C: Nightly Testing</h3>
+<p>CRCR doesn&#8217;t dispatch for nightlies, but HUD supports reporting nightly results. The workflow runs on a schedule and resolves the SHA directly as below:</p>
+<pre><code class="language-bash">NIGHTLY_SHA=$(curl -fsSL "https://api.github.com/repos/pytorch/pytorch/commits?sha=nightly&amp;per_page=1" | jq -r '.[0].sha // empty')
+
+if [ -z "${NIGHTLY_SHA}" ]; then
+ echo "::error::Could not resolve HEAD for pytorch/pytorch/nightly"
+ exit 1
+fi
+
+COMMIT_MSG=$(curl -fsSL "https://api.github.com/repos/pytorch/pytorch/commits/${NIGHTLY_SHA}" | jq -r '.commit.message')
+
+SOURCE_SHA=$(echo "$COMMIT_MSG" | grep -oP '\(([a-f0-9]{40})\)' | tr -d '()')
+
+if [ -z "${SOURCE_SHA}" ]; then
+ echo "::error::Could not extract source SHA from latest nightly commit"
+ exit 1
+fi
+
+echo "Resolved pytorch/pytorch@nightly -&gt; source ${SOURCE_SHA}"
+</code></pre>
+<h3>Use Case D: Release Testing</h3>
+<p>No dispatch is sent when a PyTorch release is published. Release testing is triggered manually via workflow_dispatch, passing the release branch (e.g., release/v2.14). SHA resolution follows the same pattern as nightly.</p>
+<p>Note: HUD does not currently have a dedicated view for release test results.</p>
+<h3>Challenge 5: Enhancing workflow efficiency: test splitting, parallelization, and build-once-test-many</h3>
+<p>Thousands of tests need to run fast. Splitting happens at three levels:</p>
+<ol>
+<li>Feature-level: Group tests by area (Inductor, operators, eager, etc.)</li>
+<li>Duration-constrained: Subdivide each feature group to stay under a max job duration (e.g., 30 min)</li>
+<li>Balanced bucketing: Distribute individual tests across splits to avoid stragglers</li>
+</ol>
+<p>Building this requires a one-pass timing run to measure per-test and per-feature durations. The result: parallel jobs that finish together, not one job holding up the rest.</p>
+<p>Build-once, test-many: Build PyTorch and backend wheels once, publish as workflow artifacts, consume across all test splits. This avoids redundant builds and guarantees all jobs test the same artifacts.</p>
+<h3>Challenge 6: Enhancing workflow reliability</h3>
+<p><!-- IMAGE PLACEHOLDER: Figure 4 - Retry Patterns for Workflow Resilience (re-upload via WP media library) --></p>
+<p><em><img decoding="async" class="alignnone wp-image-171235 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/image-4-1.png" alt="" width="1374" height="822" srcset="https://pytorch.org/wp-content/uploads/2026/09/image-4-1.png 1374w, https://pytorch.org/wp-content/uploads/2026/09/imag…
+<p><em>Figure 4: Retry Patterns for Workflow Resilience</em></p>
+<p>The final state of the workflow in HUD should focus on surfacing regressions introduced by the evolving PyTorch and OOT accelerator codebases, rather than workflow stability failures such as transient infrastructure issues that may not be meaningful to the broader community.</p>
+<p>While infrastructure and platform layers could be equipped with multiple reliability measures to improve workflow stability, developers can further strengthen the workflow by adopting additional reliability patterns, as shown in Table 2. The patterns listed here are not intended to be exhaustive,…
+<table>
+<thead>
+<tr>
+<th>Reliability Dimension</th>
+<th>Pattern</th>
+<th>Impact from CRCR PoV</th>
+<th>Implementation Detail</th>
+</tr>
+</thead>
+<tbody>
+<tr>
+<td>Fault tolerance</td>
+<td>Parallel and isolated execution pattern</td>
+<td>Non-faulty test jobs continue to complete and report their status even when another job fails.</td>
+<td>GitHub Actions matrix jobs with pods/containers at the platform layer for execution isolation.</td>
+</tr>
+<tr>
+<td>Fault tolerance</td>
+<td>Continue on error pattern</td>
+<td>Failures in non-critical paths do not cause the overall workflow to be reported as failed in HUD.</td>
+<td>Use GitHub Actions <code>continue-on-error</code> for non-critical jobs or steps.</td>
+</tr>
+<tr>
+<td>Resilience</td>
+<td>Retry pattern</td>
+<td>Transient failures can be recovered without surfacing persistent failures in HUD.</td>
+<td>As shown in Figure 4, use workflow-level retries and finer-grained in-job retries based on test logs and failure heuristic.</td>
+</tr>
+<tr>
+<td>Resilience</td>
+<td>Logging and failure classification</td>
+<td>Distinguishes actionable test regressions from transient infrastructure or platform failures not recoverable from retry pattern.</td>
+<td>Verbose logging with summaries can help classify failures (that passed through retry pattern) and act for recovery.</td>
+</tr>
+<tr>
+<td>Consistency</td>
+<td>Build-once, use-many</td>
+<td>Parallel jobs test the same PyTorch and OOT accelerator artifacts, reducing inconsistencies across splits.</td>
+<td>Build wheels once, publish them as GitHub workflow artifacts, and consume the same artifacts across all test jobs.</td>
+</tr>
+</tbody>
+</table>
+<p><em>Table 2: Workflow Reliability Patterns</em></p>
+<h3>Challenge 7: CRCR Callbacks with Matrix Jobs</h3>
+<p>Github provides the ability to run a set of jobs based on a combination of variables and refers to that as a matrix strategy. Individual jobs as part of a matrix strategy are called legs.</p>
+<p>The constraint: A matched in_progress/completed callback pair must originate from the same job. CRCR uses check_run_id to track retries, and each matrix leg gets its own check_run_id. This means callbacks cannot be sent from a parent job that spawns a matrix.</p>
+<p>Two workarounds:</p>
+<p>Option A: Per-leg callbacks. Each matrix job sends its own in_progress and completed. Simple, but can clutter HUD for large matrices and may not be ideal for dynamic jobs.</p>
+<p>Option B: Dispatch + poll. The parent job dispatches a separate workflow containing the matrix, then polls for completion. This pattern also applies to heterogeneous CI such as triggering Jenkins or other platforms from a GitHub Actions parent job.</p>
+<h3>Challenge 8: Making a run reproducible weeks later</h3>
+<p>Debugging a failure from two weeks ago requires knowing exactly what ran. Capture this metadata as workflow artifacts:</p>
+<ol>
+<li>Dispatch payload (PR number, SHA, action)</li>
+<li>PyTorch and backend wheel versions (or commit SHAs)</li>
+<li>Runtime environment (Python, CUDA, OS, CPU arch)</li>
+<li>Per-test results (pass/fail/skip, duration)</li>
+</ol>
+<p>Combined with build-once artifacts, this ensures any run can be reproduced or bisected weeks later.</p>
+<h2>Conclusion</h2>
+<p>CRCR handles distribution and reporting. The decisions it leaves to the downstream repo are the right ones to own: which dispatches matter, which tests are meaningful, and what &#8220;green&#8221; means on a given platform.</p>
+<p>Owning those decisions meant keeping three moving targets in sync: the backend, PyTorch core, and the test suite. That&#8217;s where the engineering in this post went: an agentic test-selection pipeline, a declarative test-reuse framework, build-once-promote-many, and layered retries.</p>
+<p>None of this is hardware-specific. The test-reuse framework targets the generic privateuse1 device. We intend to work with the PyTorch team to upstream it — most naturally alongside Torch OpenReg, the in-tree reference backend that accelerator authors already use to learn the extension points.</p…
+<p>If any of these patterns would help your integration, we&#8217;d like to hear which ones matter most before shaping that proposal. Reach out by opening an issue at <a href="https://github.com/torch-spyre/torch-spyre">torch-spyre/torch-spyre</a>.</p>
+<h2>Acknowledgements</h2>
+<p>We would like to thank the IBM Spyre Team for their support in this effort.</p>
+]]></content:encoded>
+
+
+
+ </item>
+ <item>
<title>Accelerate Your AI Journey with new Introduction Track at PyTorch Conference NA 2026 and PyTorch Associate Training</title>
<link>https://pytorch.org/blog/accelerate-your-ai-journey-with-new-introduction-track-at-pytorch-conference-na-2026-and-pytorch-associate-training/</link>
@@@ -105,7 +416,7 @@
<guid isPermaLink="false">https://pytorch.org/?p=170102</guid>
<description><![CDATA[Some of the most interesting work being built with PyTorch starts in universities, research labs, student groups, and academic institutions. A new model architecture. A library created to support a...]]></description>
- <content:encoded><![CDATA[<p><span style="font-weight: 400;"><img fetchpriority="high" decoding="async" class="aligncenter wp-image-170129 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/PyTorch-logo-to-PyTorch-Foundation-logo-1.png" alt="" width="1983" height="793" srcset="…
+ <content:encoded><![CDATA[<p><span style="font-weight: 400;"><img decoding="async" class="aligncenter wp-image-170129 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/PyTorch-logo-to-PyTorch-Foundation-logo-1.png" alt="" width="1983" height="793" srcset="https://pytorch.org/w…
<p><span style="font-weight: 400;">Some of the most interesting work being built with PyTorch starts in universities, research labs, student groups, and academic institutions.</span></p>
<p><span style="font-weight: 400;">A new model architecture. A library created to support a paper. A benchmarking framework. A research tool that solves a problem no existing project addresses. An educational toolkit developed for a university course. A system that begins with three graduate student…
<p><b>What happens after the research project works?</b></p>
@@@ -1198,361 +1509,8 @@ Tue 20 Oct, 16:20–16:30, LL20AB</p>
- </item>
- <item>
- <title>Helion x 🤗 HF Kernels: Building and Shipping Out-of-the-box Performant Kernels</title>
- <link>https://pytorch.org/blog/helion-x-%f0%9f%a4%97-hf-kernels-building-and-shipping-out-of-the-box-performant-kernels/</link>
-
- <dc:creator><![CDATA[Sayak Paul (Hugging Face), Dunfan Lu (Meta), Tarindu Jayatilaka (Meta), Jongsok Choi (Meta)]]></dc:creator>
- <pubDate>Fri, 11 Sep 2026 17:45:57 +0000</pubDate>
- <category><![CDATA[Blog]]></category>
- <guid isPermaLink="false">https://pytorch.org/?p=165024</guid>
-
- <description><![CDATA[TL;DR The HuggingFace Kernels project now has Helion support. This blog walks through how to build, autotune, and ship performant and portable Helion kernels via the Hugging Face Kernels project,...]]></description>
- <content:encoded><![CDATA[<h3><img decoding="async" class="alignnone wp-image-164371 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/All-PyTorch-Blog-Social-Images-21.png" alt="" width="1920" height="1080" srcset="https://pytorch.org/wp-content/uploads/2026/09/All-PyTorch-Bl…
-<h3>TL;DR</h3>
-<p>The HuggingFace Kernels project now has Helion support. This blog walks through how to build, autotune, and ship performant and portable Helion kernels via the Hugging Face Kernels project, allowing users to consume these kernels seamlessly.</p>
-<h2>Introduction</h2>
-<p><a href="https://github.com/pytorch/helion">Helion</a> is a high-level DSL for writing high-performance, portable kernels for machine learning. The <a href="https://huggingface.co/docs/kernels/en/index"><img src="https://s.w.org/images/core/emoji/17.0.2/72x72/1f917.png" alt="🤗" class="wp-smiley"…
-<p>In this post, we will discuss how Helion is supported within the Kernels project, how users can benefit from first-class autotuning support in Helion, and how to ship pre-tuned kernel configs to reduce cold-start times. We will also show examples of Helion kernels and how tuning them for specific…
-<p>P.S.: Throughout the rest of the post, we will refer to the <strong>Kernels</strong> project with “k” in capital letters to distinguish it from actual “kernels”.</p>
-<h2>Intro: Helion</h2>
-<p>Helion is a tiled DSL for writing performant ML kernels. The programming model is often described as &#8220;PyTorch with tiles” – the kernel operates on PyTorch tensors, and tile-level operations are specified via ordinary PyTorch tensor operators. As a quick example, the following function shows…
-<pre><code class="language-python">import torch, helion, helion.language as hl
-
-@helion.kernel()
-def matmul(x: torch.Tensor, y: torch.Tensor) -&gt; torch.Tensor:
- m, k = x.size()
- k, n = y.size()
- out = torch.empty([m, n], dtype=x.dtype, device=x.device)
-
- for tile_m, tile_n in hl.tile([m, n]):
- acc = hl.zeros([tile_m, tile_n], dtype=torch.float32)
- for tile_k in hl.tile(k):
- acc = torch.addmm(acc, x[tile_m, tile_k], y[tile_k, tile_n])
- out[tile_m, tile_n] = acc
-
- return out
-</code></pre>
-<p>What makes Helion desirable is not just its concise syntax but what it leaves deliberately unspecified. When you write <span style="font-weight: 400;"><code>hl.tile</code></span>, you say only that the iteration space should be tiled &#8211; not how large the tiles are, or how their data is fetch…
-<p>This autotuning process is why a Helion kernel can often outperform a hand-written kernel in a lower-level language, when benchmarked on a large set of shapes. With that said, autotuning is sometimes a lengthy process, so it is beneficial to have an established approach for shipping a kernel bund…
-<h2>Intro: Kernels</h2>
-<p>The current landscape of kernel packaging and distribution is fragmented, characterized by inconsistent source structures, disparate tooling, and limited compatibility support. Consequently, users often face arduous build times, even when pre-built wheels are available.</p>
-<p>The Kernels project addresses these challenges by establishing a standardized, unified packaging and build process for both AoT and JIT kernels. The project is divided into two primary components:</p>
-<ul>
-<li><span style="font-weight: 400;"><strong><code>kernel-builder</code></strong>: </span> A tool for developers to reliably package and distribute kernels across different framework versions and system configurations. It enforces standards to ensure predictable source structures, build reproducibili…
-<li><span style="font-weight: 400;"><b><code>kernels</code></b>: </span>A consumer-facing Python library that allows users to effortlessly load ready-to-use kernels without dependency management issues via a simple command like <code>get_kernel("org/name", version=1)</code>, much like pulling a mode…
-</ul>
-<p>For kernel users, we want to provide a seamless experience of loading kernels and getting them ready to use right away. Let’s take a look at an example of how one could load the popular Flash-Attention 3 kernel:</p>
-<pre><code class="language-python">from kernels import get_kernel
-
-kernel_module = get_kernel("kernels-community/flash-attn3", version=1)
-flash_attn_func = kernel_module.flash_attn_func
-
-flash_attn_func(...)
-</code></pre>
-<p>We provide prebuilt binaries for a comprehensive compatibility matrix of ahead-of-time kernels, such as Flash Attention 3. This is quite beneficial to end users, particularly when the kernel&#8217;s upstream repository may not have a specific build available.</p>
-<p>Users can browse a wide variety of kernels on the Hugging Face Hub platform: <a href="http://hf.co/kernels">hf.co/kernels</a>:</p>
-<p><img decoding="async" class="alignleft wp-image-164983 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Screenshot-2026-09-09-161052.png" alt="" width="1280" height="774" srcset="https://pytorch.org/wp-content/uploads/2026/09/Screenshot-2026-09-09-161052.png 1280w, https://pytorch.o…
-<p>We refer to this collection of kernels as the <strong>Kernels Hub</strong>.</p>
-<h2>Packaging and using Helion in Kernels</h2>
-<p>Helion kernels are plain Python. They compile themselves the first time you call them, so there is nothing for <span style="font-weight: 400;"><code>kernel-builder</code> </span>to compile ahead of time. Helion provides utilities to tune these kernels for specific workloads and hardware (more on …
-<p>In this section, we discuss how to scaffold and structure a Helion kernel for building with <span style="font-weight: 400;"><code>kernel-builder</code></span><span style="font-weight: 400;">. </span></p>

Diff display stops at 400 lines. The line counts above are from the whole diff. 47 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.