Change
7aec857
7aec857fb9bae8638dbe44456719bff10c26c30c · commit on GitHub
pytorch-blog-feed: changed (204930 bytes, HTTP 200)
raw/pytorch-blog-feed/response.xml modified
- Source
- pytorch-blog-feed
- Lines added
- +260
- Lines removed
- -344
- Stored bytes at this commit
- 204,930
- Timestamp
- origin
- Raw artifact at this commit
- raw/pytorch-blog-feed/response.xml
Recorded headers
| observed_at | 2026-10-03T05:18:56.230Z |
|---|---|
| origin_date | 2026-10-03T05:11:04.000Z |
| status | 200 |
| final URL | https://pytorch.org/blog/feed/ |
| etag | "7413570e4dbbf53684980b13b562fe9c" |
| last-modified | Fri, 02 Oct 2026 19:57:16 GMT |
| date | Sat, 03 Oct 2026 05:18:56 GMT |
| age | 472 |
| cache-control | public, max-age=60, s-maxage=43200, stale-while-revalidate=86400, stale-if-error=604800 |
| cf-cache-status | null |
| content-encoding | null |
| content-length | 204930 |
@
@@ -12,7 +12,7 @@ <atom:link href="https://pytorch.org/blog/feed/" rel="self" type="application/rss+xml" /> <link>https://pytorch.org</link> <description></description>-
<lastBuildDate>Thu, 01 Oct 2026 22:32:14 +0000</lastBuildDate>+
<lastBuildDate>Fri, 02 Oct 2026 19:57:16 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod>@
@@ -28,6 +28,263 @@ <height>32</height></image> <item>+
<title>Building a High-Performance and Portable vLLM Linear Backend with Helion</title>+
<link>https://pytorch.org/blog/building-a-high-performance-and-portable-vllm-linear-backend-with-helion/</link>+
+
<dc:creator><![CDATA[Sean Chen (Red Hat) and Shangdi Yu (PyTorch, Meta Platforms)]]></dc:creator>+
<pubDate>Fri, 02 Oct 2026 19:55:07 +0000</pubDate>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=171767</guid>+
+
<description><![CDATA[TL;DR We integrated Helion into vLLM’s linear backend to explore how an autotuned, high-level kernel DSL can improve LLM inference performance while reducing kernel implementation complexity. A single Helion general...]]></description>+
<content:encoded><![CDATA[<h2><span style="font-weight: 400;">TL;DR</span></h2>+
<p><span style="font-weight: 400;">We integrated </span><a href="https://helionlang.com/index.html"><span style="font-weight: 400;">Helion</span></a><span style="font-weight: 400;"> into vLLM’s </span><a href="https://docs.vllm.ai/en/latest/api/vllm/model_executor/kernels/linear/"><span style="font-…+
<p><span style="font-weight: 400;">On NVIDIA Hopper GPUs, the Helion linear backend combines per-shape tuning with hybrid dispatch to outperform the vLLM default CUTLASS and DeepGEMM backends across the evaluated models, delivering consistent end-to-end performance gains and more than 10% throughput…+
<h2><span style="font-weight: 400;">Brief Background on vLLM and Helion</span></h2>+
<p><a href="https://docs.vllm.ai/en/latest/"><span style="font-weight: 400;">vLLM</span></a><span style="font-weight: 400;"> is a high-performance inference and serving framework for large language models (LLMs). For quantized linear layers, such as FP8, INT8, INT4, and NVFP4, vLLM provides speciali…+
<p><a href="https://helionlang.com/index.html"><span style="font-weight: 400;">Helion</span></a><span style="font-weight: 400;"> is a PyTorch-native hardware agnostic kernel DSL designed for writing high-performance kernels using a tile-programming model. Its high-level abstraction allows developers…+
<h2><span style="font-weight: 400;">Helion Adoption Opportunities and Challenges</span></h2>+
<p><span style="font-weight: 400;">This section summarizes the key opportunities and challenges of adopting Helion kernels in inference engines such as vLLM, providing a high-level overview of Helion’s value proposition and tradeoffs.</span></p>+
<h3><span style="font-weight: 400;">Opportunities</span></h3>+
<p><b>Performance</b><span style="font-weight: 400;">: Our </span><a href="https://pytorch.org/blog/portable-vllm-model-inference-kernels-in-helion/"><span style="font-weight: 400;">previous work</span></a><span style="font-weight: 400;"> demonstrated Helion’s potential to achieve SOTA performance f…+
<p><b>Systematic Tuning Framework</b><span style="font-weight: 400;">: Compared with more open-ended agentic approaches that iteratively generate, profile, and refine kernel implementations, Helion formulates kernel tuning as a structured numerical optimization problem over a well-defined search spa…+
<p><b>Portability and Abstraction</b><span style="font-weight: 400;">: Helion is not only a portable DSL. It is possible to maintain a single kernel implementation while optimizing it for different workload patterns and hardware platforms. </span></p>+
<p><b>Client-Side Kernel Optimization</b><span style="font-weight: 400;">: Default kernels in inference engines such as vLLM are typically optimized for general workloads and commonly used models. With Helion, users can further tune kernel performance for their specific deployments without requiring…+
<h3><span style="font-weight: 400;">Challenges</span></h3>+
<p><span style="font-weight: 400;">Helion still has some challenges integrating into inference engines, with ongoing work from the Helion team to address and mitigate them.</span></p>+
<p><b>Ahead-of-time kernel tuning overhead</b><span style="font-weight: 400;">: Helion automates and systematizes kernel tuning, but fine-grained tuning can still take hours. This overhead comes primarily from the granularity of the tuning strategy rather than from Helion itself. Achieving the same …+
<p><b>Startup-time overhead</b><span style="font-weight: 400;">: CUDA Graph capture during vLLM startup triggers Helion JIT compilation, increasing cold-start latency. This overhead can be largely eliminated on warm starts by caching compiled artifacts.</span></p>+
<p><b>Inference-runtime overhead</b><span style="font-weight: 400;">: Outside the CUDA Graph capture range, Helion kernel dispatch and launch during inference can introduce additional CPU overhead, potentially offsetting the performance gains from fine-grained tuning. In practice, Helion is most eff…+
<p><b>Maintenance overhead</b><span style="font-weight: 400;">: Shipping pre-tuned configs for popular models creates an ongoing upstream maintenance burden. Large config files are difficult to maintain and impractical to validate exhaustively through unit tests and CI. </span></p>+
<h3><span style="font-weight: 400;">The Tradeoff Triangle</span></h3>+
<p><span style="font-weight: 400;">Fine-grained kernel tuning with Helion presents a tradeoff among </span><b>performance</b><span style="font-weight: 400;">, </span><b>usability</b><span style="font-weight: 400;">, and </span><b>maintainability</b><span style="font-weight: 400;">. This tradeoff is …+
<p><img fetchpriority="high" decoding="async" class="aligncenter wp-image-171928 " src="https://pytorch.org/wp-content/uploads/2026/10/1.png" alt="" width="587" height="587" srcset="https://pytorch.org/wp-content/uploads/2026/10/1.png 1254w, https://pytorch.org/wp-content/uploads/2026/10/1-300x300.p…+
<p style="text-align: center;"><i><span style="font-weight: 400;">Fig. 1: The Performance-Usability-Maintainability tradeoff triangle</span></i></p>+
<p><span style="font-weight: 400;">Higher performance generally requires more fine-grained config tuning. This either increases client-side autotuning overhead or requires maintainers to provide and maintain more pre-tuned configs upstream. The goal, therefore, is to strike the right balance for the…+
<h2><span style="font-weight: 400;">Helion Linear Backend </span></h2>+
<h3><span style="font-weight: 400;">Scope</span></h3>+
<p><span style="font-weight: 400;">This work adds a Helion linear backend to vLLM and focuses on NVIDIA Hopper GPUs using Helion’s Triton backend. FP8 and INT8 are the primary quantization formats for efficient inference on Hopper, so we target quantized GEMM with the following formats:</span></p>+
<ul>+
<li style="font-weight: 400;" aria-level="1"><a href="https://docs.vllm.ai/en/latest/features/quantization/llm_compressor/fp8/"><b>FP8_Dynamic</b></a><span style="font-weight: 400;">: FP8 per-token activation and per-channel weight scaling.</span></li>+
<li style="font-weight: 400;" aria-level="1"><a href="https://docs.vllm.ai/en/latest/features/quantization/llm_compressor/int8_w8a8/"><b>W8A8_INT8</b></a><span style="font-weight: 400;">: INT8 per-token activation and per-channel weight scaling.</span></li>+
<li style="font-weight: 400;" aria-level="1"><a href="https://docs.vllm.ai/en/latest/features/quantization/llm_compressor/fp8/"><b>Block_FP8</b></a><span style="font-weight: 400;">: FP8 with 1×128 activation scaling and 128×128 weight scaling. </span></li>+
</ul>+
<p><a href="https://github.com/pytorch/helion/pull/3377#issuecomment-5345673133"><span style="font-weight: 400;">Initial results</span></a><span style="font-weight: 400;"> with Helion’s CuteDSL backend show competitive GEMM performance on NVIDIA Blackwell GPUs. As the CuteDSL backend matures, this w…+
<h3><span style="font-weight: 400;">Kernel Implementation</span></h3>+
<p><span style="font-weight: 400;">Several GEMM algorithm variants can improve performance for specific input shapes. In particular:</span></p>+
<ul>+
<li style="font-weight: 400;" aria-level="1"><b>Split-K</b><span style="font-weight: 400;">: Partitions the K dimension across multiple thread blocks to increase parallelism when the M and/or N dimensions are too small.</span></li>+
<li style="font-weight: 400;" aria-level="1"><b>Swap-AB</b><span style="font-weight: 400;">: Rewrites </span><span style="font-weight: 400;">A@B</span> <span style="font-weight: 400;">as</span> <span style="font-weight: 400;">(B.T@A.T).T</span><span style="font-weight: 400;"> to improve performance …+
</ul>+
<p><span style="font-weight: 400;">Traditionally, kernel authors need to implement multiple GEMM variants, benchmark the standard GEMM against these specialized variants, and develop heuristics to determine which implementation to dispatch to based on the input shape. For example, vLLM’s current Blo…+
<p><span style="font-weight: 400;">With Helion, a single unified GEMM implementation can cover all three variants—Standard, Split-K, and Swap-AB. Instead of implementing separate kernels and manually designing dispatch heuristics, the algorithmic choices are exposed as tunable parameters. The AOT au…+
<p><span style="font-weight: 400;">For simplicity, we use a basic matrix multiplication kernel below to illustrate the approach. The quantized GEMM kernels used in this work follow the same structure, with additional quantization and scaling logic.</span></p>+
<pre><code class="language-python">def matmul(
+
out: torch.Tensor, # [M, N]
+
a: torch.Tensor, # [M, K]
+
b: torch.Tensor, # [K, N]
+
) -> None:
+
M, K = a.shape
+
N = b.shape[1]
+
hl.specialize(K)
+
hl.specialize(N)
+
+
out_dtype = out.dtype
+
acc_dtype = torch.float32
+
+
split_k = hl.register_tunable(
+
"split_k", PowerOfTwoFragment(1, 256)
+
)
+
k_block_size = helion.next_power_of_2(helion.cdiv(K, split_k))
+
if split_k > 1:
+
out.zero_()
+
+
swap_ab = hl.register_tunable(
+
"swap_ab", BooleanFragment()
+
)
+
+
for tile_m, tile_n, outer_k in hl.tile(
+
[M, N, K],
+
block_size=[None, None, k_block_size]
+
):
+
acc = hl.zeros([tile_m, tile_n], acc_dtype)
+
acc_t = acc.t()
+
for tile_k in hl.tile(outer_k.begin, outer_k.end):
+
if swap_ab:
+
a_blk = hl.load(a, [tile_m.index[None, :], tile_k.index[:, None]])
+
b_blk = hl.load(b, [tile_k.index[None, :], tile_n.index[:, None]])
+
acc_t = hl.dot(
+
b_blk,
+
a_blk,
+
acc=acc_t,
+
out_dtype=acc_dtype,
+
)
+
else:
+
acc = hl.dot(
+
a[tile_m, tile_k],
+
b[tile_k, tile_n],
+
acc=acc,
+
out_dtype=acc_dtype,
+
)
+
+
if swap_ab:
+
out_blk = acc_t.t().to(out_dtype)
+
else:
+
out_blk = acc.to(out_dtype)
+
+
if split_k == 1:
+
out[tile_m, tile_n] = out_blk
+
else:
+
hl.atomic_add(out, [tile_m, tile_n], out_blk)</code></pre>+
<p><span style="font-weight: 400;">This implementation exposes both algorithmic choices as tunable parameters: </span><span style="font-weight: 400;"><code>split_k</code></span><span style="font-weight: 400;"> is constrained to power-of-two values up to 256, while </span><span style="font-weight: 40…+
<h3><span style="font-weight: 400;">Hybrid Dispatch</span></h3>+
<p><img decoding="async" class="aligncenter wp-image-171936 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/2.png" alt="" width="1448" height="284" srcset="https://pytorch.org/wp-content/uploads/2026/10/2.png 1448w, https://pytorch.org/wp-content/uploads/2026/10/2-300x59.png 300w, htt…+
<p style="text-align: center;"><i><span style="font-weight: 400;">Fig. 2: Helion linear backend hybrid dispatch strategy based on runtime num_tokens and CUDA Graph coverage.</span></i></p>+
<p>We adopt a hybrid dispatch strategy based on runtime <code>num_tokens</code> and CUDA Graph coverage. For small shapes up to <code>max_helion_size</code>, the linear backend dispatches to Helion under CUDA Graph replay. For larger shapes beyond <code>max_helion_size</code>, it falls back to the d…+
<p><span style="font-weight: 400;">This hybrid dispatch strategy provides three practical benefits:</span></p>+
<ul>+
<li style="font-weight: 400;" aria-level="1"><b>Eliminates Helion runtime overhead</b><span style="font-weight: 400;">. Helion kernels execute only through CUDA Graph replay, avoiding the additional CPU overhead from kernel dispatch and launch.</span></li>+
<li style="font-weight: 400;" aria-level="1"><b>Reduces kernel tuning overhead</b><span style="font-weight: 400;">. Fine-grained Helion tuning is limited to the small </span><span style="font-weight: 400;"><code>num_tokens</code></span><span style="font-weight: 400;"> range that dominates decoding w…+
<li style="font-weight: 400;" aria-level="1"><b>Reduces config maintenance overhead</b><span style="font-weight: 400;">. The smaller set of tuned configs makes pre-tuned configs more practical to validate, ship, and maintain over time.</span></li>+
</ul>+
<h3><span style="font-weight: 400;">Kernel Autotuning</span></h3>+
<p><span style="font-weight: 400;">We autotune the Helion linear kernels using the utility script available in vLLM:</span></p>+
<pre><code>HELION_AUTOTUNER=LLMSeededLFBOTreeSearch \
+
HELION_BENCHMARK_CUDAGRAPH =1 \
+
python scripts/autotune_helion_kernels.py \
+
--kernels scaled_mm block_scaled_mm \
+
--autotune-effort "full"</code></pre>+
<p><span style="font-weight: 400;">The following sections describe the key tuning strategies and setups used in this work.</span></p>+
<h4><span style="font-weight: 400;">Per-shape Config Tuning</span></h4>+
<p><span style="font-weight: 400;">To compete with highly optimized GEMM libraries such as CUTLASS, DeepGEMM, and FlashInfer, we tune the Helion kernel individually for each input shape within the Helion dispatch range.</span></p>+
<p><span style="font-weight: 400;">As described in the hybrid dispatch strategy, we set </span><span style="font-weight: 400;"><code>max_helion_size = 32</code></span><span style="font-weight: 400;"> for this work. vLLM captures CUDA Graphs for the following num_tokens values:</span></p>+
<pre><code>[1, 2, 4] + range(8, 256, 8) + range(256, max_graph_size + 1, 16)</code></pre>+
<p>With <code>max_helion_size = 32</code>, the Helion kernels are therefore autotuned for:</p>+
<pre><code>num_tokens = [1, 2, 4, 8, 16, 24, 32]</code></pre>+
<p><span style="font-weight: 400;">Each num_tokens value is tuned individually for the corresponding GEMM shapes used by the model.</span></p>+
<h4><span style="font-weight: 400;">Enable CUDA Graph for Autotuner Benchmarking</span></h4>+
<p><span style="font-weight: 400;">Due to the additional CPU overhead from kernel dispatch and launch, Helion kernels are used only under CUDA Graph execution. Benchmarking with CUDA Graph enabled therefore allows the autotuner to evaluate configs under conditions that more closely match actual infe…+
<p><span style="font-weight: 400;">Helion exposes this feature through the <code>HELION_BENCHMARK_CUDAGRAPH</code></span><span style="font-weight: 400;">environment variable and is turned on for this work.</span></p>+
<h4><span style="font-weight: 400;">Use LLM-Guided Search</span></h4>+
<p><span style="font-weight: 400;">The kernel configs in this work are generated using the </span><a href="https://pytorch.org/blog/from-minutes-to-seconds-llm-guided-autotuning-for-helion-kernels/"><span style="font-weight: 400;">LLMSeededLFBOTreeSearch</span></a><span style="font-weight: 400;"> au…+
<p><span style="font-weight: 400;">Starting the search from higher-quality candidates helps the autotuner discover better-performing configs. It may also reduce overall autotuning time by directing the search toward more promising regions of the search space.</span></p>+
<h2><span style="font-weight: 400;">Performance Evaluation</span></h2>+
<p><span style="font-weight: 400;">We evaluate the Helion linear backend at both the kernel and end-to-end serving levels across a range of dense models and quantization formats. To understand how performance scales with model size, we benchmark Qwen3-1.7B, Qwen3-4B, Qwen3-8B, Qwen3-14B, and Qwen3-3…+
<p><span style="font-weight: 400;">For each model, we evaluate three commonly used 8-bit quantization formats on NVIDIA Hopper GPUs. For example, for Qwen3.8-27B, we benchmark: </span></p>+
<ul>+
<li style="font-weight: 400;" aria-level="1"><a href="https://huggingface.co/AzatAI/Qwen3.8-27B-FP8-dynamic"><span style="font-weight: 400;">AzatAI/Qwen3.8-27B-FP8-dynamic</span></a><span style="font-weight: 400;"> (FP8_Dynamic)</span></li>+
<li style="font-weight: 400;" aria-level="1"><a href="https://huggingface.co/Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8"><span style="font-weight: 400;">Freaksterz/Qwen3.8-27B-SmoothQuant-W8A8-INT8</span></a><span style="font-weight: 400;"> (W8A8_INT8)</span></li>+
<li style="font-weight: 400;" aria-level="1"><a href="https://huggingface.co/Qwen/Qwen3.8-27B-FP8"><span style="font-weight: 400;">Qwen/Qwen3.8-27B-FP8</span></a><span style="font-weight: 400;"> (Block_FP8) </span></li>+
</ul>+
<p><span style="font-weight: 400;">All benchmarks are performed on an NVIDIA H100 80GB HBM3 GPU.</span></p>+
<h3><span style="font-weight: 400;">Kernel-Level Evaluation</span></h3>+
<p><span style="font-weight: 400;">We first evaluate the Helion quantized GEMM kernels in isolation to understand the performance benefit of fine-grained tuning independent of the rest of the inference stack. Each Helion kernel is compared against the corresponding kernel used by the default vLLM li…+
<ul>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">FP8_Dynamic: Helion vs. CUTLASS</span></li>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">W8A8_INT8: Helion vs. CUTLASS</span></li>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Block_FP8: Helion vs. FlashInfer/DeepGEMM</span></li>+
</ul>+
<p><img decoding="async" class="aligncenter wp-image-171943 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/3.png" alt="" width="2048" height="1035" srcset="https://pytorch.org/wp-content/uploads/2026/10/3.png 2048w, https://pytorch.org/wp-content/uploads/2026/10/3-300x152.png 300w, h…+
<p style="text-align: center;"><i><span style="font-weight: 400;">Fig. 3: Helion quantized GEMM kernel speedup distribution over the default vLLM kernel libraries across all input shapes used by the end-to-end benchmarked models. Diamonds indicate the geometric mean speedup.</span></i></p>+
<p><span style="font-weight: 400;">Helion achieves geometric mean speedups of </span><b>1.110×</b><span style="font-weight: 400;"> for FP8_Dynamic over CUTLASS, </span><b>1.178×</b><span style="font-weight: 400;"> for W8A8_INT8 over CUTLASS, </span><b>1.149×</b><span style="font-weight: 400;"> for B…+
<h3><span style="font-weight: 400;">End-to-End Evaluation</span></h3>+
<p><span style="font-weight: 400;">We next evaluate whether the kernel-level improvements translate into end-to-end serving performance.</span></p>+
<h4><span style="font-weight: 400;">Server Setup</span></h4>+
<p><span style="font-weight: 400;">We use the following command to start the vLLM server:</span></p>+
<pre><code>vllm serve \
+
--model "$MODEL" \
+
--max-num-seqs 32 \
+
--tensor-parallel-size 1 \
+
--no-enable-prefix-caching \
+
--linear-backend helion</code></pre>+
<p>The relevant options are:</p>+
<ul>+
<li><code>--max-num-seqs 32</code>: The Helion linear backend currently uses a hybrid dispatch threshold of 32 <code>num_tokens</code>. We therefore focus the evaluation on batch sizes up to 32, where Helion kernels are active and can directly affect end-to-end performance.</li>+
<li><code>--no-enable-prefix-caching</code>: Prefix caching is intentionally disabled to avoid its impact on benchmark results.</li>+
<li><code>--linear-backend helion</code>: Enables the Helion linear backend. Omitting this flag uses the default linear backend.</li>+
</ul>+
<h3><span style="font-weight: 400;">Benchmark Setup</span></h3>+
<p><span style="font-weight: 400;">We run end-to-end serving benchmarks using the ShareGPT dataset with:</span></p>+
<pre><code>vllm bench serve \
+
--backend vllm \
+
--model "${MODEL}" \
+
--endpoint /v1/completions \
+
--dataset-name sharegpt \
+
--dataset-path "${DATASET}" \
+
--max-concurrency "${BATCH_SIZE}" \
+
--num-warmups "${NUM_WARMUPS}" \
+
--num-prompts "${PROMPTS}" \
+
--ignore-eos</code></pre>+
<p><span style="font-weight: 400;">Each workload is benchmarked with both the default and Helion linear backends. The default vLLM </span><a href="https://docs.vllm.ai/en/latest/api/vllm/model_executor/kernels/linear/scaled_mm/"><span style="font-weight: 400;">linear backends</span></a><span style="…+
<ul>+
<li style="font-weight: 400;" aria-level="1"><b>FP8_Dynamic</b><span style="font-weight: 400;">: CutlassFP8ScaledMMLinearKernel</span></li>+
<li style="font-weight: 400;" aria-level="1"><b>W8A8_INT8</b><span style="font-weight: 400;">: CutlassInt8ScaledMMLinearKernel</span></li>+
<li style="font-weight: 400;" aria-level="1"><b>Block_FP8</b><span style="font-weight: 400;">: FlashInferFp8DeepGEMMDynamicBlockScaledKernel</span></li>+
</ul>+
<h4><span style="font-weight: 400;">End-to-End Benchmark Results</span></h4>+
<p><span style="font-weight: 400;">The following Figure shows the end-to-end throughput speedup of the Helion linear backend over the corresponding default backend across different model sizes, batch sizes, and quantization formats.</span></p>+
<p><img decoding="async" class="aligncenter wp-image-171947 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/4.png" alt="" width="2048" height="625" srcset="https://pytorch.org/wp-content/uploads/2026/10/4.png 2048w, https://pytorch.org/wp-content/uploads/2026/10/4-300x92.png 300w, htt…+
<p style="text-align: center;"><i><span style="font-weight: 400;">Fig. 4: Helion linear backend end-to-end throughput speedup.</span></i></p>+
<p><span style="font-weight: 400;">Overall, the Helion linear backend delivers consistent end-to-end performance gains across the evaluated models and quantization formats, with more than 10% throughput improvement for some workloads.</span></p>+
<h2><span style="font-weight: 400;">Toward a Practical Adoption Model</span></h2>+
<p><span style="font-weight: 400;">The Helion linear backend presented in this work is currently available in our </span><a href="https://github.com/redhat-et/vllm-helion"><span style="font-weight: 400;">vLLM fork</span></a><span style="font-weight: 400;"> and ready for production use. The fork also…+
<p><span style="font-weight: 400;">For latency-critical kernels such as the GEMM kernels studied in this work, we are exploring a model in which the Helion kernels and integration framework are maintained upstream with a default config for functional testing and CI, while workload-specific autotunin…+
<p><span style="font-weight: 400;">Fine-grained tuning, however, is not the only practical adoption strategy for Helion kernels. For smaller auxiliary kernels, such as quantization, activation, and normalization kernels, it can be preferable to trade some peak performance for a much smaller config s…+
<p><span style="font-weight: 400;">We encourage interested users to try the current implementation and share their experience in the </span><a href="https://github.com/vllm-project/vllm/issues/46526"><span style="font-weight: 400;">vLLM Helion linear backend RFC</span></a><span style="font-weight: …+
<h2><span style="font-weight: 400;">Future work</span></h2>+
<p><span style="font-weight: 400;">Future work will focus on expanding Helion kernel coverage across inference workloads and hardware platforms.</span></p>+
<p><b>Hardware coverage</b><span style="font-weight: 400;">. The Helion team is continuing to expand and optimize backend support across hardware platforms, including the CuteDSL backend for NVIDIA Blackwell GPUs and ongoing performance work for AMD GPUs and </span><a href="https://pytorch.org/blog/…+
<p><b>Model coverage</b><span style="font-weight: 400;">. This work primarily targets dense models, where the linear backend accounts for a significant portion of inference latency. For MoE models, the MoE backend becomes the more important optimization target. We are working on a Helion MoE backend…+
<h2><span style="font-weight: 400;">Conclusion</span></h2>+
<p><span style="font-weight: 400;">Helion’s high-level abstraction makes it possible to express and maintain a single kernel implementation while optimizing it across different workloads and hardware targets. As demonstrated by the quantized GEMM kernels in this work, even algorithmic variants such …+
<p><span style="font-weight: 400;">These performance gains, however, come with tradeoffs. Fine-grained tuning for latency-critical kernels improves performance but increases AOT tuning effort, while shipping pre-tuned configs improves out-of-the-box usability at the cost of ongoing maintenance. This…+
<p><span style="font-weight: 400;">Today, we ship the pre-tuned configs with our </span><a href="https://github.com/redhat-et/vllm-helion"><span style="font-weight: 400;">vLLM fork</span></a><span style="font-weight: 400;">. For upstream adoption, we are exploring a different model: maintain the Hel…+
<h2><span style="font-weight: 400;">Acknowledgments</span></h2>+
<p><span style="font-weight: 400;">This work was supported by many contributors across the OCTO and vLLM teams at Red Hat, as well as the Helion team at Meta. In particular, we would like to thank our colleagues: Richard Zou and Jongsok Choi for their feedback and support throughout this work. </spa…+
]]></content:encoded>+
+
+
+
</item>+
<item>+
<title>New Pathway to PyTorch Certified Associate (PTCA) Certification</title>+
<link>https://pytorch.org/blog/new-pathway-to-pytorch-certified-associate-ptca-certification/</link>+
+
<dc:creator><![CDATA[PyTorch Foundation]]></dc:creator>+
<pubDate>Fri, 02 Oct 2026 19:33:47 +0000</pubDate>+
<category><![CDATA[Announcements]]></category>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=171978</guid>+
+
<description><![CDATA[We’re excited to bring you the new PyTorch Certified Associate (PTCA) Certification Pathway, combining focused learning modules and the PyTorch Certified Associate (PTCA) certification exam together in one structured experience...]]></description>+
<content:encoded><![CDATA[<p><span style="font-size: 18.72px;"><img decoding="async" class="alignnone size-large wp-image-171984" src="https://pytorch.org/wp-content/uploads/2026/10/PyTorch-Certification-Pathway-1024x536.png" alt="PyTorch Certification Pathway" width="1024" height="536" sr…+
<p>We’re excited to bring you the new <a href="https://training.linuxfoundation.org/certification/ptca-certification-pathway/">PyTorch Certified Associate (PTCA) Certification Pathway</a>, combining focused learning modules and the PyTorch Certified Associate (PTCA) certification exam together…+
<p>The pathway includes four self-paced learning modules plus the PTCA certification exam:</p>+
<ul>+
<li>PyTorch Fundamentals</li>+
<li>Data Handling in PyTorch</li>+
<li>PyTorch Model Development</li>+
<li>Optimizing PyTorch</li>+
</ul>+
<p>Together, the modules cover core skills across the same key areas measured by the PTCA certification, from working with tensors and devices to preparing data, building neural networks, and improving model performance.</p>+
<h4><strong>Learn, Build, Validate, Certify</strong></h4>+
<p>Throughout the pathway, you’ll build skills to:</p>+
<ul>+
<li><strong>Work with tensors and devices</strong> across the model lifecycle, from training through prediction</li>+
<li><strong>Prepare and manage data</strong> using Datasets, DataLoaders, and transforms</li>+
<li><strong>Build neural networks</strong> using PyTorch layers, activations, loss functions, and optimizers</li>+
<li><strong>Measure and optimize performance</strong> using tools and techniques including torch.compile, Automatic Mixed Precision, the PyTorch Profiler, and distributed training concepts</li>+
</ul>+
<p>The full pathway includes 15–17 hours of self-paced learning and hands-on labs designed to help you build a strong foundation. Certification readiness takes more than completing coursework, so additional hands-on practice and real-world experience with PyTorch are recommended as you prepare for t…+
<h4><strong>Put Your PyTorch Skills on a Clear Path</strong></h4>+
<p>PTCA is designed for early-stage practitioners with Python and machine learning experience who are beginning to work with PyTorch. The certification validates your understanding of the core concepts and workflows used by engineers and researchers to design, train, optimize, and work with machine …+
<p>Build your PyTorch skills and get certified today.</p>+
]]></content:encoded>+
+
+
+
</item>+
<item> <title>Optimizing Jagged Flash Attention with TLX: The Road Toward SOTA FA4 on Blackwell</title> <link>https://pytorch.org/blog/optimizing-jagged-flash-attention-with-tlx-the-road-toward-sota-fa4-on-blackwell/</link> @
@@ -44,7 +301,7 @@<p><span style="font-weight: 400;">Code available at: </span><a href="https://github.com/facebookresearch/ads_model_kernel_library/tree/main/tlx_jfa"><span style="font-weight: 400;">https://github.com/facebookresearch/ads_model_kernel_library/tree/main/tlx_jfa</span></a></p><h2><span style="font-weight: 400;">Introduction</span></h2><p><span style="font-weight: 400;">Meta’s ads models, including the Generative Ads Model (GEM) [2] and the Kunlun architecture [5], run attention over jagged (variable-length or ragged) user sequences. As described in our </span><a href="https://engineering.fb.com/2026/08/03/ml-applications/tr…-
<p><img fetchpriority="high" decoding="async" class="aligncenter wp-image-171512 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/1.png" alt="" width="1376" height="768" srcset="https://pytorch.org/wp-content/uploads/2026/09/1.png 1376w, https://pytorch.org/wp-content/uploads/2026/09/1…+
<p><img decoding="async" class="aligncenter wp-image-171512 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/1.png" alt="" width="1376" height="768" srcset="https://pytorch.org/wp-content/uploads/2026/09/1.png 1376w, https://pytorch.org/wp-content/uploads/2026/09/1-300x167.png 300w, ht…<p><span style="font-weight: 400;">Hitting peak throughput on Blackwell means keeping the tensor cores continuously fed. Plain Triton leaves most of these decisions to the compiler; TLX [1] exposes them as first-class primitives (explicit SMEM/TMEM allocation, async_task warp specialization, barrier…<h2><span style="font-weight: 400;">1. The Challenge: Performant and Extensible Attention on Blackwell</span></h2><p><span style="font-weight: 400;">Optimizing attention on Blackwell means solving two problems at once. The first is performance: attention is the single slowest kernel in GEM, and it reaches peak throughput only when the global-memory loads, softmax, and matmuls are tightly overlapped. The second …@
@@ -924,349 +1181,8 @@ echo "Resolved pytorch/pytorch@nightly -> source ${SOURCE_SHA}" -
</item>-
<item>-
<title>PyTorch Day Japan 2026 Comes to Tokyo on December 10</title>-
<link>https://pytorch.org/blog/pytorch-day-japan-2026-comes-to-tokyo/</link>-
-
<dc:creator><![CDATA[PyTorch Foundation]]></dc:creator>-
<pubDate>Fri, 18 Sep 2026 00:25:32 +0000</pubDate>-
<category><![CDATA[Announcements]]></category>-
<category><![CDATA[Blog]]></category>-
<guid isPermaLink="false">https://pytorch.org/?p=168823</guid>-
-
<description><![CDATA[PyTorch Day Japan 2026 will bring the open source AI community together in Tokyo on December 10 for a full day of technical talks and interactive discussions designed to foster...]]></description>-
<content:encoded><![CDATA[<p><a href="https://events.linuxfoundation.org/pytorch-day-japan/">PyTorch Day Japan 2026</a> will bring the open source AI community together in Tokyo on December 10 for a full day of technical talks and interactive discussions designed to foster knowledge exchan…-
<p>Hosted by PyTorch Foundation, Hugging Face, IBM, and Mitsubishi Electric, the event will bring together PyTorch enthusiasts, machine learning engineers, AI researchers, and industry professionals working across open source AI.</p>-
<p>The program will explore PyTorch and PyTorch Foundation hosted projects, including vLLM, DeepSpeed, Ray, Helion, and Safetensors, alongside broader topics in AI and machine learning such as training, inference, responsible AI, physical and edge AI, and open model development.</p>-
<p><a href="https://events.linuxfoundation.org/pytorch-day-japan/program/cfp/">The call for proposals is now open</a>, and <a href="https://events.linuxfoundation.org/pytorch-day-japan/register/">registration is available with discounted pricing</a> through November 11.</p>-
<h2>Submit a Session Proposal</h2>-
<p>PyTorch Day Japan 2026 is accepting proposals for session presentations and lightning talks.</p>-
<p>Suggested topics include:</p>-
<p><strong>Sovereign AI and local open models:</strong> Strategies and architectures for building and running open-weights AI on local infrastructure, including open model adaptation, post-training, local inference, and privacy-first deployments using PyTorch.</p>-
<p><strong>Physical AI and edge AI:</strong> Experiences bringing PyTorch models to physical hardware, robotics, and edge systems, including real-time on-device inference, vision-language-action models, hardware acceleration, and optimization for resource-constrained environments.</p>-
<p><strong>PyTorch ecosystem:</strong> Developments across the PyTorch library and developer toolchain, including domain libraries such as TorchVision, TorchAudio, TorchRL, and PyTorch Distributed, as well as developer tooling, PyTorch 2.x and <code>torch.compile</code>, production pipelines, and co…-
<p>The CFP closes <strong>Sunday, September 27 at 11:59 PM JST</strong>.</p>-
<p>Key dates:</p>-
<ul>-
<li>CFP deadline: Sunday, September 27 at 11:59 PM JST</li>-
<li>CFP notifications: Tuesday, October 13</li>-
<li>Schedule announcement: Wednesday, October 14</li>-
<li>Presentation slides due: Wednesday, December 9</li>-
<li>PyTorch Day Japan: Thursday, December 10</li>-
</ul>-
<p><a href="https://events.linuxfoundation.org/pytorch-day-japan/program/cfp/">Submit a proposal</a></p>-
<h2>Register for PyTorch Day Japan</h2>-
<p>Registration is also open for PyTorch Day Japan 2026.</p>-
<p>Registration is <strong>¥5,000 through November 11 at 11:59 PM JST</strong>, representing ¥3,000 in savings. Discounted academic pricing is also available for students and faculty members.</p>-
<p><a href="https://events.linuxfoundation.org/pytorch-day-japan/register/">Register for PyTorch Day Japan</a></p>-
<p>PyTorch Day Japan 2026 takes place <strong>Thursday, December 10 in Tokyo, Japan</strong>.</p>-
<p>For full event details, visit the <a href="https://events.linuxfoundation.org/pytorch-day-japan/">PyTorch Day Japan 2026 website</a>.</p>-
]]></content:encoded>-
-
-
-
</item>-
<item>-
<title>Open Research, Tooling & Optimization at PyTorch Conference North America 2026</title>-
<link>https://pytorch.org/blog/open-research-tooling-optimization-at-pytorch-conference-north-america-2026/</link>-
-
<dc:creator><![CDATA[PyTorch Foundation]]></dc:creator>-
<pubDate>Wed, 16 Sep 2026 19:00:47 +0000</pubDate>-
<category><![CDATA[Blog]]></category>-
<guid isPermaLink="false">https://pytorch.org/?p=167741</guid>-
-
<description><![CDATA[TL;DR Taking place October 20 to 21 in San Jose, California, PyTorch Conference North America 2026 highlights open research, tooling, and performance optimization across compiler architecture, cross-hardware kernel domain-specific languages,...]]></description>-
<content:encoded><![CDATA[<h2>TL;DR</h2>-
<p><span style="font-weight: 400;">Taking place October 20 to 21 in San Jose, California, PyTorch Conference North America 2026 highlights open research, tooling, and performance optimization across compiler architecture, cross-hardware kernel domain-specific languages, exascale distributed training…-
<h2><span style="font-weight: 400;">Introduction</span></h2>-
<p><span style="font-weight: 400;">Every year, the beating heart of PyTorch Conference is the work that happens <em>underneath</em> the flashy headlines- the compilers, kernels, autotuners, distributed runtimes, profilers, and open-source libraries that make the rest of the ecosystem possible. At Py…-
<p><span style="font-weight: 400;">This is where you’ll find the engineers rewriting torch.compile’s internals for speed, the teams building DSLs (Helion, CuteDSL, FlyDSL) that let a single kernel target NVIDIA, AMD, Intel, and TPU silicon, the researchers pushing quantization down to FP…-
<p><span style="font-weight: 400;">In this blog we have picked out the sessions that fall into this theme – open source tooling, systems research, and hard-won performance optimization.</span></p>-
<p><a href="https://hubs.ly/Q04tDx8f0">View the full conference schedule</a></p>-
<p><a href="https://hubs.ly/Q04tDx8f0">Register for PyTorch Conference North America 2026</a></p>-
<h2><span style="font-weight: 400;">Compilers, Graph Capture & Dynamic Shapes</span></h2>-
<p><span style="font-weight: 400;">Everything to do with torch.compile, Dynamo, and the machinery that turns eager PyTorch code into a graph format that can be optimized at runtime, or used to generate artifacts for later execution</span></p>-
<p><strong>Beyond Size and Stride: Unleash Performance with Device-Aware Tensor Layouts</strong><br />-
Olivier Tardieu; Matthew Arnold (IBM)<br />-
Tue 20 Oct, 16:20–16:30, LL20AB</p>-
<p>A lightweight tensor layout extension that lets torch.compile (Inductor) automatically adapt tiling and NUMA-aware placement to the target device, without changing existing PyTorch model code.</p>-
<p><b>Nested Graph Breaks: Reducing the Cost of Graph Breaks in torch.compile</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">William Wen (Meta) </span></p>-
<p><span style="font-weight: 400;">Tue 20 Oct, 16:35–16:45, LL20AB </span></p>-
<p><span style="font-weight: 400;">New Dynamo support reduces the cost of a graph break nested O(N) layers deep from O(N) duplicate breaks and O(N²) frame traces down to O(1) duplicate breaks and O(N) frame traces – yielding larger captured graphs, fewer breaks, and faster, more debuggable com…-
<p><b>Static Tensor Shape Checking for PyTorch with Pyrefly</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Steven Troxler; Avik Chaudhuri (Meta) </span></p>-
<p><span style="font-weight: 400;">Tue 20 Oct, 17:30–17:55, LL20AB </span></p>-
<p><span style="font-weight: 400;">Shape mismatches cause roughly 45% of deep learning program failures and drive slow torch.compile recompilations. Pyrefly brings static, near-instant tensor shape checking to the type checker itself, evaluated across 28 real LLM, vision, recommender, and RL models.…-
<p><b>Parametrized Dynamic Shape CUDA Graphs</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Elias Ellison (Meta); Daniel Galvez (NVIDIA) </span></p>-
<p><span style="font-weight: 400;">Wed 21 Oct, 11:45–12:10, LL21ABC </span></p>-
<p><span style="font-weight: 400;">Building on parametrized CUDA Graphs and torch.compile’s symbolic tracing and guard infrastructure, this work captures and re-parametrizes a single CUDA Graph across dynamic shapes, eliminating the need for whole-model rewrites, padding, or re-recording acros…-
<p><b>From Backed to Unbacked: Sound, Predictable, and Controllable Dynamic Shapes in PyTorch</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Laith Sakka (Meta) </span></p>-
<p><span style="font-weight: 400;">Wed 21 Oct, 16:20–16:45, LL21ABC </span></p>-
<p><span style="font-weight: 400;">This session explains why explicit graph-capture workflows (vLLM, export, pre-compilation) need unbacked dynamic shapes, which disallow implicit guards and force the compiler to prove general validity. </span></p>-
<p><b>Speeding Up torch.compile: A New FakeTensor</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Angel Li (Meta) </span></p>-
<p><span style="font-weight: 400;">Wed 21 Oct, 16:55–17:05, LL21ABC </span></p>-
<p><span style="font-weight: 400;">FakeTensor propagation eats roughly 20% of Dynamo’s tracing time. This talk introduces a new C++ FakeTensor implementation that delivers a 30x speedup on operations like </span><span style="font-weight: 400;">aten.mm</span><span style="font-weight: 400;">, me…-
<p><b>Lightweight FX Tracing in PyTorch</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Richard Zou; Yidi Wu (Meta) </span></p>-
<p><span style="font-weight: 400;">Wed 21 Oct, 17:10–17:20, LL21ABC </span></p>-
<p><span style="font-weight: 400;">A new FX tracer based on </span><span style="font-weight: 400;">make_fx</span><span style="font-weight: 400;">, targeting functionally pure PyTorch code instead of trying to capture full Python semantics like Dynamo. This means trading flexibility for a much simple…-
<p><b>Unlocking the Full Potential of TorchDynamo: Accelerating, Comparing, and Debugging ML Systems</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Yi Pan (UC Berkeley); Megan Frisella (University of Washington); Stephanie Wang (Paul Allen School, University of Washington) </span></p>-
<p><span style="font-weight: 400;">Wed 21 Oct, 17:30–17:55, LL21ABC </span></p>-
<p><span style="font-weight: 400;">Research using TorchDynamo’s graph-interception capability well beyond torch.compile itself. DynaFlow and Piper for transparent parallelism acceleration, Magneton for automatically comparing equivalent operations across competing systems, and numerical debugg…-
<h2><span style="font-weight: 400;">Kernel Engineering & Domain-Specific Languages</span></h2>-
<p><span style="font-weight: 400;">The DSLs, autotuners, and hand- and agent-written kernels that squeeze more performance out of every GPU cycle, spanning Triton, Helion, CUTLASS, FlyDSL, and the agentic pipelines now writing and optimizing kernels themselves.</span></p>-
<p><b>Extending TorchInductor with FlyDSL: A New MLIR-Native Backend for High-Performance GEMMs</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Liz Li (AMD) </span></p>-
<p><span style="font-weight: 400;">Tue 20 Oct, 11:10–11:35, 210BF </span></p>-
<p><span style="font-weight: 400;">AMD’s FlyDSL, a Python-native, MLIR-based GPU kernel DSL, gets integrated into TorchInductor’s GEMM compilation pipeline. The talk covers how FlyDSL plugs into autotuning, coexists with existing backends, and delivers measured speedups over Triton on AM…-
<p><b>Helion: CuteDSL and TPU Backends for Heterogeneous Hardware, and Why It Suits Agents</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Oguz Ulgen, Dunfan Lu, Jason Ansel (Meta) </span></p>-
<p><span style="font-weight: 400;">Tue 20 Oct, 11:45–12:10, 210BF </span></p>-
<p><span style="font-weight: 400;">Helion, PyTorch’s high-level kernel-authoring DSL, gains two new compiler backends – CuteDSL for recent NVIDIA GPUs and Pallas for TPUs – letting one kernel source target very different hardware. The second half explores why Helion’s abstrac…-
<p><b>Practical GPU Programming with Triton for PyTorch Developers</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Suman Debnath, JanakiRam Goteti (Crusoe AI) </span></p>-
<p><span style="font-weight: 400;">Tue 20 Oct, 11:45–12:10, LL20CD</span></p>-
<p><span style="font-weight: 400;">A friendly, from-scratch introduction to writing GPU kernels in Triton – no CUDA or C++ required. Building from a simple vector-add up to a small matrix multiplication, with an emphasis on reading and understanding what PyTorch already generates for you.</sp…-
<p><b>High-Velocity GPU Kernel Authoring with CUTLASS Python</b><span style="font-weight: 400;"> </span><br />-
<span style="font-weight: 400;">Michael Goldfarb, Guray Ozen (NVIDIA) </span></p>-
<p><span style="font-weight: 400;">Tue 20 Oct, 12:20–12:45, 210BF </span></p>-
<p><span style="font-weight: 400;">New Python-first capabilities in CUTLASS’s CuTe DSL – high-level building blocks, low-level hardware-instruction primitives, and a zero-cost Resource and Task Scheduler – that reduce boilerplate for kernel authors while keeping the DSL’s zer…-
<p><b>Smarter Autotuning for Kernels: From Bayesian Optimization to LLM-Guided Search in Helion DSL</b><span style="font-weight: 400;"> </span><br />Diff display stops at 400 lines. The line counts above are from the whole diff. 62 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.