Change
e06633b
e06633b5c25bf7a1f8e607356329861f5cd108bb · commit on GitHub
pytorch-blog-feed: changed (284325 bytes, HTTP 200)
raw/pytorch-blog-feed/response.xml modified
- Source
- pytorch-blog-feed
- Lines added
- +179
- Lines removed
- -367
- Stored bytes at this commit
- 284,325
- Timestamp
- origin
- Raw artifact at this commit
- raw/pytorch-blog-feed/response.xml
Recorded headers
| observed_at | 2026-09-08T04:36:28.784Z |
|---|---|
| origin_date | 2026-09-08T04:34:40.000Z |
| status | 200 |
| final URL | https://pytorch.org/blog/feed/ |
| etag | "96fc700a42a76343f3bd97a9ee10b52e" |
| last-modified | Tue, 08 Sep 2026 01:14:31 GMT |
| date | Tue, 08 Sep 2026 04:36:28 GMT |
| age | 108 |
| cache-control | public, max-age=60, s-maxage=43200, stale-while-revalidate=86400, stale-if-error=604800 |
| cf-cache-status | null |
| content-encoding | null |
| content-length | 284325 |
@
@@ -12,7 +12,7 @@ <atom:link href="https://pytorch.org/blog/feed/" rel="self" type="application/rss+xml" /> <link>https://pytorch.org</link> <description></description>-
<lastBuildDate>Fri, 04 Sep 2026 13:59:00 +0000</lastBuildDate>+
<lastBuildDate>Tue, 08 Sep 2026 01:14:31 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod>@
@@ -28,6 +28,182 @@ <height>32</height></image> <item>+
<title>Alibaba Cloud, Ant Group, Cambricon and Huawei Come Together in Shanghai to Advance the Open Source AI Stack at PyTorch Conference China</title>+
<link>https://pytorch.org/blog/alibaba-cloud-ant-group-cambricon-and-huawei-come-together-in-shanghai-to-advance-the-open-source-ai-stack-at-pytorch-conference-china/</link>+
+
<dc:creator><![CDATA[PyTorch Foundation]]></dc:creator>+
<pubDate>Tue, 08 Sep 2026 01:00:02 +0000</pubDate>+
<category><![CDATA[Announcements]]></category>+
<category><![CDATA[Blog]]></category>+
<category><![CDATA[Press Release]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=162723</guid>+
+
<description><![CDATA[China’s leading AI technology companies Alibaba Cloud, Cambricon and Ant Group join the PyTorch Foundation as members Summary Alibaba Cloud and Cambricon joined the PyTorch Foundation as Platinum members, and...]]></description>+
<content:encoded><![CDATA[<p><i><span style="font-weight: 400;">China’s leading AI technology companies Alibaba Cloud, Cambricon and Ant Group join the PyTorch Foundation as members</span></i></p>+
<p><b>Summary</b></p>+
<ul>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Alibaba Cloud and Cambricon joined the PyTorch Foundation as Platinum members, and Ant Group joined as a Gold member.</span></li>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">The foundation announced the new memberships Sept. 8 at the PyTorch Conference China 2026 in Shanghai. Together these organizations are mapping the future of the PyTorch ecosystem.</span></li>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Representatives from the three new member companies and Huawei are delivering keynotes on advancing the open AI stack, covering hardware, models, and infrastructure.</span></li>+
</ul>+
<p><b>SHANGHAI – KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China 2026 – September 8, 2026 – </b><a href="https://hubs.la/Q03PC0k70"><span style="font-weight: 400;">The PyTorch Foundation</span></a><span style="font-weight: 400;">, a community-driven hub for open source AI und…+
<p><span style="font-weight: 400;">Together these organizations are mapping the future of the PyTorch ecosystem. During the conference keynotes, </span><a href="https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1317079"><span style="fon…+
<p><span style="font-weight: 400;">“Open source has become the default way the world builds AI, and the PyTorch Foundation is a leading hub for this innovation,” said Mark Collier, Executive Director of the PyTorch Foundation. “Welcoming Alibaba Cloud, Cambricon, and Ant Group as members alongside H…+
<p><span style="font-weight: 400;">The keynotes coincide with today’s announcement that the PyTorch Foundation has welcomed Alibaba Cloud and Cambricon as Platinum members and Ant Group as a Gold member — joining Huawei, a long-standing supporter and contributor to the PyTorch ecosystem. More …+
<p><span style="font-weight: 400;">Alibaba Cloud is a global leader in full-stack AI services, offering state-of-the-art intelligent capabilities and a worldwide AI cloud computing network. Qwen, the family of large language and multimodal AI models developed by Alibaba Cloud, has become one of the …+
<p><span style="font-weight: 400;">Cambricon is a China-based designer of AI chips and accelerators supporting large-scale model training and inference. Cambricon hosts many open source toolkits, drivers, and machine learning pipelines for its Machine Learning Unit (MLU) hardware ecosystem, supporti…+
<p><span style="font-weight: 400;">Ant Group is a global digital technology company driving innovation in AI and digital services, actively contributing to open-source ecosystems across cloud-native and AI frameworks. As a Gold member, Ant Group joins the Foundation’s efforts to advance production-g…+
<p><b>Member Keynotes at PyTorch Conference China include:</b></p>+
<ul>+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">“</span><a href="https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1317079"><span style="font-weight: 400;">Serving Qwen at Scale: Multi-Cluster AI Infrastruct…+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">“</span><a href="https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1308539"><span style="font-weight: 400;">What AI Agents Need from Open Infrastructure</span>…+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">“</span><a href="https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1309274"><span style="font-weight: 400;">Building an Agent Runtime with Open Infrastructure<…+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">“</span><a href="https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1313965"><span style="font-weight: 400;">Towards Device-agnostic PyTorch: Building Unified I…+
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">“</span><a href="https://www.lfopensource.cn/kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/program/schedule/?id=1317073"><span style="font-weight: 400;">Ascend & PyTorch: Pioneer New AI Open Ecosystem…+
</ul>+
<p><span style="font-weight: 400;">Developers and contributors interested in participating in the PyTorch Foundation project ecosystem are encouraged to join the community onsite at </span><a href="https://hubs.la/Q049GBK60"><span style="font-weight: 400;">KubeCon + CloudNativeCon + OpenInfra Summit…+
<p><b>Supporting Quotes</b></p>+
<p><span style="font-weight: 400;">“We believe the future of AI is built on open, production-proven infrastructure — and PyTorch sits at the heart of that future. Joining the PyTorch Foundation is a natural step given our years of running PyTorch at scale across heterogeneous hardware on Alibaba Clo…+
<p><b>– Dr. Feifei Li, Chief Technology Officer, Alibaba Cloud</b></p>+
<p><span style="font-weight: 400;">“AI proves its value through real tasks and real users. We want more people to be able to build with AI and afford to use it. Joining the PyTorch Foundation is another step in Ant Group’s long-term participation in global open-source collaboration. We l…+
<p><b>– Zhengyu He, Chief Technology Officer, Ant Group</b></p>+
<p><span style="font-weight: 400;">“We believe that realizing the full potential of AI computing requires both software-hardware co-optimization and a thriving open ecosystem. Joining the PyTorch Foundation gives us an opportunity to contribute our experience from large-scale AI deployments. We want…+
<p><b>– Elton Gong, Vice President of Software Engineering, Cambricon</b></p>+
<p><span style="font-weight: 400;">“When Huawei joined the PyTorch Foundation in 2023, we came with a goal we still hold today: to help make diverse computing power ubiquitous, and to do that work upstream first. That is what led us to help start the Accelerator Integration Working Group, and …+
<p><b>— Fred Li, Head of Computing Open Source Development Team, Huawei</b></p>+
<p><span style="font-weight: 400;">++++</span></p>+
<h3><b>About the PyTorch Foundation</b></h3>+
<p><span style="font-weight: 400;">The PyTorch Foundation is the vendor-neutral home for the open source intelligence layer developers use for training, optimizing, serving, orchestrating, and running models on any chip in any cloud for any agent. As a community-driven hub hosted by the Linux Founda…+
<h3><b>About the Linux Foundation</b></h3>+
<p><span style="font-weight: 400;">The Linux Foundation is the world’s leading home for collaboration on open source software, hardware, standards, and data. Linux Foundation projects, including Linux, Kubernetes, Model Context Protocol (MCP), OpenChain, OpenSearch, OpenSSF, OpenStack, PyTorch, Ray,…+
<p><i><span style="font-weight: 400;">The Linux Foundation has registered trademarks and uses trademarks. For a list of trademarks of the Linux Foundation, please see its trademark usage page: www.linuxfoundation.org/trademark-usage. Linux is a registered trademark of Linus Torvalds.</span></i></p>+
<p> </p>+
<p><b>Media Contact</b></p>+
<p><span style="font-weight: 400;">Grace Lucier</span><span style="font-weight: 400;"><br />+
</span><span style="font-weight: 400;">The Linux Foundation</span><span style="font-weight: 400;"><br />+
</span><a href="mailto:[email redacted]"><span style="font-weight: 400;">[email redacted]</span></a><span style="font-weight: 400;"> </span></p>+
]]></content:encoded>+
+
+
+
</item>+
<item>+
<title>Cambricon Joins the PyTorch Foundation as a Platinum Member</title>+
<link>https://pytorch.org/blog/cambricon-joins-the-pytorch-foundation-as-a-platinum-member/</link>+
+
<dc:creator><![CDATA[PyTorch Foundation]]></dc:creator>+
<pubDate>Tue, 08 Sep 2026 00:59:13 +0000</pubDate>+
<category><![CDATA[Announcements]]></category>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=162761</guid>+
+
<description><![CDATA[The PyTorch Foundation, a community-driven hub for open source AI under the Linux Foundation, is announcing today that Cambricon has joined as a Platinum member. Founded in 2016, Cambricon is...]]></description>+
<content:encoded><![CDATA[<p><img fetchpriority="high" decoding="async" class="alignnone size-large wp-image-162764" src="https://pytorch.org/wp-content/uploads/2026/09/Cambricon-PyTorch-Foundation-Member-1024x576.png" alt="Cambricon PyTorch Foundation Member" width="1024" height="576" src…+
<p>The PyTorch Foundation, a community-driven hub for open source AI under the Linux Foundation, is announcing today that Cambricon has joined as a Platinum member.</p>+
<p>Founded in 2016, Cambricon is an early pioneer in AI chips, dedicated to the research and development of AI chip products and software. Through a decade of continuous iteration, Cambricon has built a mature, high-performance hardware-software product portfolio that is developer-friendly and highl…+
<p>With its cutting-edge chip technologies and comprehensive foundational software ecosystem, Cambricon enables efficient training and inference of large language models at scale. Beyond that, it has helped turn AI computing into a commercial reality across industries. Its products have achieved bro…+
<p>“We believe that the full potential of AI computing can only be realized through tight software-hardware co-optimization and a thriving open ecosystem,” said Elton Gong, Vice President of Software Engineering at Cambricon, “For us, PyTorch is not just a framework, it is the core of our software e…+
<p>Cambricon follows an “Upstream First” approach and has consistently contributed to the open-source software ecosystem as a long-standing PyTorch contributor. Over recent years, Cambricon’s contributions to PyTorch have spanned several core areas, including torch.compile, Eager Operators, Device R…+
<p>In the meantime, Cambricon has worked closely with the vLLM community to enable Day 0 support for leading open-source large language models, including DeepSeek -V4 and GLM-5.</p>+
<p>Looking ahead, Cambricon will increase its investment in the open-source ecosystem, deepening its collaboration with the PyTorch Foundation in areas such as compile infrastructure enhancement, CI/CD enhancement, device-agnostic support for PyTorch domain-specific libraries. This reflects a more s…+
<p>“For any AI accelerator to succeed at scale, it has to meet developers across the AI lifecycle, from building and optimizing models with PyTorch to serving them efficiently with vLLM,” said Mark Collier, Executive Director of the PyTorch Foundation. “Cambricon’s sustained upstream contributions t…+
<p>As a platinum member, Cambricon is granted one seat to the PyTorch Foundation Governing Board. The Board sets policy through our bylaws, mission and vision statements, describing the overarching scope of foundation initiatives, technical vision, and direction.</p>+
<p>We’re happy to welcome Jin Wang, Senior Director of AI Frameworks and Infrastructure at Cambricon, to our board. Jin Wang leads Cambricon’s AI frameworks and software stack, with his team focused on providing software support for efficient inference and large-scale model training. The team’s work…+
<p>Under Wang’s leadership, the team has been a long-standing contributor to open-source projects, enabling the integration of Cambricon’s products with mainstream AI frameworks and optimizing them for different workloads.</p>+
<p>We’re also pleased to welcome Jing Zhu, Lead Maintainer at Cambricon working on PyTorch, to the PyTorch Foundation’s Technical Advisory Council (TAC). Jing Zhu serves as a lead maintainer on Cambricon’s PyTorch team, focusing on integrating Cambricon’s products with the PyTorch ecosys…+
<p>To learn more about how your organization can join the PyTorch Foundation, visit our <a href="https://pytorch.org/join/">website</a>.</p>+
<h3>About PyTorch Foundation</h3>+
<p>The PyTorch Foundation is the vendor-neutral home for the open source intelligence layer developers use for training, optimizing, serving, orchestrating, and running models on any chip in any cloud for any agent. As a community-driven hub hosted by the Linux Foundation, the PyTorch Foundation sup…+
]]></content:encoded>+
+
+
+
</item>+
<item>+
<title>PyTorch x Hugging Face in Bengaluru: Building India’s Next Generation of ML Systems Contributors</title>+
<link>https://pytorch.org/blog/pytorch-x-hugging-face-in-bengaluru-building-indias-next-generation-of-ml-systems-contributors/</link>+
+
<dc:creator><![CDATA[Sumantro Mukherjee, Red Hat]]></dc:creator>+
<pubDate>Mon, 07 Sep 2026 13:05:37 +0000</pubDate>+
<category><![CDATA[Blog]]></category>+
<category><![CDATA[Community]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=158011</guid>+
+
<description><![CDATA[TL;DR More than 170 students, engineers, researchers, and open-source contributors gathered in Bengaluru for a technical evening hosted by Red Hat and Hugging Face around PyTorch, large-scale inference, reinforcement learning...]]></description>+
<content:encoded><![CDATA[<h3>TL;DR</h3>+
<p><span style="font-weight: 400;">More than 170 students, engineers, researchers, and open-source contributors gathered in Bengaluru for a technical evening hosted by Red Hat and Hugging Face around PyTorch, large-scale inference, reinforcement learning environments, distributed training, and next-…+
<p><span style="font-weight: 400;">What stood out most was not only the technical range of the talks, but the shared conviction behind them. India has no shortage of talent using AI and ML systems. The deeper opportunity now is to help more students and practitioners become builders and maintainers …+
<p><img decoding="async" class="alignnone size-large wp-image-159047" src="https://pytorch.org/wp-content/uploads/2026/08/Bengaluru-Event-1024x345.jpg" alt="Bengaluru Event" width="1024" height="345" srcset="https://pytorch.org/wp-content/uploads/2026/08/Bengaluru-Event-1024x345.jpg 1024w, https://p…+
<h2><span style="font-weight: 600;">Setting the Tone: From AI Users to AI Infrastructure Builders</span></h2>+
<p><a href="https://www.linkedin.com/in/sudhir-dharanendraiah-80a0867/"><span style="font-weight: 400;">Sudhir Dharanendraiah</span></a><span style="font-weight: 400;"> opened the evening by framing a challenge that resonated across the room: India should not remain merely a large consumer base for …+
<p><span style="font-weight: 400;">That framing mattered. It shifted the event away from product demos and toward systems thinking. The conversation was not just about how to call an API or fine-tune a model, but about how the underlying machinery works: what makes inference efficient, what makes re…+
<p><span style="font-weight: 400;">For students, early-career engineers, and startup teams in the room, this was an important signal. The next wave of innovation in AI will not belong only to those consuming models. It will also belong to those improving the compiler paths, the kernel libraries, the…+
<h2><span style="font-weight: 600;"><img decoding="async" class="wp-image-163687 size-large alignleft" src="https://pytorch.org/wp-content/uploads/2026/08/DSC06649-1024x683.jpg" alt="" width="1024" height="683" srcset="https://pytorch.org/wp-content/uploads/2026/08/DSC06649-1024x683.jpg 1024w, https…+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2></h2>+
<h2><span style="font-weight: 600;">Profiling in PyTorch: Making Performance Visible</span></h2>+
<p><a href="https://www.linkedin.com/in/arig23498/"><span style="font-weight: 400;">Aritra Roy Gosthipaty</span></a><span style="font-weight: 400;"> from Hugging Face opened the technical program with a practical talk on profiling in PyTorch built around a simple but durable principle: what you cann…+
<p><span style="font-weight: 400;">Rather than treating performance as a vague outcome, the session broke profiling into a repeatable workflow. Aritra showed how to annotate regions of interest with `torch.profiler.record_function`, wrap execution with `torch.profiler.profile`, and use schedules to …+
<p><span style="font-weight: 400;">One particularly useful thread in the talk was the idea of being “overhead bound.” Small workloads can easily create the illusion that GPU acceleration is underperforming, when in reality the CPU-side launch and orchestration costs dominate the run. By …+
<p><span style="font-weight: 400;">For an audience full of people building or debugging real systems, this was a strong starting point. Profiling is often the difference between disciplined optimization and superstition.</span><span style="font-weight: 400;"><br />+
</span><span style="font-weight: 400;">Slides can be found </span><a href="https://drive.google.com/file/d/11ZAxSbV6sB1nrEstA-dxS5xrLNTRhqDB/view?usp=sharing"><span style="font-weight: 400;">here</span></a><span style="font-weight: 400;">. More reading materials can be found </span><a href="https://…+
<h2><span style="font-weight: 600;"><img decoding="async" class="alignnone size-large wp-image-159200" src="https://pytorch.org/wp-content/uploads/2026/08/PyTorch-x-Hugging-Face-in-Bengaluru-Photo-1024x683.jpg" alt="" width="1024" height="683" srcset="https://pytorch.org/wp-content/uploads/2026/08/P…+
<h2><span style="font-weight: 600;">SGLang, Transformers, and Kernels: The New Shape of Inference</span></h2>+
<p><span style="font-weight: 400;">The next Hugging Face talk by </span><a href="https://www.linkedin.com/in/adarshxs/"><span style="font-weight: 400;">Adarsh</span></a><span style="font-weight: 400;"> focused on modern LLM inference through the lens of SGLang, the Transformers backend, and the eme…+
<p><span style="font-weight: 400;">The talk unpacked why inference is structurally difficult in large language models. Prefill is compute-bound and highly parallel, while decode is sequential, memory-sensitive, and dominated by the cost of repeatedly interacting with KV cache. From there, the sessio…+
<p><span style="font-weight: 400;">A major concept in the presentation was RadixAttention. Instead of discarding KV cache state once a request is complete, SGLang keeps previously seen prefixes in a radix-tree-based cache with LRU behavior. That design is especially compelling in workloads with shar…+
<p><span style="font-weight: 400;">The talk also highlighted a productive division of labor between Hugging Face Transformers and serving engines like SGLang. Transformers remains the source of truth for model definitions, configuration parsing, tokenizers, templates, and weight formats. SGLang then…+
<p><span style="font-weight: 400;">The final segment on Hugging Face Kernels widened the picture further. As custom operators and accelerator-specific kernels become more central to ML performance, build fragmentation has become a real problem. Different toolchains, backend combinations, and compati…+
<p><span style="font-weight: 400;">Together, these ideas showed how inference is evolving: not as a single monolithic stack, but as a layered collaboration between model definitions, serving runtimes, compiler-friendly execution paths, and reusable kernel infrastructure. Slides can be found </span><…+
<h2><span style="font-weight: 600;"><img decoding="async" class="alignnone size-large wp-image-159201" src="https://pytorch.org/wp-content/uploads/2026/08/PyTorch-x-Hugging-Face-in-Bengaluru-3-1024x683.jpg" alt="" width="1024" height="683" srcset="https://pytorch.org/wp-content/uploads/2026/08/PyTor…+
<h2><span style="font-weight: 600;">RL Environments 101: Why the Next Scaling Axis Is the Environment</span></h2>+
<p><a href="https://www.linkedin.com/in/adithya-s-kolavi/"><span style="font-weight: 400;">Adithya S Kolavi</span></a><span style="font-weight: 400;">’s session on RL environments brought the post-training story into focus. The talk traced a familiar arc from pretraining to supervised fine-tuning to…+
<p><span style="font-weight: 400;">The key insight of the talk was that once a task can be graded by a program, it can become an environment in which a model learns. That shift sounds abstract, but the presentation made it concrete. An RL environment was described not as a black box, but as a struct…+
<p><span style="font-weight: 400;">This framing helps explain why reinforcement learning for LLMs is both powerful and difficult. Classical RL environments standardized interaction for control problems years ago, but agentic LLM training introduces many more moving parts. A model may need tools, san…+
<p><span style="font-weight: 400;">That is where </span><a href="https://github.com/huggingface/OpenEnv"><span style="font-weight: 400;">OpenEnv</span></a><span style="font-weight: 400;"> entered the discussion. Presented as a common shape for LLM environments, OpenEnv extends the spirit of Gym-styl…+
<p><span style="font-weight: 400;">The later sections of the talk pushed on the ecosystem implication: if better environments lead to better models, then generating many high-quality environments becomes a strategic advantage. Coding tasks are especially attractive because they are verifiable, deter…+
<p><span style="font-weight: 400;">This was one of the most energizing talks of the evening because it gave students and practitioners a tangible frontier to contribute to. Not everyone will build a foundation model, but many can help create the environments, verifiers, tools, and benchmarks that ma…+
<h2><span style="font-weight: 600;"><img decoding="async" class="alignnone size-large wp-image-159202" src="https://pytorch.org/wp-content/uploads/2026/08/PyTorch-x-Hugging-Face-in-Bengaluru-2-1024x683.jpg" alt="" width="1024" height="683" srcset="https://pytorch.org/wp-content/uploads/2026/08/PyTor…+
<h2><span style="font-weight: 600;">Scaling Up, One Dimension at a Time</span></h2>+
<p><a href="https://www.linkedin.com/in/mansi-agarwal-a72bbab2/"><span style="font-weight: 400;">Mansi Agarwal</span></a><span style="font-weight: 400;"> from the Red Hat PyTorch engineering team brought the audience into the heart of modern distributed training with a talk on DeviceMesh, DTensor, a…+
<p><span style="font-weight: 400;">The talk began by naming a pain point that anyone who has worked on large training jobs will recognize: combining different forms of parallelism has historically required too much manual plumbing. Data parallelism, tensor parallelism, and pipeline parallelism often…+
<p><span style="font-weight: 400;">The promise of the newer PyTorch abstractions, as Mansi argued, is composability. DeviceMesh lets engineers describe a cluster as an n-dimensional topology. DTensor makes tensors aware of how they are distributed across that topology. FSDP2 then rebuilds sharded da…+
<p><span style="font-weight: 400;">This matters because it changes the developer experience as much as the runtime behavior. Instead of hand-crafting process groups and injecting custom communication into model code, engineers can reason in terms of mesh dimensions and placement rules. Adding tensor…+
<p><span style="font-weight: 400;">The session also did not hide the trade-offs. DTensor’s eager-mode overhead, incomplete operator coverage, and the limits of greedy sharding propagation are real constraints. But that honesty made the overall message stronger: composability in distributed training …+
<p><span style="font-weight: 400;">For many attendees, this talk was a window into a level of systems design they may not encounter in day-to-day model usage, but absolutely will encounter if they choose to contribute to the core stack. Slides can be found </span><a href="https://drive.google.com/fi…+
<p><img decoding="async" class="alignnone size-large wp-image-159203" src="https://pytorch.org/wp-content/uploads/2026/08/PyTorch-x-Hugging-Face-in-Bengaluru-1024x683.jpg" alt="" width="1024" height="683" srcset="https://pytorch.org/wp-content/uploads/2026/08/PyTorch-x-Hugging-Face-in-Bengaluru-1024…+
<h2><span style="font-weight: 600;">Zero-Copy GPU-to-GPU Communication in PyTorch</span></h2>+
<p><a href="https://www.linkedin.com/in/arkadip-maitra/"><span style="font-weight: 400;">Arkadip Maitra</span></a><span style="font-weight: 400;"> closed the evening with a deep systems talk on zero-copy GPU-to-GPU communication in PyTorch, moving the discussion down to the communication substrate t…+
<p><span style="font-weight: 400;">The talk began with `c10d`, PyTorch’s default distributed communication layer, and why it served the ecosystem well for a long time. It offered a general-purpose abstraction across CPU and GPU backends and fit the era in which most distributed workloads were bulk-s…+
<p><span style="font-weight: 400;">But the assumptions around communication are changing. Network interfaces have evolved, GPUDirect RDMA has matured, NVLink paths have strengthened, and training fabrics have become more topology-aware and specialized. As cluster sizes and communication patterns cha…+
<p><span style="font-weight: 400;">That is why the zero-copy path discussed in the talk is so important. By avoiding unnecessary copy steps, PyTorch can reduce thread-block consumption and deliver meaningful communication speedups, especially in message-size regimes that matter in real training and …+
<p><span style="font-weight: 400;">This was a fitting close to the event because it reinforced a recurring lesson from the evening: high-level model performance often depends on low-level engineering choices that most users never see. Helping more practitioners understand those layers is part of wha…+
<h2><span style="font-weight: 600;"><img decoding="async" class="alignnone wp-image-163688 size-large" src="https://pytorch.org/wp-content/uploads/2026/09/DSC06880-1024x683.jpg" alt="" width="1024" height="683" srcset="https://pytorch.org/wp-content/uploads/2026/09/DSC06880-1024x683.jpg 1024w, https…+
<h2><span style="font-weight: 600;">Why This Collaboration Matters</span></h2>+
<p><span style="font-weight: 400;">What made the event distinctive was not just that it featured speakers from both Red Hat and Hugging Face, but that the collaboration surfaced a coherent view of the stack.</span></p>+
<p><span style="font-weight: 400;">Hugging Face brought perspectives from profiling, inference infrastructure, and RL post-training workflows. Red Hat’s PyTo</span>rch engineering team brought perspectives from distributed training internals and communication primitives. Put together, the talks form…+
<p><span style="font-weight: 400;">For Indian students and AI practitioners, that kind of ecosystem view is invaluable. It shortens the distance between “using AI” and “contributing to AI systems.” It shows that open-source contribution is not confined to model releases or ap…+
<p><span style="font-weight: 400;">At a time when many people are asking how India can participate more deeply in the future of AI, this event offered a credible answer: by joining the communities that build the core layers, and by treating technical collaboration as a way to widen the pipeline from…+
<h2><span style="font-weight: 600;"> Looking Ahead</span></h2>+
<p><span style="font-weight: 400;">With more than 170 attendees, the evening made one thing clear: there is real appetite in India for technically serious, systems-oriented ML community events. The energy in the room suggested that students want more than introductions, practitioners want more than …+
<p><span style="font-weight: 400;">If this event is any indication, collaborations between Red Hat, Hugging Face, and the wider PyTorch community can do more than host good meetups. They can help build a local culture of contribution around the open ML stack itself, one where the next generation of …+
]]></content:encoded>+
+
+
+
</item>+
<item> <title>Your Guide to Hardware Acceleration & Compute Infrastructure at PyTorch Conference North America 2026</title> <link>https://pytorch.org/blog/your-guide-to-hardware-acceleration-compute-infrastructure-at-pytorch-conference-north-america-2026/</link> @
@@ -1208,7 +1384,7 @@ <guid isPermaLink="false">https://pytorch.org/?p=157323</guid> <description><![CDATA[The keynote lineup is set for PyTorch Conference North America 2026, taking place October 20–21 in San Jose, California. The program includes sessions on PyTorch updates, native PyTorch on Trainium,...]]></description>-
<content:encoded><![CDATA[<p><img fetchpriority="high" decoding="async" class="aligncenter wp-image-119439 size-full" src="https://pytorch.org/wp-content/uploads/2025/12/PyTorch-Conference-North-America-2026.png" alt="PyTorch Conference North America 2026" width="1200" height="630" srcset=…+
<content:encoded><![CDATA[<p><img decoding="async" class="aligncenter wp-image-119439 size-full" src="https://pytorch.org/wp-content/uploads/2025/12/PyTorch-Conference-North-America-2026.png" alt="PyTorch Conference North America 2026" width="1200" height="630" srcset="https://pytorch.org/…<p>The keynote lineup is set for <strong>PyTorch Conference North America 2026</strong>, taking place October 20–21 in San Jose, California. The program includes sessions on PyTorch updates, native PyTorch on Trainium, agent workloads, model customization, open science, and more.</p><p>Featured keynote speakers and their scheduled sessions include:</p><ul>@
@@ -1234,372 +1410,8 @@ -
</item>-
<item>-
<title>Harnessing AI for Day-One Model Enablement</title>-
<link>https://pytorch.org/blog/harnessing-ai-for-day-one-model-enablement/</link>-
-
<dc:creator><![CDATA[IBM Spyre Team]]></dc:creator>-
<pubDate>Thu, 20 Aug 2026 15:45:59 +0000</pubDate>-
<category><![CDATA[Blog]]></category>-
<guid isPermaLink="false">https://pytorch.org/?p=154257</guid>-
-
<description><![CDATA[TL;DR The AI model landscape never stops moving, and the software stack that runs those models is always a step behind: even on a mature compilation stack, a new model...]]></description>-
<content:encoded><![CDATA[<p><strong>TL;DR</strong></p>-
<p>The AI model landscape never stops moving, and the software stack that runs those models is always a step behind: even on a mature compilation stack, a new model family often arrives with some novel module that doesn’t lower well, and on new hardware – where a young stack meets a whol…-
<h2>The problem: catching up to an evolving model landscape</h2>-
<p>The model landscape never stops moving. New architectures and checkpoints appear constantly — each a particular composition of operations, shapes, and numerical ranges — and the software that has to run them is always a step behind. Even on the most mature, widely-deployed stack, a brand-new mode…-
<p> </p>-
<p><em><strong><img decoding="async" class="alignnone size-large wp-image-154262" src="https://pytorch.org/wp-content/uploads/2026/08/model-landscape-timeline-1024x715.png" alt="model-landscape-timeline" width="1024" height="715" srcset="https://pytorch.org/wp-content/uploads/2026/08/model-landscape…-
<p>Running a model on any device requires a software stack – a compiler, a set of operator lowerings, and a runtime – that maps it onto the hardware’s cores, memory hierarchy, and number formats. That stack is a large, evolving piece of software, and a new architecture can exercise…-
<p>New hardware is where the gap is at its widest. A new accelerator is not just a chip: it ships with a young stack still growing into the needs of different models. Here it is not one new model meeting a mature stack, but a whole ecosystem of models meeting a stack that is still being built –…-
<p> </p>-
<p><em><strong><img decoding="async" class="alignnone size-large wp-image-154261" src="https://pytorch.org/wp-content/uploads/2026/08/accelerator_timeline-1024x508.png" alt="accelerator_timeline" width="1024" height="508" srcset="https://pytorch.org/wp-content/uploads/2026/08/accelerator_timeline-10…-
<p>New hardware is the extreme case, but the challenge is a general one: keeping the software in step with the model landscape, whether the stack is mature or brand new. Below we describe a strategy for bridging that gap, and demonstrate it through a concrete instance – running stock HuggingFa…-
<h2>The strategy: adapters as a bridge between models and stack</h2>-
<p>The strategy we describe is a thin layer of runtime patches – adapters – that allow a stock model to run on a given platform today, without waiting for every underlying gap in the stack to be closed first. We assume that the platform already provides a PyTorch compiler that lowers<br …-
core tensor operations in ordinary torch code – including matrix multiplications, elementwise operations, and reductions – onto the target hardware, so most of the model logic runs through it unchanged. When some operation in a model doesn’t yet have a clean path through the stack,…-
<p>Rather than blocking on the state of the stack, adapters provide a practical bridge across it as it stands today. A useful analogy is a large construction project. Around a building with active construction – whether it is still going up or already standing and being renovated — there is al…-
<p>Adapters play the same role. An individual adapter is usually transitional: its job is to carry a model across one specific gap in the stack as it stands today, and it is designed to be removed once that gap is closed — as the platform matures, new optimizations are integrated, more architectures…-
<p>The adapter layer, however, is permanent in a way no individual adapter is. Because the model landscape never stops moving, there is always some new gap the stack hasn’t caught up to yet — so even as old scaffolding comes down, new scaffolding goes up elsewhere. That layer connects two thin…-
<p>What makes this strategy practical at the scale of thousands of models is AI. Historically, enabling each new model family was a slow, specialist effort, requiring a great deal of work for every hardware platform and chip architecture. Coding agents change the economics, turning what used to be a…-
<h2>The platform: Spyre and torch-spyre</h2>-
<p>The hardware in our example is <strong>Spyre</strong>, IBM’s AI accelerator, built on the AIU (Artificial Intelligence Unit). It is made of small cores connected by a high-bandwidth ring, each with its own local scratchpad memory and arrays of processing elements that carry out the matrix m…-
<p>Spyre’s memory and compute operate on fixed-size chunks called sticks – 128 bytes, or 64 values in fp16 – and it expects tensor dimensions to line up on stick boundaries (see the <a href="https://github.com/torch-spyre/RFCs/blob/main/0047-TiledTensors/0047-TiledTensorsRFC.md">Ti…-
<p>The software stack that maps a model onto this hardware is <strong><a href="https://github.com/torch-spyre/torch-spyre">torch-spyre</a></strong>, a PyTorch backend that compiles and runs ordinary torch code on Spyre: a model is compiled into a plan the hardware can execute. And the adapters ̵…-
<h2>How AI helps build adapters</h2>-
<p>An adapter lives in the gap between two codebases. On one side is the model as Transformers expresses it – the modules, the attention blocks, the way a particular architecture wires its RoPE and its norms together. On the other is torch-spyre, the compiler and runtime that lower that comput…-
<p>Coding agents can now follow the internal logic of an entire transformer model at every level, from a single module up through the attention blocks to the full forward pass. They can trace the same computation down through torch-spyre’s lowering logic to see how it is meant to run on the ha…-
<p> </p>-
<p><img decoding="async" class="alignnone size-large wp-image-154260" src="https://pytorch.org/wp-content/uploads/2026/08/how-ai-helps-cycleagent-1024x588.png" alt="how-ai-helps-cycleagent" width="1024" height="588" srcset="https://pytorch.org/wp-content/uploads/2026/08/how-ai-helps-cycleagent-1024x…-
<p>A few things make this work in practice. We keep a knowledge base that describes both the hardware and the software stack as it evolves, so an agent can ground its reasoning in how Spyre actually behaves rather than in generic assumptions. We give it direct access to the source and GitHub of both…-
<p>Importantly, every adapter we add is also a record of which adaptations work, and for what reason. A new model is rarely a completely new problem: it is usually a variation on an architecture we have already brought up, and the closest existing adapter is both a template to imitate and a source o…-
<p> </p>-
<p><em><strong><img decoding="async" class="alignnone size-large wp-image-154259" src="https://pytorch.org/wp-content/uploads/2026/08/spyre-embedding-adapter-coverage-1024x481.png" alt="spyre-embedding-adapter-coverage" width="1024" height="481" srcset="https://pytorch.org/wp-content/uploads/2026/08…-
<p>At the same time, human expertise and supervision remain essential. Two failure modes recur.</p>-
<p>The first is that <strong>localization is genuinely difficult</strong>. When a model that is correct on CPU or GPU produces the wrong output on the device, narrowing down the exact source of the mislowering is rarely a matter of reading the code once. It means reproducing different parts of the m…-
<p>The second is that <strong>the agent will often draw the wrong conclusion from a targeted experiment</strong>. A large numerical discrepancy somewhere deep inside the model is a clue, not a verdict. It is easy — for a person and an agent alike — to find an alarming difference in some intermediate…-
The division of labor follows from this. AI handles the wide, repetitive, cross-referencing part of the work: reading two large codebases at once, recognizing an architecture, and drafting a first adapter by analogy to the ones that came before. Human expertise handles the diagnosis: localizing the …-
<h2>Adapters as a validation tool</h2>-
<p>Each adapter does more than carry its model onto the device. In doing so, it also exposes exactly where the stack underneath still needs work.</p>-
<p>This is a general property of bridging a mature ecosystem to a young stack. The gaps in such a stack rarely live in a single operation; they live in combinations no unit test anticipated, and the only reliable way to surface them is to run real workflows end to end against a trusted reference. Ad…-
<p>In our example, the Torch-Spyre team works primarily at the lower levels of the software stack – the compiler, the operator lowerings, the runtime that places work on the device. At that level it is genuinely hard to predict how a real model will behave. A production model is a particular c…-
<p>Running stock HuggingFace models end-to-end is what makes these gaps visible. An end-to-end divergence from a CPU/GPU reference is a strong signal that something in the platform needs attention. In practice the issues that surface fall into a few recurring families:</p>-
<ul>-
<li><strong>Missing lowering paths</strong> – a block or fused shape the stack cannot yet lower, even when the individual ops are already working in isolation.</li>-
<li><strong>Device-only numerical behavior</strong> – values that overflow or turn into NaNs on the device where the CPU/GPU reference stays finite.</li>-
<li><strong>Alignment and padding assumptions</strong> – shapes that quietly corrupt results when they don’t match what the hardware expects.</li>-
</ul>-
<p>What ties these together is that they are largely fusion-dependent. When a full model is compiled, the stack fuses different ops and shapes into combined kernels, and which things get fused depends on the surrounding graph. That is why an operator can pass low-level testing on its own and still f…-
<p>Real models also bring real data. Low-level tests typically feed operators random tensors, but trained weights and the activations they produce have structure that random inputs don’t — particular value ranges, near-constant rows, occasional large outliers. That structure is often exactly w…-
<p>This also builds on itself. The temporary adaptations already in place each get a model past a known crash or mislowering at one point in the stack. Doing so is what lets execution reach further in and expose the next gap, which was invisible while the model was still failing earlier. So each ada…-
<p>So the adapters play a double role – they are an integration layer that lets Transformer models run on Spyre today, and they are a validation tool that continuously exercises the platform against the messiness of real architectures. Every model we bring up is also a test case, and the failu…-
<h2>Adaptation examples</h2>-
<p>To illustrate the range of adaptations we implement in <a href="https://github.com/torch-spyre/hf-adapters">HF-adapters</a>, we’ll walk through two examples, and then step back to see what they have in common, and where they differ.</p>-
<h3>Replacing an operator</h3>-
<p>The smallest patches replace one currently unsupported operation with a mathematically identical one.</p>-
<p>A concrete example is the <code>gelu_new</code> activation used by decoders like GPT-2 and GPT-Neo. Its tanh-approximation form, <code>0.5 x (1 + tanh(sqrt(2/pi) (x + 0.044715 x³)))</code>, computes the cube with <code>torch.pow(x, 3.0)</code>. This raise to a power op doesn’t yet lower wel…-
<pre><code class="language-python">def gelu_forward(self, x):
-
if x.device.type == "spyre":
-
# x*x*x lowers on Spyre; torch.pow(x, 3) does not
-
inner = COEFF * (x + 0.044715 * (x * x * x))
-
else:
-
# CPU keeps the stock form → bit-identical to HuggingFace
-
inner = COEFF * (x + 0.044715 * torch.pow(x, 3.0))
-
return 0.5 * x * (1.0 + torch.tanh(inner))
-
</code></pre>-
<p>These fixes are satisfying precisely because they’re invisible: the model’s math is untouched, and the only thing that changes is the shape of the instruction the stack sees.</p>-
<h3>Reshaping the data</h3>-
<p>Other issues can’t be fixed by swapping a single operator. They come from <em>how much</em> data is being processed and <em>how it’s laid out</em>, and the fix is to reshape the data before it reaches the hardware — through changes related to work division, padding, or other structura…-
<p>A concrete example is the model’s final output projection – the LM head – that turns the model’s internal representation into a score for every word in its vocabulary. On Spyre this is one large matrix multiply, and the device tackles it by slicing the vocabulary dimension…-
<p>The fix is to pad the vocabulary: we round it up, adding a small number of unused entries, until the block count factors into pieces the work division can distribute evenly. The padding entries carry no meaning and are ignored, so the model’s output is unchanged – but the matrix multi…-
<pre><code class="language-python">def pad_lm_head(model):
-
"""Grow the vocabulary dimension so that the output matrix multiply can be
-
efficiently divided across the device cores.
-
-
The extra rows are unused padding and do not affect the model output.
-
"""
-
padded_vocab = round_up_to_divisible_size(model.vocab_size)
-
...
-
</code></pre>-
<h3>What these examples show</h3>-
<p>Taken together, these two examples span the range of adaptations we deal with. Replacing an operator is a low-level change to a single instruction: the problem is that one specific operation doesn’t lower well, and the fix is to rewrite that exact operation; there’s no gap between the…-
<p>The common thread is that in neither case do we change <em>what</em> is computed. The model’s math is preserved exactly – we only change the form it takes so the existing computation can succeed on the stack as it stands today.</p>-
<h2>Conclusion</h2>-
<p>The model landscape never stops moving, and every backend that runs it – each GPU architecture, each accelerator, each maturing compiler stack – is left playing catch-up. Historically this was slow, specialist work that kept new models from reaching cross-platform deployment. What cha…-
<p>We demonstrated this on Spyre, at the hardest end of the spectrum – a brand-new accelerator with a stack still being built. But nothing about the loop is specific to Spyre. Whenever a new model family lands ahead of the stack that must run it – whether that stack is brand new or matur…-
]]></content:encoded>-
-
-
-
</item>-
<item>-
<title>FP8 Training on AMD GPUs with TorchTitan and TorchAO: Upstreaming Performance Improvements</title>-
<link>https://pytorch.org/blog/fp8-training-on-amd-gpus-with-torchtitan-and-torchao-upstreaming-performance-improvements/</link>-
-
<dc:creator><![CDATA[AMD: Rishi Sinha, Yuankai Chen, Liz Li, Shekhar Pandey, Wen Chen, Xiaobo Chen, Yao Fu, Zhenyu Gu, Andy Luo, Peng Sun META: Matthias Reso, Hamid Shojanazeri, TorchAO team, TorchTitan team]]></dc:creator>-
<pubDate>Thu, 13 Aug 2026 16:00:31 +0000</pubDate>-
<category><![CDATA[Blog]]></category>-
<guid isPermaLink="false">https://pytorch.org/?p=154222</guid>-
-
<description><![CDATA[At the PyTorch Conference 2025, we demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters using Primus-Turbo, an AMD optimization library for training frameworks such as TorchTitan. We have since upstreamed those AMD optimizations so TorchTitan supports AMD…-
<content:encoded><![CDATA[<p><span style="font-weight: 400;">At the PyTorch Conference 2025, we demonstrated linear scaling beyond 1,000 GPUs on AMD Instinct clusters using Primus-Turbo, an AMD optimization library for training frameworks such as TorchTitan. We have since upstreamed those …-
<p><span style="font-weight: 400;">On dense models, FP8 training delivers a 13.4% throughput gain over BF16 on Llama3-8B (</span><a href="https://github.com/pytorch/ao/pull/2736"><span style="font-weight: 400;">#2736</span></a><span style="font-weight: 400;">) as shown in Figure 1. On MOE architectu…-
<p><span style="font-weight: 400;">This blog covers the FP8 optimizations that deliver these gains, from major kernel acceleration to the Triton fusion pipeline that narrowed the quantization overhead gap on MoE models. Getting there took three pieces of work: </span></p>-
<ol>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Adding native support for AMD’s FP8 number format </span></li>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Enabling grouped GEMM for Mixture-of-Experts models on ROCm</span></li>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Building a Triton fusion pipeline that reduced the quantization overhead</span></li>-
</ol>-
<p> </p>-
<table>-
<tbody>-
<tr>-
<td><b>Workload</b></td>-
<td><b>Optimization</b></td>-
<td><b>Result</b></td>-
<td><b>PR</b></td>-
</tr>-
<tr>-
<td><span style="font-weight: 400;">Llama3-8B (dense)</span></td>-
<td><span style="font-weight: 400;">Rowwise FP8 vs BF16</span></td>-
<td><span style="font-weight: 400;">+13.4% throughput</span></td>-
<td><a href="https://github.com/pytorch/ao/pull/2736"><span style="font-weight: 400;">#2736</span></a></td>-
</tr>-
<tr>-
<td><span style="font-weight: 400;">DeepSeek-MoE-16B</span></td>-
<td><span style="font-weight: 400;">Backward transpose removal + fusion</span></td>-
<td><span style="font-weight: 400;">4.2x backward pass</span></td>-
<td><a href="https://github.com/pytorch/ao/pull/3972"><span style="font-weight: 400;">#3972</span></a><span style="font-weight: 400;">, </span><a href="https://github.com/pytorch/ao/pull/4069"><span style="font-weight: 400;">#4069</span></a></td>-
</tr>-
<tr>-
<td><span style="font-weight: 400;">DeepSeek-V3 671B</span></td>-
<td><span style="font-weight: 400;">Colwise scales coalescing</span></td>-
<td><span style="font-weight: 400;">6.2x per MoE layer (7,290→1,170µs)</span></td>-
<td><a href="https://github.com/pytorch/ao/pull/4113"><span style="font-weight: 400;">#4113</span></a></td>-
</tr>-
<tr>-
<td><span style="font-weight: 400;">DeepSeek-V3 671B</span></td>-
<td><span style="font-weight: 400;">Forward pass fusion</span></td>-
<td><span style="font-weight: 400;">+17% end-to-end; recovers 89% of FP8 gap</span></td>-
<td><a href="https://github.com/pytorch/ao/pull/4311"><span style="font-weight: 400;">#4311</span></a></td>-
</tr>-
</tbody>-
</table>-
<p><i><span style="font-weight: 400;"><img decoding="async" class="alignnone size-large wp-image-154235" src="https://pytorch.org/wp-content/uploads/2026/08/Rowwise-FP8-training-throughput-on-8×MI300X-GPU-with-Llama3-8B-e1786571048745-1024x544.png" alt="Rowwise FP8 training throughput on 8×MI300X GP…-
<h2><b>AMD FP8 format in TorchAO</b></h2>-
<p><span style="font-weight: 400;">Each linear layer performs three matrix multiplications: the forward pass, gradient input, and gradient weight update. FP8 training quantizes these operations from 16 bits to 8 bits, greatly improving throughput. AMD Instinct GPUs implement a variant of the FP8 for…-
<table>-
<thead>-
<tr>-
<th><span style="font-weight: 400;">Property</span></th>-
<th><span style="font-weight: 400;">e4m3fnuz (AMD)</span></th>-
</tr>-
</thead>-
<tbody>-
<tr>-
<td><span style="font-weight: 400;">Max value</span></td>-
<td><span style="font-weight: 400;">240</span></td>-
</tr>-
<tr>-
<td><span style="font-weight: 400;">NaN/Inf encodings</span></td>-
<td><span style="font-weight: 400;">No</span></td>-
</tr>-
<tr>-
<td><span style="font-weight: 400;">Hardware</span></td>-
<td><span style="font-weight: 400;">MI300X, MI325X, MI350X</span></td>-
</tr>-
</tbody>-
</table>-
<p><span style="font-weight: 400;">The FP8 capabilities demonstrated in Primus-Turbo were upstreamed directly to TorchAO and TorchTitan as displayed in Figure 2, spanning three areas: hardware-aware FP8 format support, MoE grouped GEMM enablement on ROCm, a Triton kernel fusion pipeline that reduced…-
<p><i><span style="font-weight: 400;"><img decoding="async" class="alignnone size-large wp-image-154236" src="https://pytorch.org/wp-content/uploads/2026/08/TorchTitan-FP8-training-software-stack-on-ROCm-1024x626.png" alt="TorchTitan FP8 training software stack on ROCm" width="1024" height="626" src…-
<p><span style="font-weight: 400;">The TorchAO library initially did not support the same numerical format as AMD so it was computing scales against a different max value. On AMD Instinct GPU, where e4m3fnuz has a max of 240, this produced silently wrong results: tensors were scaled into a range tha…-
<ul>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Auto-detect the platform and select the correct FP8 dtype and max value, instead of hardcoding NVIDIA e4m3fn: TorchAO </span><a href="https://github.com/pytorch/ao/pull/1142"><span style="font-weight: 400;">#1142</span></a>…-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Report correct MI300X peak FLOPS so MFU numbers are accurate : TorchTitan </span><a href="https://github.com/pytorch/torchtitan/pull/920"><span style="font-weight: 400;">#920</span></a></li>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Add platform-specific loss baselines for FNUZ numerics : TorchTitan </span><a href="https://github.com/pytorch/torchtitan/pull/2156"><span style="font-weight: 400;">#2156</span></a></li>-
</ul>-
<p><span style="font-weight: 400;">FP8 scaling can be applied at different granularities: a single scale per tensor (tensorwise, fastest but coarsest), a scale per row (rowwise, better accuracy), per fixed-size tile (blockwise), or per group packed alongside the data (MXFP8). TorchAO and TorchTitan …-
<h2><b>Scaling FP8 to MoE Architectures</b></h2>-
<p><span style="font-weight: 400;">Mixture-of-Experts (MoE) models like DeepSeek V3 and Llama 4 route each token to a subset of experts, producing variable-size batches that must be processed through a grouped GEMM (Figure 3). Unlike dense models, where every linear layer has the same shape, grouped…-
<p><span style="font-weight: 400;">We enabled FP8 grouped GEMM on ROCm by adapting the quantization pipeline to use the correct dtype and dispatch for AMD via the Composable Kernel backend (#3955).</span></p>-
<p><i><span style="font-weight: 400;"><img decoding="async" class="alignnone size-large wp-image-154237" src="https://pytorch.org/wp-content/uploads/2026/08/MoE-FP8-grouped-GEMM-pipeline-on-ROCm-1024x412.png" alt="MoE FP8 grouped GEMM pipeline on ROCm" width="1024" height="412" srcset="https://pytor…-
<h2><b>Triton Kernel Optimization</b></h2>-
<p><span style="font-weight: 400;">With correctness established, we turned to performance. The FP8 quantization pipeline in TorchAO converts tensors to FP8 through a multi-step chain:</span></p>-
<ol>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Compute per-row/column absolute max (absmax)</span></li>-
<li style="font-weight: 400;" aria-level="1"><span style="font-weight: 400;">Derive the scale factor and apply it</span></li>Diff display stops at 400 lines. The line counts above are from the whole diff. 125 lines shown here cut at 300 characters. 2 email addresss masked in this display. See the About page on personal data. The raw artifact at this commit is linked above.