Change
ac8d7c9
ac8d7c9d415a6af73f2ac043796e90f7bec93480 · commit on GitHub
pytorch-blog-feed: changed (201668 bytes, HTTP 200)
raw/pytorch-blog-feed/response.xml modified
- Source
- pytorch-blog-feed
- Lines added
- +168
- Lines removed
- -172
- Stored bytes at this commit
- 201,668
- Timestamp
- origin
- Raw artifact at this commit
- raw/pytorch-blog-feed/response.xml
Recorded headers
| observed_at | 2026-10-06T06:21:09.578Z |
|---|---|
| origin_date | 2026-10-06T05:18:28.000Z |
| status | 200 |
| final URL | https://pytorch.org/blog/feed/ |
| etag | "6e359a17025d75ce914cbb99201711c6" |
| last-modified | Mon, 05 Oct 2026 20:45:50 GMT |
| date | Tue, 06 Oct 2026 06:21:09 GMT |
| age | 3761 |
| cache-control | public, max-age=60, s-maxage=43200, stale-while-revalidate=86400, stale-if-error=604800 |
| cf-cache-status | null |
| content-encoding | null |
| content-length | 201668 |
@
@@ -12,7 +12,7 @@ <atom:link href="https://pytorch.org/blog/feed/" rel="self" type="application/rss+xml" /> <link>https://pytorch.org</link> <description></description>-
<lastBuildDate>Fri, 02 Oct 2026 19:57:16 +0000</lastBuildDate>+
<lastBuildDate>Mon, 05 Oct 2026 13:44:04 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod>@
@@ -28,6 +28,171 @@ <height>32</height></image> <item>+
<title>Evolution of the PyTorch Media Processing Landscape</title>+
<link>https://pytorch.org/blog/evolution-of-the-pytorch-media-processing-landscape/</link>+
+
<dc:creator><![CDATA[Nicolas Hug (Meta), Scott Schneider (Meta)]]></dc:creator>+
<pubDate>Mon, 05 Oct 2026 20:45:50 +0000</pubDate>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=172290</guid>+
+
<description><![CDATA[TL;DR If you need to decode or encode media, whether it’s images, video, or audio, use TorchCodec. If you need to transform media, use TorchVision for images and video, and...]]></description>+
<content:encoded><![CDATA[<h3>TL;DR</h3>+
<p>If you need to <em>decode or encode</em> media, whether it’s images, video, or audio, use <a href="https://github.com/pytorch/torchcodec"><strong>TorchCodec</strong></a>. If you need to <em>transform</em> media, use <a href="https://github.com/pytorch/vision"><strong>TorchVision</strong></a…+
<p>Images, video and audio are now central to a lot of model development, from vision-language models to diffusion models that generate images and video. Training these models means decoding media into tensors and transforming them, and generative models then need to encode their outputs back into m…+
<h2>Before and now</h2>+
<p><!-- IMAGE PLACEHOLDER: image1 (Before/now diagram of the media stack) - download from the Doc and upload via the WP media library --></p>+
<h3><img fetchpriority="high" decoding="async" class="alignnone wp-image-172323 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/media_libraries_2026-ezgif.com-svg-to-png-converter.png" alt="" width="960" height="560" srcset="https://pytorch.org/wp-content/uploads/2026/10/media_librari…+
<h3>TorchCodec: the one place for decoding and encoding</h3>+
<p>A few years ago, media decoding and encoding capabilities were scattered and partially duplicated across TorchVision and TorchAudio, with each one building its own stack. TorchVision had multiple entry points (<code>io.read_video()</code> and <code>io.VideoReader()</code>), spread across three ba…+
<p>In response to this, we have consolidated all decoding and encoding capabilities for media to live in TorchCodec: images, video, and audio, on CPU and CUDA. The APIs that used to live in TorchVision or TorchAudio are deprecated, or removed. We did this for three reasons:</p>+
<p><strong>One centralized location for users.</strong> With all media I/O in TorchCodec, users no longer have to work out which library happens to offer the decoder they need, and they get a library designed as a whole rather than a set of independent APIs.</p>+
<p><strong>Performance optimizations.</strong> A single centralized place also means that our development effort goes into one library, instead of being spread across many. When decoding and encoding lived in several implementations, it was never obvious which one deserved the optimization work. Now…+
<p><strong>Maintenance and dependencies.</strong> Media I/O involves a lot of dependencies: six major FFmpeg versions (4 through 9 at the time of writing), NVIDIA’s codec SDK for GPU decoding, and one library per image format (<code>libjpeg</code>, <code>libpng</code>, etc.), each with its own…+
<h3>TorchVision and TorchAudio: now focused on transforming media</h3>+
<p>TorchVision and TorchAudio supported a broad product surface, ranging across models, datasets, pipelines, transforms, and the I/O above. Today, both libraries are focused around what they do best: their transforms. The rest is not under active development.</p>+
<p>The driving factor here was the observation that the transforms in TorchVision and TorchAudio are the most active usage areas in the community. In contrast, models, datasets, and pipelines have lagged in usage, as great alternatives have risen, such as the HuggingFace libraries. Narrowing the sco…+
<p>This transition was particularly disruptive for TorchAudio: a lot of APIs were deprecated and eventually removed. We appreciate the community’s patience and feedback, which allowed us to reconsider our initial plan. Several popular APIs we had originally slated for removal <a href="https://…+
<h2>How they fit together</h2>+
<p>In practice the split is simple: TorchCodec turns media files or encoded bytes into tensors (decoding), and tensors back into media files (encoding). TorchVision and TorchAudio transform the tensors in between.</p>+
<p>For video, with TorchCodec’s <a href="https://meta-pytorch.org/torchcodec/main/generated/torchcodec.decoders.VideoDecoder.html">VideoDecoder </a>and TorchVision’s <a href="https://docs.pytorch.org/vision/stable/transforms.html">v2 transforms</a>:</p>+
<pre><code class="language-python">import torch
+
from torchcodec.decoders import VideoDecoder
+
from torchvision.transforms import v2
+
+
clip = VideoDecoder("video.mp4")[10:20] # FrameBatch, uint8, (10, C, H, W)
+
+
transform = v2.Compose([
+
v2.RandomResizedCrop(224),
+
v2.ToDtype(torch.float32, scale=True),
+
])
+
batch = transform(clip.data)
+
</code></pre>+
<p>For audio, with TorchCodec’s <a href="https://meta-pytorch.org/torchcodec/main/generated/torchcodec.decoders.AudioDecoder.html">AudioDecoder</a> and TorchAudio’s <a href="https://docs.pytorch.org/audio/stable/transforms.html">transforms</a>:</p>+
<pre><code class="language-python">from torchcodec.decoders import AudioDecoder
+
from torchaudio.transforms import MelSpectrogram
+
+
samples = AudioDecoder("audio.mp3").get_all_samples() # AudioSamples
+
+
mel = MelSpectrogram(sample_rate=samples.sample_rate)(samples.data)
+
</code></pre>+
<p>Images work the same way: <a href="https://meta-pytorch.org/torchcodec/main/generated/torchcodec.decoders.decode_image.html#torchcodec.decoders.decode_image"><code>decode_image</code></a> and friends return plain tensors that go straight into <code>torchvision.transforms.v2</code>. Encoding goes …+
<p>In terms of releases, all three libraries are now <a href="https://docs.pytorch.org/docs/2.14/notes/libtorch_stable_abi.html">ABI stable</a>. This means that a given version of TorchCodec, TorchVision or TorchAudio isn’t tied to one single version of PyTorch anymore: it keeps working with t…+
<p>If you’re still using the decoding or encoding APIs in TorchVision or TorchAudio, now is a good time to move to TorchCodec: start with our <a href="https://meta-pytorch.org/torchcodec/main/generated_examples/migration/torchvision_migration.html">migration guide</a>, and let us know on <a hr…+
<p>There is a lot more to say about the media processing libraries, and we’ll do it in a follow-up post: what’s new in each library, and how to get the most performance out of a full decoding and transform pipeline. Stay tuned!</p>+
]]></content:encoded>+
+
+
+
</item>+
<item>+
<title>PyTorch Hardware Enablement: Updates from the Accelerator Integration Working Group</title>+
<link>https://pytorch.org/blog/pytorch-hardware-enablement-updates-from-the-acceleration-integration-working-group/</link>+
+
<dc:creator><![CDATA[PyTorch TAC Accelerator Integration Working Group: Anisha Kushwaha, Atharva Kshirsagar, Jiahao Chen, Jiahao Tan, Jiawei Li, Jewel K. M., Mansi Agarwal, Parshant Sharma, Riya Punia, Subin George, Tanmay Kumar, Vishal Goyal, Guangye Yu, Zesheng Zong]]></dc:creator>+
<pubDate>Mon, 05 Oct 2026 13:12:59 +0000</pubDate>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=171617</guid>+
+
<description><![CDATA[TL;DR The Accelerator Integration Working Group plays a vital role in standardizing how new hardware architectures connect to the open source AI ecosystem. As compute platforms diversify across cloud, edge,...]]></description>+
<content:encoded><![CDATA[<h3>TL;DR</h3>+
<p>The Accelerator Integration Working Group plays a vital role in standardizing how new hardware architectures connect to the open source AI ecosystem. As compute platforms diversify across cloud, edge, and specialized silicon, the Accelerator Integration Working Group establishes clear, vendor-neu…+
<h2>Introduction</h2>+
<p>The PyTorch hardware ecosystem keeps diversifying, with a growing range of AI compute platforms adopting native framework integration. This expanding adoption creates a shared community challenge: streamlining integration paths, lowering complexity, standardizing integration mechanisms, and estab…+
<p>That is the core mission of the <a href="https://github.com/pytorch-fdn/accelerator-integration-wg">Accelerator Integration Working Group</a>. With a long‑term vision for a scalable, inclusive PyTorch hardware ecosystem, cross‑community contributors work on upstream framework improvements, shared…+
<p>In <a href="https://docs.google.com/document/d/1O5YBzMqH0kJ5Xe7bkxs38utXF2ksBietPC5sPffykc4/edit?usp=sharing">2026 H1</a>, we made tangible progress across multiple core workstreams: refactored test suites for broader cross‑backend reuse, the Cross‑Repository CI Relay (CRCR) mechanism, profiling …+
<h2>Cross-Repository CI Relay</h2>+
<p><em>Authors: Subin George, Jiahao Chen, Jiahao Tan, Jiawei Li, Jewel K. M.</em></p>+
<h3>Goals</h3>+
<p>The PyTorch community has introduced <a href="https://github.com/pytorch/pytorch/issues/175022">Cross-Repository CI Relay (CRCR)</a>, a new system designed to close a long-standing visibility gap in its CI ecosystem. PyTorch sits at the center of a large network of dependent projects – acce…+
<h3>Benefits</h3>+
<p>CRCR solves this with a fully automated pipeline. When a PR is opened or a commit is pushed to pytorch/pytorch, a webhook triggers dispatch events to all registered downstream repositories in parallel.</p>+
<p><!-- IMAGE PLACEHOLDER: image1 - Figure 1: CRCR Architecture --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172273 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/figure1.png" alt="" width="1048" height="848" srcset="https://pytorch.org/wp-content/uploads/2026/10/figure1.png 1048w, https://pytorch.org/wp-content/uploads/2026/10/figure1-…+
<p><em>Figure 1: CRCR Architecture</em></p>+
<p>Each repo runs its own CI workflow and reports status back via an authenticated callback, using a GitHub OIDC token that cryptographically verifies the calling repository’s identity. Results then flow into the PyTorch CI HUD (<a href="http://hud.pytorch.org/crcr">hud.pytorch.org/crcr</a>) w…+
<p><!-- IMAGE PLACEHOLDER: image2 - Figure 2: CRCR Hud Dashboard --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172274 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/figure2.png" alt="" width="1416" height="388" srcset="https://pytorch.org/wp-content/uploads/2026/10/figure2.png 1416w, https://pytorch.org/wp-content/uploads/2026/10/figure2-…+
<p><em>Figure 2: CRCR Hud Dashboard</em></p>+
<p>The system uses a tiered <a href="https://github.com/pytorch/pytorch/blob/main/.github/allowlist.yml">allowlist</a> with four participation levels (<a href="https://github.com/pytorch/pytorch/issues/175022">L1–L4</a>), letting downstream repos progress from simple dispatch notifications, to full …+
<h2>Test Refactoring</h2>+
<p><em>Authors: Riya Punia, Tanmay Kumar, Jiahao Chen, Jiahao Tan, Jiawei Li</em></p>+
<h3>Goals</h3>+
<p>Validating a new accelerator backend against PyTorch’s existing test suite is one of the most effective ways to ensure implementation correctness and catch upstream changes early. PyTorch maintains over 600,000 test cases covering operators, autograd, profiling, distributed training, and mo…+
<p>In H1 2026, the working group launched a systematic effort to decouple PyTorch’s test suite from specific hardware. This work is tracked through a <a href="https://github.com/pytorch/pytorch/issues/185590">central tracking issue</a>, an <a href="https://github.com/pytorch/pytorch/issues/174…+
<p><!-- IMAGE PLACEHOLDER: image3 - Figure 3: Test Refactoring Architecture --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172275 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/figure3.png" alt="" width="1906" height="1188" srcset="https://pytorch.org/wp-content/uploads/2026/10/figure3.png 1906w, https://pytorch.org/wp-content/uploads/2026/10/figure3…+
<p><em>Figure 3: Test Refactoring Architecture</em></p>+
<p>The effort has three dimensions:</p>+
<p><strong>Device-agnostic test migration:</strong> Replacing hardcoded device references (device=”cuda”, torch.cuda.synchronize(), @onlyCUDA) with parameterize equivalents (device=device, getattr(torch, device.type), instantiate_device_type_tests()). Each test class is classified as acc…+
<p><strong>Hardware classification metadata:</strong> Building on the <a href="https://github.com/pytorch/pytorch/issues/185142">test class classification</a> proposed by <a href="https://github.com/alband">Alban</a>, the group implemented a hw_classification class attribute (GENERIC, DEVICE_GENERIC…+
<p><strong>CI guardrails:</strong> A HW_CLASSIFICATION <a href="https://github.com/pytorch/pytorch/pull/190173">linter</a> enforces that every new test class declares its classification. Existing unclassified files (1,191 at launch) are allowlisted and being driven to zero incrementally —matching th…+
<h3>Benefits</h3>+
<p>For accelerator developers, these changes fundamentally shift the onboarding experience. Instead of patching numerous test files to adapt them for their hardware, vendors can rely on instantiate_device_type_tests() to automatically generate test variants for their backend. Tests that were previou…+
<p>For the PyTorch project itself, the hw_classification system provides the infrastructure foundation, which proposes class-level hardware and frequency metadata to reduce double-testing, improve CI observability, and rationalize test scheduling. The refactoring work done in H1 – migrating te…+
For the broader ecosystem, the multi-level test skipping mechanism (by feature, class, test, and operator) combined with hardware classification gives backend vendors fine-grained control over which tests to run and which to skip based on their implementation maturity – without modifying upstr…+
<h3>Next Steps</h3>+
<p>In H2, the priority is driving the unclassified allowlist toward zero, extending device-agnostic coverage to distributed, JIT, and autograd modules, and wiring hw_classification into CI scheduling for automatic test selection per runner. The group is also exploring a declarative per-operator capa…+
<h2>Profiling</h2>+
<p><em>Authors: Vishal Goyal, Anisha Kushwaha</em></p>+
<h3>Goals</h3>+
<p>Integrating a new hardware accelerator with PyTorch profiling requires more than operator-level timing. Vendors need a clear contract for kernel-level timelines, memory and runtime events, and CPU–device correlation in <a href="https://perfetto.dev/docs/">Chrome/Perfetto traces</a>. Without a wor…+
<p>In H1 2026, profiling followed OpenReg’s usual minimal, stub-like approach: not a production profiler, but just enough to demonstrate the integration mechanics. Building on the PyTorch core registration API (<a href="https://github.com/pytorch/pytorch/pull/172154">REGISTER_PRIVATEUSE1_PROFI…+
<h3>Benefits</h3>+
<p>For accelerator developers, the OpenReg profiling path provides a concrete starting point for kernel-level bring-up. Instead of reverse-engineering CUPTI-style plugins inside Kineto or settling for coarse ProfilerStubs timing, new accelerator teams can follow a single, minimal stub implementation…+
<p>For the broader PyTorch ecosystem, this turns PrivateUse1 profiling into a validated integration surface rather than a per-vendor addition to Kineto’s source tree. End-to-end torch.profiler tests catch registration, session-lifecycle, and fallback regressions early, and the stub reference stays a…+
<h2>OpenReg</h2>+
<p><em>Authors: Mansi Agarwal, Jiahao Chen, Jiahao Tan</em></p>+
<h3>Goals</h3>+
<p>Integrating a new hardware accelerator with PyTorch requires implementing a wide surface of functionality – device registration, operator dispatch, stream and event management, autograd integration, and more. Without a working reference, backend developers often resort to reverse-engineerin…+
<p><a href="https://github.com/pytorch/pytorch/tree/main/test/cpp_extensions/open_registration_extension/torch_openreg">OpenReg</a> is PyTorch’s in-tree reference backend for <a href="https://docs.pytorch.org/tutorials/advanced/privateuseone.html">PrivateUse1</a>-based accelerator integration.…+
<h3>Benefits</h3>+
<p>For accelerator developers, OpenReg provides a concrete starting point for bring-up work. Instead of piecing together integration patterns from scattered examples and production backends, new accelerator teams can follow a single, minimal implementation that demonstrates how an out-of-tree backen…+
<p>For the broader PyTorch ecosystem, OpenReg serves as a regression guard. Exercised through CI, it validates that PrivateUse1 integration mechanisms remain stable as PyTorch evolves, catching breakages before they reach downstream hardware vendors. The tighter alignment between OpenReg and the acc…+
<h3>Next Steps</h3>+
<p>In H2, OpenReg will continue to expand into additional PyTorch subsystems, while validating that PrivateUse1 integration paths remain stable across releases.</p>+
<p><!-- IMAGE PLACEHOLDER: image4 - Figure 4: OpenReg Integration Architecture --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172278 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/figure4.png" alt="" width="957" height="638" srcset="https://pytorch.org/wp-content/uploads/2026/10/figure4.png 957w, https://pytorch.org/wp-content/uploads/2026/10/figure4-30…+
<p><em>Figure 4: OpenReg Integration Architecture</em></p>+
<h2>Distributed Support</h2>+
<p><em>Authors: Mansi Agarwal, Atharva Kshirsagar</em></p>+
<h3>Goals</h3>+
<p>Distributed training is a core requirement for large-scale model development, but integrating a new accelerator with PyTorch’s distributed stack requires navigating a broad surface area, including custom transport, <a href="https://docs.pytorch.org/docs/stable/distributed.html#backends">ProcessGr…+
<p>In 2026 H1, the working group began addressing this gap by building <a href="https://github.com/pytorch/pytorch/tree/main/test/cpp_extensions/open_registration_extension/torch_openreg/csrc/distributed/c10d">OCCL</a> (OpenReg Collective Communications Library), a minimal reference implementation o…+
<p>This effort also helped improve the surrounding integration experience. Alongside OCCL, the contributors from working group added the <a href="https://docs.pytorch.org/docs/main/accelerator/distributed.html">documentation</a> to the <a href="https://docs.pytorch.org/docs/main/accelerator/index.ht…+
<h3>Benefits</h3>+
<p><!-- IMAGE PLACEHOLDER: image5 - Figure 5: Distributed OCCL Architecture --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172279 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/figure5.png" alt="" width="951" height="656" srcset="https://pytorch.org/wp-content/uploads/2026/10/figure5.png 951w, https://pytorch.org/wp-content/uploads/2026/10/figure5-30…+
<p><em>Figure 5: Distributed OCCL Architecture</em></p>+
<p>For accelerator developers, OCCL provides a much clearer starting point for <a href="https://docs.pytorch.org/docs/2.14/accelerator/distributed.html">adding distributed support</a> to a new backend. Until now, vendors often had to study production backends such as NCCL to infer the c10d integrati…+
<p>OCCL makes that contract easier to understand by presenting a standalone, hardware-independent path through the core pieces of distributed integration. Backend authors can follow a simpler model for how ProcessGroup registration, collective execution, and completion semantics fit together inside …+
<h2>Compile Backend Support</h2>+
<p><em>Authors: Parshant Sharma</em></p>+
<p>Compiler integration is one of the important pieces for any accelerator that wants to work well within PyTorch. Integrating a new accelerator with <code>torch.compile</code> requires navigating a broad surface area, including graph capture and device management in Dynamo, and optionally schedulin…+
<p>Rather than targeting production performance, the work is designed to make the Dynamo and Inductor integration points explicit. Using <a href="https://github.com/pytorch/pytorch/tree/main/test/cpp_extensions/open_registration_extension">OpenReg</a>, it shows how a device backend can register with…+
<p>Beyond the code itself, this work helped surface gaps in the existing documentation and developer experience. The working group contributed a <a href="https://github.com/pytorch/pytorch/pull/185700">compiler integration guide</a> to PyTorch’s accelerator docs, covering both Dynamo and Induc…+
<h3>Benefits</h3>+
<p>This work provides a much clearer starting point for adding compiler support to a new backend. Until now, vendors often had to study CUDA and Triton codepaths to infer the registration contracts. While those backends are essential in production, they are tightly coupled to hardware-specific sched…+
<p>Through the <a href="https://github.com/pytorch/pytorch/pull/181254">Dynamo backend</a> and <a href="https://github.com/pytorch/pytorch/pull/185486">Inductor integration</a>, backend authors can see how Dynamo graph capture, Inductor scheduling, wrapper codegen, and device operation overrides fit…+
<h2>PyTorch Additional Platform Page</h2>+
<p><em>Authors: Jiahao Chen, Jiawei Li</em></p>+
<h3>Goals</h3>+
<p>For a long time, users of the PyTorch community were missing an official channel to get clear information regarding PyTorch support for a new backend. The <a href="https://pytorch.org/get-started/additional-platforms/">PyTorch Additional Platforms page</a> is now the official, foundation-governed…+
<p>Applying is deliberately simple: open an <a href="https://github.com/pytorch-fdn/additional-compute-platforms/issues/new?template=application.yml">Additional Compute Platform Application</a> issue, provide evidence (links, docs, CI dashboards, security policy etc.) for each <a href="https://githu…+
<p>In H2, we’re actively encouraging more accelerator vendors, especially those who’ve already done the integration work described elsewhere in this recap, to apply and get their platform in front of the full PyTorch user base. Full details on requirements and the review process are in t…+
<h2>Conclusion</h2>+
<p>The progress achieved by the Accelerator Integration Working Group in H1 2026 demonstrates the power of vendor-neutral collaboration in open source AI. <span style="font-weight: 400;">By establishing standardized testing, profiling, compiler support, and distributed execution pathways, the commun…+
<p><span style="font-weight: 400;">To explore the ongoing workstreams or get involved in defining next-generation hardware enablement for PyTorch, visit the Accelerator Integration Working Group repository at</span><a href="https://github.com/pytorch-fdn/accelerator-integration-wg?utm_source=gemini"…+
<h2><span style="font-weight: 400;">Acknowledgements</span></h2>+
<p><span style="font-weight: 400;">We sincerely appreciate the generous support and guidance from PyTorch maintainers and community members throughout 2026H1. Special thanks to @alband, @afrittoli, @atalman, @jbschlosser, @marco-s, @matthew-d-white,@mikaylagawarecki, @ZainRizvi, @zxiiro … for their …+
<p><span style="font-weight: 400;">Find more information about Accelerator Integration Working Group: </span><a href="https://github.com/pytorch-fdn/accelerator-integration-wg"><span style="font-weight: 400;">https://github.com/pytorch-fdn/accelerator-integration-wg</span></a></p>+
<p><span style="font-weight: 400;">PyTorch TAC Working Groups: </span><a href="https://pytorch.org/working-groups/"><span style="font-weight: 400;">https://pytorch.org/working-groups/</span></a></p>+
]]></content:encoded>+
+
+
+
</item>+
<item> <title>Building a High-Performance and Portable vLLM Linear Backend with Helion</title> <link>https://pytorch.org/blog/building-a-high-performance-and-portable-vllm-linear-backend-with-helion/</link> @
@@ -58,7 +223,7 @@<p><b>Maintenance overhead</b><span style="font-weight: 400;">: Shipping pre-tuned configs for popular models creates an ongoing upstream maintenance burden. Large config files are difficult to maintain and impractical to validate exhaustively through unit tests and CI. </span></p><h3><span style="font-weight: 400;">The Tradeoff Triangle</span></h3><p><span style="font-weight: 400;">Fine-grained kernel tuning with Helion presents a tradeoff among </span><b>performance</b><span style="font-weight: 400;">, </span><b>usability</b><span style="font-weight: 400;">, and </span><b>maintainability</b><span style="font-weight: 400;">. This tradeoff is …-
<p><img fetchpriority="high" decoding="async" class="aligncenter wp-image-171928 " src="https://pytorch.org/wp-content/uploads/2026/10/1.png" alt="" width="587" height="587" srcset="https://pytorch.org/wp-content/uploads/2026/10/1.png 1254w, https://pytorch.org/wp-content/uploads/2026/10/1-300x300.p…+
<p><img decoding="async" class="aligncenter wp-image-171928 " src="https://pytorch.org/wp-content/uploads/2026/10/1.png" alt="" width="587" height="587" srcset="https://pytorch.org/wp-content/uploads/2026/10/1.png 1254w, https://pytorch.org/wp-content/uploads/2026/10/1-300x300.png 300w, https://pyto…<p style="text-align: center;"><i><span style="font-weight: 400;">Fig. 1: The Performance-Usability-Maintainability tradeoff triangle</span></i></p><p><span style="font-weight: 400;">Higher performance generally requires more fine-grained config tuning. This either increases client-side autotuning overhead or requires maintainers to provide and maintain more pre-tuned configs upstream. The goal, therefore, is to strike the right balance for the…<h2><span style="font-weight: 400;">Helion Linear Backend </span></h2>@
@@ -1012,177 +1177,8 @@ echo "Resolved pytorch/pytorch@nightly -> source ${SOURCE_SHA}" -
</item>-
<item>-
<title>How Shopify built a continual learning loop with PyTorch and vLLM</title>-
<link>https://pytorch.org/blog/how-shopify-built-a-continual-learning-loop-with-pytorch-and-vllm/</link>-
-
<dc:creator><![CDATA[Cody Mazza-Anthony, Sr. Staff Machine Learning Engineer, @cmazzaanthony & Andrew McNamara, VP Machine Learning, @drewch]]></dc:creator>-
<pubDate>Tue, 22 Sep 2026 13:45:18 +0000</pubDate>-
<category><![CDATA[Blog]]></category>-
<category><![CDATA[Case Studies]]></category>-
<guid isPermaLink="false">https://pytorch.org/?p=169848</guid>-
-
<description><![CDATA[TL;DR: This case study explores how Shopify compresses production failures into model weights every day, beats frontier-model quality, and cuts serving costs 96% by building a continual learning loop with...]]></description>-
<content:encoded><![CDATA[<p><span style="font-weight: 400;"><img decoding="async" class="alignnone size-large wp-image-169894" src="https://pytorch.org/wp-content/uploads/2026/09/How-Shopify-built-a-continual-learning-loop-with-PyTorch-vLLM-1024x536.png" alt="How Shopify built a continual…-
<p>Frontier models are often the fastest way to launch a new AI product. A small team can get something useful in front of users quickly and learn from real-world usage. But as usage grows, the economics change: frontier models can be too slow and expensive to serve every request at scale.</p>-
<p>Frontier models are general-purpose, not tailored to your product. More importantly, they do not learn from production on their own. A user correction, rejected output, or recurring failure does not make the next response better. But each failure is hard-won knowledge about your product, and cont…-
<p>Additionally, the deployed frontier model weights are frozen. It has no mechanism for internalizing what production teaches it. Instead, improvements accumulate in the discrete artifacts around it: prompt edits, retrieval examples, routing rules, and harness code. Production knowledge piles up in…-
<p>Shopify’s GraphQL agent is our clearest example of that loop running in production. That flywheel delivers higher quality than frontier models while reducing latency and cutting costs by 96%.</p>-
<h2><img decoding="async" class="alignnone wp-image-169864 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Flywheel-Optimization-Process-e1790031034727.png" alt="Flywheel Optimization Process" width="920" height="387" srcset="https://pytorch.org/wp-content/uploads/2026/09/Flywheel-Opt…-
<p>Defining quality is the most important step in the loop—and the one that teams most often rush. It begins as a specification of what good looks like and becomes the reward signal that drives learning. When you get the evals or specs wrong, everything downstream optimizes the wrong behavior.</p>-
<p>Quality starts with a rubric, and that rubric turns your product requirements into a few scored criteria: completeness, execution, response quality, and safety. Each one has concrete anchors for what every score means. Think of it as your product team’s definition of good and bad. It’s the …-
<p>Once you’re happy with the rubric, have your two best annotators/product experts blindly annotate 25 random samples and record their inter-annotator agreement. We use <a href="https://en.wikipedia.org/wiki/Cohen%27s_kappa">Cohen’s kappa</a>, which measures agreement above chance. If it’s ve…-
<p>That agreement is the judge’s ceiling: even expert annotators do not agree 100% of the time, because some conversations are genuinely ambiguous. The goal is not a judge that is “perfect,” but one that matches humans about as well as humans match each other.</p>-
<p>When you collect annotations, push for detail. A score and one sentence is not enough for the calibration algorithms to learn from. You want the <em>why</em> behind every score, and that reasoning is gold.</p>-
<h2><img decoding="async" class="alignnone size-large wp-image-169867" src="https://pytorch.org/wp-content/uploads/2026/09/Judges-1024x604.png" alt="Judges" width="1024" height="604" srcset="https://pytorch.org/wp-content/uploads/2026/09/Judges-1024x604.png 1024w, https://pytorch.org/wp-content/uplo…-
<p>The rubric is just the starting point. It’s the judge’s first prompt, but it hasn’t learned anything from your ground truth yet. Calibration will take that rubric and turn it into a judge that can be run on infinite production datapoints.</p>-
<p>We’re big fans of DSPy for this, and we calibrate with reflection-based optimizers like GEPA and Agentic Context Engineering [1,2]. GEPA evolves the prompt by reflecting on natural-language failure traces and keeps a Pareto frontier of candidates instead of greedily choosing a single winner. ACE …-
<p>The judge is your offline metric, the thing you optimize against before you ship. But it’s only a proxy, so you need to establish that it reflects performance on real traffic and is aligned with your online metric. Backtest it against previous A/B tests: can it recover the direction of known wins…-
<p>Then run targeted degradation tests, either offline or on a carefully controlled slice of traffic. Deliberately make one behavior worse and confirm that the corresponding criterion responds. If the system stops trying to fulfil the user’s goal, for example, the goal-fulfilment score should fall s…-
<p>Keep each judge small and targeted rather than cramming all of your product’s behavior into one. Focused judges make these tests easier to interpret and the resulting metrics easier to trust. You can always add more.</p>-
<h2>Improve the frontier baseline with autoresearch</h2>-
<p>Now that we have a reliable judge, we can use it to improve the initial frontier-powered product we launched, the baseline built to get in front of users quickly. At this stage, we push that system as far as it will go without touching the weights. Every improvement lands in prompts, tool definit…-
<p>But improving this baseline system is a different problem from building the judge. It’s already a production application, with dynamically assembled prompts, custom control loops, and bespoke orchestration spread across a large codebase. No single prompt determines its behavior, so prompt tuning …-
<p>So we treat it as an <a href="https://shopify.engineering/autoresearch">autoresearch</a> problem, in the spirit of <a href="https://github.com/karpathy/autoresearch">Karpathy’s recent project</a>: an agent proposes a change to a prompt, a tool definition, or the harness; evaluates it agains…-
<p>We configure the whole thing in one readable markdown file: where to get data, which directories the agent may edit, the judge as the metric, the optimizer to use, and the propose-evaluate-keep-or-discard loop.</p>-
<h2><img decoding="async" class="alignnone size-large wp-image-169868" src="https://pytorch.org/wp-content/uploads/2026/09/program.md_-1024x902.png" alt="program.md" width="1024" height="902" srcset="https://pytorch.org/wp-content/uploads/2026/09/program.md_-1024x902.png 1024w, https://pytorch.org/w…-
<p>Once harness improvements plateau, we begin optimizing in parameter space by mining anonymized production traffic for hard negatives: conversations the judge correctly scores low and that expose where the model is weakest.</p>-
<p>Across millions of diverse merchants, real traffic produces a continual stream of difficult cases: partial context, ambiguous requests, business-specific workflows, tool failures, and many ways of expressing the same intent. In a traditional workflow, each failure becomes a bug report or Slack th…-
<p> </p>-
<p><img decoding="async" class="alignnone size-large wp-image-169873" src="https://pytorch.org/wp-content/uploads/2026/09/Failed-convo-flow-1024x406.png" alt="Failed convo flow" width="1024" height="406" srcset="https://pytorch.org/wp-content/uploads/2026/09/Failed-convo-flow-1024x406.png 1024w, htt…-
<p>Training proceeds in two stages. First, we distill the healed trajectories into a smaller model through supervised fine-tuning. We train on the complete trajectories—including the reasoning that produced them—not just their final answers (see <a href="https://toloka.ai/blog/fine-tuning-for-agenti…-
<p>Second, we apply GRPO, using the calibrated judge as the reward signal. For each prompt, the model samples a group of responses, the judges score them, and GRPO reinforces the responses that perform best. Supervised fine-tuning teaches the model successful trajectories to imitate; GRPO optimizes …-
<p>The self-healing pipeline runs daily, continually adding new trajectories to the training corpus. On the same cadence, we run a full-parameter fine-tune over the accumulated data and then repeat GRPO. Training on both new and previous trajectories limits drift and catastrophic forgetting across c…-
<blockquote><p>We use PyTorch to distribute training across GPUs with tensor, context and data parallelism, making full parameter fine-tuning practical at scale.</p></blockquote>-
<h2>Compress the prompt to serve it faster</h2>-
<p>A better model still has to run, and an agent’s system prompt is long and static. Attention scales with sequence length, so every generated token attends over that entire prefix. A long prompt is a fixed tax on latency and serving cost, paid on every request.</p>-
<p>Gist compression removes most of that tax. We run the same model two ways: a teacher with the full system prompt, and a student with a short sequence of learned gist tokens instead. We built a custom PyTorch trainer that learns the gist token embeddings by matching the teacher’s output distributi…-
<h2><img decoding="async" class="alignnone size-full wp-image-169874" src="https://pytorch.org/wp-content/uploads/2026/09/system-prompt.png" alt="system prompt" width="936" height="176" srcset="https://pytorch.org/wp-content/uploads/2026/09/system-prompt.png 936w, https://pytorch.org/wp-content/uplo…-
<p><img decoding="async" class="alignnone size-large wp-image-169875" src="https://pytorch.org/wp-content/uploads/2026/09/Merchant-Sidekick-convo-1024x552.png" alt="Merchant Sidekick convo" width="1024" height="552" srcset="https://pytorch.org/wp-content/uploads/2026/09/Merchant-Sidekick-convo-1024x…-
<p><img decoding="async" class="alignnone size-large wp-image-169880" src="https://pytorch.org/wp-content/uploads/2026/09/GraphQL-distillation-1024x544.png" alt="GraphQL distillation" width="1024" height="544" srcset="https://pytorch.org/wp-content/uploads/2026/09/GraphQL-distillation-1024x544.png 1…-
<p><strong>It made the model better.</strong> The self-healing pipeline turns low-scoring production conversations into successful trajectories, giving the model a continual stream of lessons drawn from real merchant needs. Together, SFT and RL enable the specialized model to surpass the frontier mo…-
<p><strong>It made the model far cheaper to serve. </strong>We serve our models through vLLM, an inference engine built on PyTorch. vLLM keeps our throughput high with continuous batching and works well for our tool-call heavy workloads. Serving this traffic on a frontier model could easily cost an …-
<p><strong>It made the model faster, and the gap grows under load.</strong> Gisting compressed the agent’s long, static system prompt from roughly 6,000 tokens down to about 1,500 learned gist tokens. In a load test at 350 requests per minute, time-to-first-token dropped about 19%, and end-to-…-
<p><strong>It freed up hardware. </strong>The same compression raises throughput: about 16% more requests per second and about 12% more output tokens per second on identical GPUs, which works out to roughly 14% fewer GPUs for the same traffic.</p>-
<h2>Beyond the harness: continual learning that compounds</h2>-
<p>Frontier models help you launch, and the first improvements live in the discrete artifacts around them: prompts, context, tool definitions, and control flow. Those changes strengthen the harness but leave the model unchanged.</p>-
<p>Continual learning goes further, translating those lessons into updates in the model’s continuous parameter space. Each cycle begins with a more capable model, not merely a more elaborate harness. That is how a smaller model becomes faster, cheaper, and better at your task than the frontier basel…-
<h2>References</h2>-
<ol>-
<li>Agrawal et al. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. arXiv:2507.19457</li>-
<li>Zhang et al. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. arXiv:2510.04618</li>-
<li>Hsieh et al. Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes. arXiv:2305.02301</li>-
<li>Shuttleworth et al. LoRA vs Full Fine-tuning: An Illusion of Equivalence. arXiv:2410.21228</li>-
<li>Wingate et al. Prompt Compression and Contrastive Conditioning for Controllability and Toxicity Reduction in Language Models. arXiv:2210.03162</li>-
<li>Mu et al. Learning to Compress Prompts with Gist Tokens. arXiv:2304.08467</li>-
</ol>-
<p><em>Originally published on the </em><a href="https://shopify.engineering/"><em>Shopify Engineering Blog</em></a><em>. This version has been adapted for the PyTorch community.</em></p>-
]]></content:encoded>-
-
-
-
</item>-
<item>-
<title>TinyTorch: Don’t Just Import PyTorch. Build It.</title>-
<link>https://pytorch.org/blog/tinytorch-dont-just-import-pytorch-build-it/</link>-
-
<dc:creator><![CDATA[Vijay Janapa Reddi, Harvard University and ETH Zurich · Andrea Mattia Garavagno, ETH Zurich]]></dc:creator>-
<pubDate>Mon, 21 Sep 2026 20:41:14 +0000</pubDate>-
<category><![CDATA[Blog]]></category>-
<guid isPermaLink="false">https://pytorch.org/?p=167810</guid>-
-
<description><![CDATA[A framework you write yourself, tensors through transformers TL;DR Every mature systems project eventually needs a teaching version. TinyTorch is a free, open-source curriculum where you build a working ML...]]></description>-
<content:encoded><![CDATA[<p><em>A framework you write yourself, tensors through transformers</em></p>-
<h2><img decoding="async" class="alignleft wp-image-169836 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/All-PyTorch-Blog-Social-Images-25.png" alt="" width="1920" height="1080" srcset="https://pytorch.org/wp-content/uploads/2026/09/All-PyTorch-Blog-Social-Images-25.png 1920w, https…-
<h2>TL;DR</h2>-
<p>Every mature systems project eventually needs a teaching version. TinyTorch is a free, open-source curriculum where you build a working ML framework from scratch, tensors through transformers, in pure Python, using PyTorch’s own API. Twenty modules. Runs on a laptop with 4 GB of RAM and no …-
<p>The rest of this post is why we think it needed to exist, and what six years of running it taught us that might be useful to anyone else doing open curriculum work.</p>-
<h2>Every Systems Field Eventually Builds Its Teaching Version</h2>-
<p>Unix got too big to hold in your head, so Andrew Tanenbaum wrote MINIX. Small enough that a student could actually finish it. It went on to shape a generation of systems engineers and famously inspired Linux.</p>-
<p>Compilers went the same way. LLVM and GCC are decades of excellent engineering and close to unreadable as a first text, so courses teach the Tiger compiler instead. MIT rewrote xv6 from x86 to RISC-V for the same reason, stripping out historical complexity to expose clean abstractions. Before all…-
<p>None of these were trying to replace the production system. They taught what production systems have to hide.</p>-
<p>PyTorch is at that point now, which is a compliment. There is genuinely good writing on its internals, most of it by the people who wrote them, and nearly all of it assumes you already think like a framework engineer. What has been missing is the rung below that. Something you build yourself, wit…-
<h2>Why This Matters to PyTorch, Not Only to Students</h2>-
<p>There is a version of this argument that only educators care about. This is not that version.</p>-
<p>Every framework runs on a small population of people who can reason about it from the inside. The ones who spot a memory leak in tensor caching, who know when gradient checkpointing is worth the recompute, who can look at a slow training run and name the bottleneck before they open a profiler.</p…-
<p>That population does not grow by itself. Right now most people get to PyTorch’s internals by accident, because something broke badly enough to force the trip. They learn the codebase under deadline pressure, from the outside in, in whatever order the bug happened to demand. Anyone who has o…-
<p>Building the thing yourself changes the arrival path, and it changes it permanently. Once you have implemented autograd you cannot unsee the computational graph. Once you have profiled your own memory allocation you cannot unknow the cost.</p>-
<p><!-- IMAGE PLACEHOLDER: Figure 1 - Side-by-side comparison graphic captioned "Building systems creates irreversible understanding." Left panel, "Traditional ML Education": a short torch.nn.Linear snippet with "Problem: You can't debug what you don't understand." Right panel, "TinyTorch: Build → U…-
<p><img decoding="async" class="alignnone wp-image-167831 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Figure1.png" alt="" width="1656" height="658" srcset="https://pytorch.org/wp-content/uploads/2026/09/Figure1.png 1656w, https://pytorch.org/wp-content/uploads/2026/09/Figure1-300x…-
<p><em><strong>Figure 1</strong>: The same linear layer, two ways. On the left, the framework call that works right up until it does not. On the right, the version you wrote and can open when something breaks. The difference is not syntax, it is whether the abstraction is a wall or a door.</em></p>-
<p>A student who has written <code>backward()</code> themselves, who allocated the momentum and variance buffers and watched Adam’s memory footprint triple, shows up to PyTorch’s real autograd with the mental model already loaded. They read the production code as a more sophisticated ver…-
<p>The part we did not anticipate is that companies want this too. Teams have used TinyTorch for new-hire onboarding as a two to three week intensive, as internal training spread across a quarter, and as targeted debugging workshops where somebody works through Module 06 on autograd or Module 12 on …-
<h2>What We Built</h2>-
<p>Twenty modules in four tiers, driven by a CLI called <code>tito</code>, delivered as Jupyter notebooks with the hard parts cut out for you to fill in. You need Python and to be comfortable with NumPy. You do not need a GPU, a cloud account, or any prior ML systems background.</p>-
<p><!-- IMAGE PLACEHOLDER: Figure 2 - Diagram of the four tiers of TinyTorch modules, each tier depending on the one below it. --></p>-
<p><img decoding="async" class="alignnone wp-image-167832 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Figure2.png" alt="" width="1656" height="528" srcset="https://pytorch.org/wp-content/uploads/2026/09/Figure2.png 1656w, https://pytorch.org/wp-content/uploads/2026/09/Figure2-300x…-
<p><em><strong>Figure 2</strong>: The four tiers. Each one depends on the tier below it, so you cannot skip ahead to optimization without having built the training loop you are optimizing. Foundation fits a half-semester module, all twenty fit a four-credit course, and self-paced learners take anywh…-
<p>Three design decisions do most of the pedagogical work.</p>-
<p><strong>Systems from day one.</strong> Module 01 ships a <code>memory_footprint()</code> method before it ships matrix multiplication. You learn that one batch of 32 ImageNet images costs 19 MB by computing it, not by reading it somewhere. Later you find out Adam needs roughly 3× the optimizer me…-
<p><strong>Progressive disclosure.</strong> The <code>Tensor</code> class stays clean through Module 05, with no gradient machinery cluttering up data layout and arithmetic. Then in Module 06 you implement <code>enable_autograd()</code>, which bolts <code>requires_grad</code>, <code>.grad</code>, an…-
<p>We went back and forth on this one. The tasteful way to do it is inheritance. What we shipped is runtime monkey-patching, which is going to offend somebody reading this. It won because it keeps one <code>Tensor</code> class across all twenty modules instead of two, and because the moment your ten…-
<p><!-- IMAGE PLACEHOLDER: Figure 3 - Screenshot of a Module 06 notebook cell showing the docstring and numbered steps for implementing the gradient of matrix multiplication, with the implementation left blank for the learner to fill in. --></p>-
<p><img decoding="async" class="alignnone wp-image-167835 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Figure3.png" alt="" width="1834" height="1284" srcset="https://pytorch.org/wp-content/uploads/2026/09/Figure3.png 1834w, https://pytorch.org/wp-content/uploads/2026/09/Figure3-300…-
<p><em><strong>Figure 3</strong>: The gradient of matrix multiplication, as a learner meets it in Module 06. The docstring gives you the mathematical rule and the numbered approach breaks it into steps, but the implementation is yours. All twenty modules look like this.</em></p>-
<p><strong>Build to validate.</strong> Six historical milestones prove your implementation works. Rosenblatt’s Perceptron in 1958, the XOR crisis in 1969, the backpropagation revival in 1986, the CNN breakthrough in 1998 where your network has to clear 75% on CIFAR-10, the transformer in 2017,…-
<p><!-- IMAGE PLACEHOLDER: Figure 4 - Timeline/ladder graphic of the six historical milestones (Perceptron 1958 through MLPerf-style benchmarking), each milestone unlocking once the modules beneath it work. --></p>-
<p><img decoding="async" class="alignnone wp-image-167836 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Figure4.png" alt="" width="1060" height="1115" srcset="https://pytorch.org/wp-content/uploads/2026/09/Figure4.png 1060w, https://pytorch.org/wp-content/uploads/2026/09/Figure4-285…-
<p><em><strong>Figure 4</strong>: The milestone ladder. A milestone unlocks only when the modules beneath it produce a working implementation, so the timeline doubles as a progress tracker and a correctness proof. Students recreate 67 years of ML history running nothing but code they wrote.</em></p>-
<p>Throughout, TinyTorch mirrors PyTorch’s API deliberately. The API is the transfer mechanism and it carries both ways. Somebody who builds <code>loss.backward()</code> here can open PyTorch’s version afterward and recognize the shape of it. A PyTorch developer can read our attention mo…-
<p><!-- IMAGE PLACEHOLDER: Figure 5 - Side-by-side code comparison of a TinyTorch training loop and the equivalent PyTorch training loop, showing only the imports differ. --></p>-
<p><img decoding="async" class="alignnone wp-image-167839 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Figure5.png" alt="" width="2034" height="566" srcset="https://pytorch.org/wp-content/uploads/2026/09/Figure5.png 2034w, https://pytorch.org/wp-content/uploads/2026/09/Figure5-300x…-
<p><em><strong>Figure 5</strong>: A TinyTorch training loop next to the equivalent PyTorch loop. The imports differ and almost nothing else does, which is the entire point. Nothing you learn here has to be unlearned later.</em></p>-
<p>Two things it is not. The resemblance stops at the API surface, so there is no dispatcher, no C++ or CUDA layer, no JIT, nothing distributed. And it is slow. Pure Python runs somewhere between 100 and 10,000 times slower than PyTorch, which we will come back to, because it turned out to matter.</…-
<h2>Where It Came From</h2>-
<p>TinyTorch did not start as a framework. It started as a course with a problem.</p>-
<p>CS 249r launched at Harvard in 2020, a graduate seminar on TinyML, and there was no textbook to assign. So the course notes became one. We put the book up as an open repository, students and educators started fixing examples and proposing chapters, and by 2024 it had outgrown its TinyML origins b…-
<p>Somewhere in there it became obvious a textbook was not enough. Reading about autograd and implementing autograd produce different kinds of knowledge, and only one of them survives contact with a production bug. So the project grew limbs. TinyTorch is the build limb. Around it now sit Marimo labs…-
<p><!-- IMAGE PLACEHOLDER: Figure 6 - Map graphic showing the TinyTorch community's 682 members clustered across 92 institutions worldwide. --></p>-
<p><img decoding="async" class="alignnone wp-image-167840 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Picture6.png" alt="" width="1783" height="1195" srcset="https://pytorch.org/wp-content/uploads/2026/09/Picture6.png 1783w, https://pytorch.org/wp-content/uploads/2026/09/Picture6-…-
<p><em><strong>Figure 6</strong>: The TinyTorch community map, 682 members across 92 institutions since it launched in December 2025. The clustering is the interesting part. Uptake is heaviest where GPU access is hardest, which is what the accessibility floor was for.</em></p>-
<p>Then there is the part that is slightly embarrassing to write down.</p>-
<p>For five years this was a slow project. Steady, word of mouth, a few hundred stars a year, the kind of thing you keep doing because the students in front of you need it. In August 2025 the repository had about 2,000 stars.</p>-
<p>In October, somebody we had never met posted it on X.</p>-
<p>Two weeks later we were at 5,000. By December we had passed 10,000, and today it is above 27,000, with 95 or more contributors and courses running at 50 or more universities. We would like to tell you we engineered that. We did not. One person with reach decided the work was worth sharing, and fi…-
<p><!-- IMAGE PLACEHOLDER: Figure 7 - Line chart of GitHub stars over time for the Machine Learning Systems repository, showing a flat multi-year plateau followed by a sharp cliff/spike starting around October 2025. --></p>-
<p><img decoding="async" class="alignnone wp-image-167845 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/Figure7.png" alt="" width="2147" height="1193" srcset="https://pytorch.org/wp-content/uploads/2026/09/Figure7.png 2147w, https://pytorch.org/wp-content/uploads/2026/09/Figure7-300…-
<p><em><strong>Figure 7</strong>: Stars on the Machine Learning Systems repository. The cliff is the part everyone looks at. The flat part is where the work happened.</em></p>-
<p>TinyTorch was built at Harvard as part of the <a href="https://mlsysbook.ai">Machine Learning Systems</a> project, and Andrea now maintains it from ETH Zurich, where both of us are currently based. Handing a curriculum to a second institution is the first honest test of whether you wrote it for a…-
<h2>What Transferred to Other Projects</h2>-
<p>These guidelines ask for lessons other institutions can apply, which is the right thing to ask for. Four of them.</p>-
<p><strong>Match the production API exactly.</strong> The highest-leverage decision we made was refusing to invent our own syntax. Every hour a learner spends translating between your teaching API and the real one is an hour that buys them nothing and costs you a fraction of them. Treat API compatib…-
<p><strong>Decide your hardware floor before your feature list.</strong> TinyTorch runs on a dual-core 2 GHz CPU with 4 GB of RAM and no network during training, because we ship two tiny offline datasets (about 1,000 grayscale digits and 350 conversational question-answer pairs, under 50 MB together…-
<p>The slowness we apologized for turned out to be the best accident in the project. When a student’s <code>Conv2d</code> takes 97 seconds on a batch that PyTorch clears in 10 milliseconds, the argument for vectorization stops being something they read and starts being something that happened …-
<p><strong>Instructor infrastructure is the actual bottleneck.</strong> Good content does not get adopted. Gradeable content gets adopted. We shipped NBGrader autograding with locked test cells and point allocations, an <code>INSTRUCTOR.md</code> covering setup and rubrics and the errors students ac…-
<p><strong>Plan succession before you need it.</strong> Academic open source dies when the PI changes focus. We are handling that in the open, with a maintenance commitment through 2027, a two-week pull request review target, and a governance transition across 2026 and 2027 that sets up an educator …-
<h2>What We Still Do Not Know</h2>-
<p>TinyTorch is in preview and aimed at classroom readiness for Fall 2026. The limits are worth more to you than a clean success story.</p>-
<p>We have not measured learning outcomes. The design leans on constructionism, cognitive apprenticeship, productive failure, threshold concepts, and five decades of evidence that build-it-yourself works in systems education. That is a strong prior. It is not a result. We do not have controlled data…-
<p>The scope is also narrower than the ambition. Single-node, CPU-only, so it teaches memory and compute well and teaches nothing about GPU kernels, distributed training, or gradient synchronization. Parallel data loading and GPU memory management are the one competency area we mapped out and then d…-
<h2>Where You Can Help</h2>-
<p>If you would rather just try it, installation is one line and everything runs locally.</p>-
<pre><code class="language-bash">curl -sSL mlsysbook.ai/tinytorch/install.sh | bash</code></pre>-
<p>Five openings, roughly in order of how much each would move things.</p>-
<ol>-
<li><strong>Review a module against real PyTorch semantics.</strong> If you work on PyTorch internals, an hour checking whether our autograd or our optimizer state handling or our KV cache teaches the right mental model is worth an enormous amount. Whatever TinyTorch teaches becomes what a cohort of…-
<li><strong>Pilot a tier and tell us what broke.</strong> Fall 2026 syllabi are mostly locked by now, so realistically that is Spring 2027 or Fall 2027, and shadowing the material this fall is a good way to decide. Foundation fits an undergraduate systems module, all twenty fit a semester, the Optim…-
<li><strong>Write the modules we cannot.</strong> Distributed training, GPU acceleration, parallel data loading. These need somebody who already teaches the material.</li>-
<li><strong>Localize it.</strong> The datasets are small and offline by design, which makes translating the conversational one into another language a weekend project with real reach.</li>-
<li><strong>Adopt it and say so.</strong> Putting your institution on the community map is what tells the next department this is a real option and not an experiment.</li>-
</ol>-
<p>Everything lives in the <a href="https://github.com/harvard-edge/cs249r_book">Machine Learning Systems repository</a>. Code is MIT, curriculum is CC BY-SA 4.0, so forking and adapting for your own institution is explicitly fine, and upstreaming what you fix is appreciated.</p>-
<p>All of it is free and it stays free. If the argument here landed and you want the cheapest way to act on it, star the repository. That number is not a scoreboard for us. It is what a department chair looks at when deciding whether an open curriculum is safe to build a course on, and what a funder…-
<h2>Join the PyTorch Academic OSPO Working Group</h2>-
<p>TinyTorch is one example of universities pushing the PyTorch ecosystem forward through open collaboration, and this post exists because of the PyTorch Academic OSPO Working Group, which has been helping us turn an enthusiastic pile of learners into something with actual governance. Several of the…-
<p>If you are interested in sharing academic projects, developing best practices, or connecting with people working where PyTorch meets academia, consider joining the <a href="https://github.com/pytorch-fdn/wg-ospo-and-academic-outreach">PyTorch Academic OSPO Working Group</a>. It welcomes researche…-
<p><img decoding="async" class="alignnone wp-image-167846 size-full" src="https://pytorch.org/wp-content/uploads/2026/09/ospoOutreach_horizontalColor-scaled.png" alt="" width="2560" height="1440" srcset="https://pytorch.org/wp-content/uploads/2026/09/ospoOutreach_horizontalColor-scaled.png 2560w, ht…-
]]></content:encoded>-
-
-
</item> </channel></rss>-
<!-- plugin=object-cache-pro client=phpredis metric#hits=4213 metric#misses=31 metric#hit-ratio=99.3 metric#bytes=1283237 metric#prefetches=245 metric#store-reads=36 metric#store-writes=6 metric#store-hits=254 metric#store-misses=15 metric#sql-queries=9 metric#ms-total=323.10 metric#ms-cache=16.43 m…+
<!-- plugin=object-cache-pro client=phpredis metric#hits=4106 metric#misses=31 metric#hit-ratio=99.3 metric#bytes=1246455 metric#prefetches=222 metric#store-reads=36 metric#store-writes=8 metric#store-hits=231 metric#store-misses=15 metric#sql-queries=12 metric#ms-total=560.29 metric#ms-cache=32.95 …129 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.