Change
7c0c275
7c0c275fefb304d488016bc7b92e7e3087c52347 · commit on GitHub
pytorch-blog-feed: changed (249928 bytes, HTTP 200)
raw/pytorch-blog-feed/response.xml modified
- Source
- pytorch-blog-feed
- Lines added
- +407
- Lines removed
- -217
- Stored bytes at this commit
- 249,928
- Timestamp
- origin
- Raw artifact at this commit
- raw/pytorch-blog-feed/response.xml
Recorded headers
| observed_at | 2026-10-09T06:08:55.223Z |
|---|---|
| origin_date | 2026-10-09T05:33:15.000Z |
| status | 200 |
| final URL | https://pytorch.org/blog/feed/ |
| etag | "570edd5a6478dafdd52cf679e888e7d7" |
| last-modified | Thu, 08 Oct 2026 19:36:43 GMT |
| date | Fri, 09 Oct 2026 06:08:55 GMT |
| age | 2140 |
| cache-control | public, max-age=60, s-maxage=43200, stale-while-revalidate=86400, stale-if-error=604800 |
| cf-cache-status | null |
| content-encoding | null |
| content-length | 249928 |
@
@@ -8,26 +8,393 @@ ><channel>-
<title>Blog – PyTorch</title>+
<title>Blog - PyTorch</title> <atom:link href="https://pytorch.org/blog/feed/" rel="self" type="application/rss+xml" />-
<link>https://pytorch.org</link>+
<link>https://pytorch.org/blog/</link> <description></description>-
<lastBuildDate>Tue, 06 Oct 2026 22:57:33 +0000</lastBuildDate>+
<lastBuildDate>Thu, 08 Oct 2026 19:36:43 +0000</lastBuildDate> <language>en-US</language> <sy:updatePeriod> hourly </sy:updatePeriod> <sy:updateFrequency> 1 </sy:updateFrequency>-
<generator>https://wordpress.org/?v=7.1.2</generator>+
<generator>https://wordpress.org/?v=7.1.3</generator><image> <url>https://pytorch.org/wp-content/uploads/2024/10/cropped-favicon-32x32.webp</url>-
<title>Blog – PyTorch</title>-
<link>https://pytorch.org</link>+
<title>Blog - PyTorch</title>+
<link>https://pytorch.org/blog/</link> <width>32</width> <height>32</height></image> <item>+
<title>Session-Aware Agentic Inference with NVIDIA Dynamo</title>+
<link>https://pytorch.org/blog/session-aware-agentic-inference-with-nvidia-dynamo/</link>+
+
<dc:creator><![CDATA[Ishan Dhanani, Jamie Li, Karen Chung]]></dc:creator>+
<pubDate>Thu, 08 Oct 2026 18:13:46 +0000</pubDate>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=172542</guid>+
+
<description><![CDATA[<p>TL;DR Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel,...</p>+
<p>The post <a href="https://pytorch.org/blog/session-aware-agentic-inference-with-nvidia-dynamo/">Session-Aware Agentic Inference with NVIDIA Dynamo</a> appeared first on <a href="https://pytorch.org">PyTorch</a>.</p>+
]]></description>+
<content:encoded><![CDATA[<h3>TL;DR</h3>+
<p>Agentic workloads change the traffic an inference server sees. Unlike single-turn chat, an agent session can involve a large initial prefill, repeated model calls, and subagents running in parallel, with KV cache remaining resident while tools execute between turns.</p>+
<p>This technical blog post details how NVIDIA Dynamo uses a unified, session level identifier to transform request-level serving infrastructure into a program-aware system, unlocking session-aware routing, shared KV cache indexing, and programmatic cache movement across vLLM and SGLang.</p>+
<h2>Introduction</h2>+
<p>Agentic workloads have changed the shape of the traffic an inference server sees. A coding session opens with a large prefill, often tens of thousands of tokens of system prompt, tool definitions, and user-specific guidance, and every turn after that appends to a context the session resends in fu…+
<p><!-- IMAGE PLACEHOLDER: Figure 1 (image1) - download from the Doc and upload via WP media library --></p>+
<p><em><img fetchpriority="high" decoding="async" class="alignnone wp-image-172935 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/Agentic-Model-Chain-GIF-2-1.gif" alt="" width="1280" height="720" /></em></p>+
<p><em>Figure 1: Compares the alternating user and model turns of a standard chatbot with an agentic workflow that includes tool calls and tool responses. In the agentic workflow, a single user request can trigger multiple inference calls, with subsequent steps depending on intermediate results.</em…+
<p>Most open-source serving stacks are still built for the model on top of the figure. They route, admit, and cache each request on its own, with no notion of the harness or the session that ties it to previous turns.</p>+
<p>Serving this shape optimally is something we have been building towards. In our previous post, <a href="https://developer.nvidia.com/blog/full-stack-optimizations-for-agentic-inference-with-nvidia-dynamo/">Full-Stack Optimizations for Agentic Inference</a>, we shared work across three layers of t…+
<p>This post covers that progress in four parts:</p>+
<ol>+
<li><strong>Session_ID:</strong> How Dynamo recognizes sessions and subagents from existing harness headers, and how custom harnesses can provide the same information.</li>+
<li><strong>Tracing and replay:</strong> How session-linked traces connect harness activity to inference performance and support offline simulation and live benchmarking.</li>+
<li><strong>Session-aware scheduling:</strong> How routing and admission control account for session working sets and apply backpressure at tool boundaries to reduce cache thrashing.</li>+
<li><strong>Shared KV cache awareness:</strong> How the experimental shared-pool indexer lets the router account for reusable KV in external stores when placing requests.</li>+
<li><strong>Programmatic KV cache:</strong> How the Session-Prefix Indexer and proposed KvHint interface support session-aware cache policies across memory tiers, with Dynamo supplying policy intent and inference engines controlling execution.</li>+
</ol>+
<h2>Identify Agent Sessions with Session ID</h2>+
<p>A Dynamo session ID is a unique stable identifier for a single agentic chain of reasoning and tool calls. Every LLM request in a given trajectory shares the same <code>session_id</code>. Usually, child sessions carry a <code>parent_session_id</code>, so traces and replay tools can establish a con…+
<p>There are a few different ways to share a stable session identifier with Dynamo, depending on what’s calling it.</p>+
<h3>Out-of-the-box support for coding agents</h3>+
<p>Dynamo recognizes the identity headers that today’s popular coding agents already emit and maps each onto the same internal <code>agent_context</code>: the session header becomes <code>session_id</code>, and the parent header, where present, becomes <code>parent_session_id</code>. Claude Co…+
<table>+
<thead>+
<tr>+
<th>Coding agent</th>+
<th>Session header</th>+
<th>Parent header</th>+
</tr>+
</thead>+
<tbody>+
<tr>+
<td>Claude Code</td>+
<td><code>x-claude-code-session-id</code> (root), <code>x-claude-code-agent-id</code> (child agents)</td>+
<td><code>x-claude-code-parent-agent-id</code></td>+
</tr>+
<tr>+
<td>Codex</td>+
<td><code>thread-id</code></td>+
<td><code>x-codex-parent-thread-id</code></td>+
</tr>+
<tr>+
<td>OpenCode</td>+
<td><code>x-opencode-session-id</code></td>+
<td><code>x-opencode-parent-session-id</code></td>+
</tr>+
</tbody>+
</table>+
<p>We’re working with harness providers across the ecosystem to make this the default everywhere, so support expands without users changing anything. Dynamo’s CI tracks each first-class harness as it changes. You can find a full list of supported first class IDs <a href="https://docs.nvi…+
<h3>Plugins for harnesses that don’t send session identity natively</h3>+
<p>For harnesses that don’t emit a compatible header on their own, the <a href="https://github.com/ai-dynamo/agent-plugins">agent-plugins repo</a> provides small integrations that translate a harness’s native session identity into <code>x-dynamo-session-id</code>. Three are available tod…+
<h3>Custom harnesses: opt in with one header</h3>+
<p>A custom harness opts in with one canonical header, <code>X-Dynamo-Session-ID</code>. Dynamo normalizes it into an internal <code>agent_context</code> struct that the rest of the stack reads, so nothing downstream needs to know which harness the request came from. If the harness supports subagent…+
<pre><code class="language-bash">curl http://localhost:8000/v1/chat/completions \
+
-H 'Content-Type: application/json' \
+
-H 'Authorization: Bearer sk-dummy' \
+
-H 'x-dynamo-session-id: research-run-42:researcher' \
+
-d '{"model":"my-model","messages":[{"role":"user","content":"..."}]}'
+
</code></pre>+
<h2>Harness <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2194.png" alt="↔" class="wp-smiley" style="height: 1em; max-height: 1em;" /> Inference Codesign</h2>+
<p>Agent observability tools can already capture what happens at the harness level, but they don’t yet show a unified picture of how the inference stack performed underneath each of those steps. We think about this gap as harness <img src="https://s.w.org/images/core/emoji/17.0.2/72x72/2194.pn…+
<h3>Capturing request traces</h3>+
<p>To enable trace collection, run your deployment with <code>DYN_REQUEST_TRACE=1</code>. Dynamo automatically records information on each request, without storing prompt, response, or tool-call content by default (these can be opt-in enabled). Instead, Dynamo emits a <code>request_end</code> record…+
<p><!-- IMAGE PLACEHOLDER: Figure 2 (image2) - download from the Doc and upload via WP media library --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172895 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/Dynosym-trace-view.png" alt="" width="1440" height="810" srcset="https://pytorch.org/wp-content/uploads/2026/10/Dynosym-trace-view.png 1440w, https://pytorch.org/wp-content/up…+
<p><em>Figure 2: Shows a screen capture of the dynamo trace, with separate tracks for agents and their tool calls. Prefill, decode, and tool execution appear on the same timeline, showing how work is sequenced and overlaps across the session.</em></p>+
<p>Because every record carries timestamps and finish reasons, we can see tool-call distributions and connect them to KV cache eviction behavior. Internally, we’ve used these techniques to optimize our model deployments end to end, using captured traces as backtesting data. Much of this data a…+
<h3>Replaying agentic traces</h3>+
<p>A captured trace is a reusable benchmark. Because it records the workload the agent sent rather than the agent’s decisions, one capture replays the same request schedule as many times as you want without rerunning the model, the tools, or any external API. Replay stays content-free the same…+
<p>One capture drives two paths. Offline, <a href="https://github.com/ai-dynamo/aisimulate">AISimulate</a> runs the request graph against Dynamo’s simulated scheduler, router, and KV cache, letting you compare worker counts, serving topologies, routing policies, and cache capacity without spen…+
<h2>Routing and admission control at the session layer</h2>+
<p>KV cache has always mattered for efficient LLM inference, but agentic inference adds real complexity: where the cache lives, how much of it can be reused at each memory tier, and whether the system is proactive about prioritization, eviction, and sharing. Throughput ends up being determined less …+
<p>Because the Dynamo router is flexible enough to act on more than just the next request, we’ve been able to iterate on session-aware routing and admission control strategies that target these characteristics directly, improving performance for serving agentic-shaped workloads (including RL p…+
<h3>The problem with request-level routing</h3>+
<p>A request-level router, Dynamo’s stock <code>KvRouter</code> included, solves placement one request at a time: given a prompt, which worker has the most cache overlap. That’s the right unit for a single turn, but it’s the wrong unit for an agent, and at high concurrency it misse…+
<ul>+
<li><strong>Cache-occupancy blowup.</strong> Between turns, an agent’s cache doesn’t get freed. It sits in HBM holding blocks while the agent runs a tool call outside the GPU. With N agents concurrently at step k, the aggregate working set is <code>N × context_k</code>, and a request-lev…+
<li><strong>No tool-boundary backpressure.</strong> When a worker is over capacity, a request-level router’s only levers are to cancel or queue in-flight requests. Both are worse than the alternative: deferring the session at the pause point right before its next turn arrives, when it’s …+
</ul>+
<h3>Session Aware admission control</h3>+
<p>Our first agent-aware routing strategy ports the scheduler from the ThunderAgent paper (Kang et al., 2026) on top of Dynamo’s KV-aware router. Written in Rust, it owns a <code>KvRouter</code> directly rather than sitting in front of one as an extra proxy hop, and it tracks real <code>prompt…+
<p>The scheduler groups requests by <code>program_id</code> (the session ID) and moves each program through two independent states: <code>(REASONING | ACTING) × (ACTIVE | PAUSED)</code>. A program enters <code>ACTING</code> at a tool boundary. Under memory pressure, the scheduler pauses <code>ACTING…+
<h3>The control loop</h3>+
<p>Pause and resume are both driven off one quantity per worker: utilization, the fraction of a worker’s KV pool occupied by its assigned programs’ working sets:</p>+
<pre><code class="language-plaintext">U_worker = (sum of token weights for ACTING programs on worker) / (worker's KV pool capacity)
+
</code></pre>+
<p>Three thresholds (defaults shown) turn that number into a control loop:</p>+
<ul>+
<li><strong>U ≥ 0.95 (pause-threshold):</strong> the worker is over-subscribed. The tick pauses the smallest ACTING programs first until U falls back to 0.80 (pause-target).</li>+
<li><strong>0.80 ≤ U < 0.95 (soft-demote band):</strong> programs aren’t paused yet, but take a -2.0 priority penalty as early backpressure before a hard pause is needed.</li>+
<li><strong>Resume only fires once U ≤ 0.85:</strong> that’s pause-threshold minus a 0.10 resume-hysteresis, a buffer below the pause line that keeps the loop from flapping between pausing and resuming right at the threshold.</li>+
</ul>+
<p>Resume itself is a bin-packing pass: paused programs are sorted by token count ascending, and the scheduler admits them smallest-first until the next one would push utilization back over threshold, accounting for a fixed per-program buffer. Resumed requests get a one-second priority boost so they…+
<h3>Results</h3>+
<p>At moderate to high concurrency, we see a large increase in throughput. the difference shows up directly in throughput. On SWE-bench, running two TP4 MiniMax-M2 replicas on a single 8xH100 node, program-aware scheduling improved throughput roughly 12-16% over KV-routing alone, with the gap coming…+
<p>The same strategy shows a similarly large boost on agentic RL rollouts. Running Uni-Agent SWE-Bench on 8xH20-3e with Qwen3-Coder-30B-A3B-Instruct against VERL’s default Global LB, our strategy stayed roughly even at low concurrency (CC32-128), where nothing needed to pause. At medium concur…+
<p><!-- IMAGE PLACEHOLDER: Figure 3 (image3) - download from the Doc and upload via WP media library --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172897 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/Dynamo-Concurrency-Regime-Chart.png" alt="" width="1440" height="810" srcset="https://pytorch.org/wp-content/uploads/2026/10/Dynamo-Concurrency-Regime-Chart.png 1440w, https:/…+
<p><em>Figure 3: Compares model-token throughput across agent admission concurrency levels for VERL Global LB, Dynamo Python TA, and Dynamo’s native session-aware scheduler. At high concurrency, the native scheduler maintains the highest throughput of the three, while Global LB throughput drops shar…+
<h3>Shared KV cache awareness in the router</h3>+
<p>KV cache placement gets harder once our working set outgrows HBM. The cache doesn’t disappear, it moves into CPU memory and then into shared stores like Mooncake or NIXL backed storage. So far the Dynamo router was only aware of GPU and native framework CPU offloaded memory. When KV moved i…+
<p>At request time, Dynamo hashes the prompt into SGLang KV pages and expands each logical page into the physical Mooncake objects implied by the worker’s TP/PP layout, K/V tensors, and optional backend tag. Mooncake publishes ordered <code>Store</code> and <code>Remove</code> events and Dynamo cons…+
<h2>Towards Programmatic KV Cache</h2>+
<p>Ultimately, performant inference on agentic workloads requires session-aware KV block management. To predict the priority value of a given KV cache block in such workloads, we need complementary information from the router and the engine. The router is aware of session lifecycle and reuse pattern…+
<p>This is the central idea of <strong>programmatic KV cache</strong>. Essentially, the Dynamo router provides the “brain” that calculates smart, session/load/cache-aware KV movement policies, while the engine “executes” the movements that Dynamo mandates. As such, we have been developing features f…+
<h2>Router-Driven KV Hints Policy</h2>+
<p>Dynamo’s programmatic KV cache movements are driven by a <strong>narrow, router-initiated hint surface (”KvHint”)</strong> which enables Dynamo to pass KV cache intent to vLLM/SGLang. This avoids invasive changes to the engine’s native scheduler and cache manager. (<a href="https://github.com/sgl…+
<p>The orchestrator holds what a request-local policy cannot see. Which sessions are still live, which token ranges are a shared prefix versus a unique tail, whether a tool-call gap will last milliseconds or minutes, when a subagent opens and closes. That knowledge belongs in the router layer, where…+
<p>The proposal is a narrow set of soft hints the router can emit to the engine. Here is a partial taxonomy of potential KvHints and their functions:</p>+
<table>+
<thead>+
<tr>+
<th>KvHint name</th>+
<th>Functionality</th>+
<th>Example use</th>+
</tr>+
</thead>+
<tbody>+
<tr>+
<td>Share</td>+
<td>Moves a cached prefix from one worker to another via p2p transfer</td>+
<td>A new engine worker warms its cache from the existing workers’ cache.</td>+
</tr>+
<tr>+
<td>Prefetch</td>+
<td>Fetches KV from a higher tier (e.g. DRAM) to lower tier, ahead of need</td>+
<td>Reloading a main agent’s context into HBM as a subagent approaches completion.</td>+
</tr>+
<tr>+
<td>Demote</td>+
<td>Pushes KV from a lower tier to a higher tier.</td>+
<td>Long tool call, or paused trajectory.</td>+
</tr>+
<tr>+
<td>Pin</td>+
<td>Pins a high-value KV cache onto its current tier for a bounded TTL</td>+
<td>Pinning a system cache shared by all requests.</td>+
</tr>+
<tr>+
<td>Retain</td>+
<td>Biases KV blocks to be more (or less) likely to be retained under cache pressure.</td>+
<td>Retaining a subagent’s KV blocks during a tool call.</td>+
</tr>+
</tbody>+
</table>+
<p>Each hint carries the session ID with it, so the tier receiving the KV knows which session it belongs to and can group, place, and reclaim it as a unit rather than as anonymous blocks. That matters most in the colder tiers, where a storage backend holding host memory or disk otherwise sees only o…+
<p>Share already has a working implementation layered into HiCache with minimal scheduler hooks, and Prefetch and Demote are designed to build on that machinery once Share lands.</p>+
<h2>Enabling KV Hint Policy Computations with Session-Prefix Indexer</h2>+
<p>For Dynamo to compose and execute these KvHint policies on the session level, it must own a continuously updated structure which bidirectionally maps between KV blocks, their session identities, and their in-session positional lineage. For this, we implemented the Session-Prefix Indexer, a radix …+
<p>The Session-Prefix Indexer is continuously hydrated from both inference engine KV events (Stored / Removed) and the prefix cache hits from Dynamo router’s main indexer (“FlashIndexer”). Consequently, through the FlashIndexer and Session-Prefix Indexer, the Dynamo router has a cumulative view of (…+
<p>With these indexers, we are actively experimenting with KvHint policies with Dynamo vLLM/SGLang on agentic workload datasets.</p>+
<p><!-- IMAGE PLACEHOLDER: Figure 4 (image4) - download from the Doc and upload via WP media library --></p>+
<p><em><img decoding="async" class="alignnone wp-image-172896 size-full" src="https://pytorch.org/wp-content/uploads/2026/10/Dynamo-Router-with-Session-Prefix-.png" alt="" width="1440" height="810" srcset="https://pytorch.org/wp-content/uploads/2026/10/Dynamo-Router-with-Session-Prefix-.png 1440w, h…+
<p><em>Figure 4: The Session-Prefix Indexer tracks shared prefixes, session lineage, and KV block locations within the Dynamo router. KV events from inference workers update this view, supporting routing decisions and policies that send optional KV hints to vLLM or SGLang.</em></p>+
<h2>Looking forward</h2>+
<p>Efficient agentic inference depends on managing context across the full session lifecycle. Dynamo’s session ID connects harness behavior to routing, admission control, and cache management, giving the serving stack a shared view of an agent’s evolving working set.</p>+
<p>Session-linked traces and replay make policies measurable, while session-aware admission control reduces cache thrashing by applying backpressure at tool boundaries. In the SWE-bench configuration tested here, program-aware scheduling improved throughput by roughly 12–16% over KV-aware routing al…+
<p>The Session-Prefix Indexer and proposed KvHint interface build toward proactive cache movement: Dynamo supplies policy intent, and engines retain control over execution. The key takeaway is that cache reuse and session concurrency must be optimized together. Keeping useful context available acros…+
<p>The post <a href="https://pytorch.org/blog/session-aware-agentic-inference-with-nvidia-dynamo/">Session-Aware Agentic Inference with NVIDIA Dynamo</a> appeared first on <a href="https://pytorch.org">PyTorch</a>.</p>+
]]></content:encoded>+
+
+
+
</item>+
<item>+
<title>Building Spyre as a Native PyTorch Device</title>+
<link>https://pytorch.org/blog/building-spyre-as-a-native-pytorch-device/</link>+
+
<dc:creator><![CDATA[Joshua Rosenkranz, Tuan Hoang Trong, Thomas Gooding, Matthew Pisano, and the IBM Spyre Team]]></dc:creator>+
<pubDate>Thu, 08 Oct 2026 12:45:49 +0000</pubDate>+
<category><![CDATA[Blog]]></category>+
<guid isPermaLink="false">https://pytorch.org/?p=172567</guid>+
+
<description><![CDATA[<p>TL;DR Spyre becomes a native PyTorch device by connecting PyTorch’s existing device, allocator, stream, and compiler abstractions through torch-spyre to the Spyre runtime and firmware. PrivateUse1 gives Spyre a real...</p>+
<p>The post <a href="https://pytorch.org/blog/building-spyre-as-a-native-pytorch-device/">Building Spyre as a Native PyTorch Device</a> appeared first on <a href="https://pytorch.org">PyTorch</a>.</p>+
]]></description>+
<content:encoded><![CDATA[<h3><strong>TL;DR</strong></h3>+
<p>Spyre becomes a native PyTorch device by connecting PyTorch’s existing device, allocator, stream, and compiler abstractions through <a href="https://github.com/torch-spyre/torch-spyre"><code>torch-spyre</code></a> to the Spyre runtime and firmware. <a href="https://docs.pytorch.org/tutorial…+
<p>Spyre is IBM’s dataflow AI accelerator, optimized for inference. It is designed for enterprise teams running AI alongside their applications and data on IBM Z, LinuxONE, and Power systems. Its reduced-precision compute is well suited to the matrix-heavy work in language generation and embed…+
<p>For developers, the challenge is using that hardware through familiar PyTorch code while keeping up with new models and inference techniques. Efficient execution needs more than a compiler: tensors must be able to stay on the device between operations, launches need to be lightweight, and host wo…+
<h2>The dataflow constraints behind the PyTorch surface</h2>+
<p>The card has 32 cores connected by a high-bandwidth ring. Each core has 2 MB of local scratchpad and arrays of processing elements. Up to 128 GB of LPDDR5 provides storage for tensors and programs.</p>+
<p>Those two levels of memory have different owners. The runtime stack allocates and manages LPDDR5. Before a program runs, the compiler has already emitted the loads and stores that move tiles between LPDDR5 and each core’s scratchpad. Data moves across that path in 128-byte sticks, so layout…+
<p>A native PyTorch device still has to honor the execution model of the hardware underneath it. For Spyre, that means mapping PyTorch’s device, allocator, stream, event, and launch abstractions onto a runtime built around compiled programs, ordered queues, explicit data movement, and fixed la…+
<p><strong>Computation is triggered by data, not by a single kernel instruction stream.</strong> A GPU kernel is typically organized around scheduled groups of threads executing the kernel’s instruction sequence. A compiled Spyre kernel is arranged differently: it contains programs for multipl…+
<p><strong>Runtime overlap comes from separate pipelines, not concurrent compute launches.</strong> Spyre has a compute pipeline and data-movement pipelines, so a transfer can run while a compiled program is computing. Those pipelines are not interchangeable: a compiled program can use many cores in…+
<p><strong>Ordered queues preserve per-stream completion order.</strong> Work arrives as a sequence of typed operations – move this in, run this, move that out – and each stream completes those operations in order. To use the separate hardware pipelines without weakening that guarantee, …+
<p><strong>Programs are compiled ahead of time against a fixed layout.</strong> The compiler fixes how values compose into the device’s native <a href="https://github.com/torch-spyre/RFCs/blob/main/0047-TiledTensors/0047-TiledTensorsRFC.md">fixed-size chunks</a> and what order dimensions appea…+
<p>In the implementation described here, PyTorch does not talk to the card directly. <code>torch-spyre</code> provides the PyTorch device integration, while the Spyre runtime and firmware manage device work and communication. That layering is why the rest of this post separates PyTorch concepts such…+
<h2>The PyTorch abstractions that Spyre maps onto</h2>+
<p>PyTorch already has concepts that map onto those constraints.</p>+
<p><strong>A device identity.</strong> <a href="https://docs.pytorch.org/tutorials/advanced/privateuseone.html">PrivateUse1</a> registration starts with two calls:</p>+
<pre><code class="language-python">torch.utils.rename_privateuse1_backend("spyre")
+
torch._register_device_module("spyre", make_spyre_module())</code></pre>+
<p>Those calls give PrivateUse1 the name <code>spyre</code> and register its device module. With the allocator, copy kernels, and dispatcher registrations behind that identity, <code>tensor.to("spyre")</code> works and the dispatcher routes Spyre tensors to Spyre kernels.</p>+
<p><strong>Device-management APIs.</strong> Backend-specific controls live under <code>torch.spyre</code>, while <a href="https://docs.pytorch.org/docs/main/accelerator/index.html"><code>torch.accelerator</code></a> gives portable code a common interface for availability, device selection, streams, …+
<p><strong>A stream abstraction.</strong> Implementing the <a href="https://github.com/pytorch/pytorch/tree/main/test/cpp_extensions/open_registration_extension/torch_openreg/csrc/runtime">device-guard hooks</a> the framework asks for is what makes native <code>torch.Stream(device="spyre")</code> an…+
<p><strong>A device allocator interface.</strong> Its core contract hands back device storage and a way to release it. Additional hooks connect the backend to PyTorch’s memory statistics, cache controls, and stream association.</p>+
<p><strong>Reference-counted tensor storage.</strong> Tensor lifetime becomes the framework’s problem: when the last reference to a Spyre tensor goes away, PyTorch calls back into the backend allocator to release or recycle the device allocation.</p>+
<p><strong>A compiler path through Inductor.</strong> <code>torch.compile(backend="inductor")</code> keeps FX graphs in PyTorch’s compiler pipeline. The same lowering, cache, and launch path can serve compiled models and registered eager operations.</p>+
<p><strong>An event model.</strong> PyTorch defines events as a way to express ordering between streams. Spyre uses the same model internally through runtime events and derived dependency edges.</p>+
<p>Laid side by side, the correspondence is close:</p>+
<table>+
<thead>+
<tr>+
<th>Runtime/Hardware needs</th>+
<th>PyTorch abstraction</th>+
</tr>+
</thead>+
<tbody>+
<tr>+
<td>Ordered queue of typed operations</td>+
<td>Stream</td>+
</tr>+
<tr>+
<td>Transfer pipelines, distinct from compute</td>+
<td>Multiple streams</td>+
</tr>+
<tr>+
<td>Cross-pipeline ordering</td>+
<td>Events, or edges the runtime derives</td>+
</tr>+
<tr>+
<td>Region-budgeted device memory</td>+
<td>Device allocator</td>+
</tr>+
<tr>+
<td>Device-resident values</td>+
<td>Tensor storage and its lifetime</td>+
</tr>+
<tr>+
<td>Ahead-of-time compiled program</td>+
<td>Compiled artifact behind <code>torch.compile</code></td>+
</tr>+
<tr>+
<td>Runtime-managed host work</td>+
<td>Ordered host operations associated with a stream</td>+
</tr>+
</tbody>+
</table>+
<p>The rest of this post walks through the four mappings that take real design work.</p>+
<h2>Mapping 1 – execution pipelines to streams</h2>+
<p>The runtime needs somewhere to put work that is ordered, asynchronous, and per-device. In PyTorch, that abstraction is a stream. For Spyre, the useful mapping is not that the hardware behaves exactly like a PyTorch stream; it is that stream-ordered work can be lowered into typed runtime operation…+
<p>Program-specific work stays above the stream boundary. On Spyre, the framework-facing layer turns the compiled artifact into a prepared launch plan: a recipe that records the host work, data movement, device compute, and synchronization needed to run that artifact. That layer can carry tensor lay…+
<p>The lower layer cannot misinterpret a program description because it never receives one. A different submission mechanism can also sit underneath without changing the framework-facing code.</p>+
<p><!-- IMAGE PLACEHOLDER: Figure 1 - "Where to put the boundary" (upload via WP media library) --></p>+
<p><em><img decoding="async" class="alignnone size-full wp-image-172576" src="https://pytorch.org/wp-content/uploads/2026/10/fig1-where-the-boundary-goes-2.svg" alt="" /></em></p>+
<p><em>Figure 1: Where to put the boundary</em></p>+
<p><strong><em>Keep the program knowledge above the queue.</em></strong> <em>Everything that understands what a compiled program is belongs in the framework-facing layer. The runtime-facing layer accepts typed operations with explicit operands and dependencies for host work, device compute, data mov…+
<p>Preparation and launch also happen at different times. Translating a compiled artifact into submittable operations is per-artifact work, not per-launch work:</p>+
<pre><code class="language-python"># once, when the compiled artifact is first seen
+
job_plan = prepare_kernel(spyrecode_dir)
+
+
# per invocation
+
launch_jobplan(job_plan, args)</code></pre>+
<p>Preparation parses the artifact, allocates and transfers the program binary, and translates the execution plan into an ordered list of typed steps. Launch walks those steps and enqueues operations. Anything that can be resolved when the artifact is first seen should be resolved there, so the laun…+
<h3>Why not another graph at runtime?</h3>+
<p>The earlier integration already started with a PyTorch FX graph through a custom backend implementation. FX was not the problem; graphs belong at the compiler boundary. The mismatch came after that handoff, when launches, copies, and tensor arguments were translated into a second backend-specific…+
<p>Running one already-compiled program meant reading and deserializing that runtime graph, adding nodes and edges for its tensor arguments, and loading and parsing the result before issuing device operations. A host-to-device copy similarly built a small graph around data conversion and transfer. T…+
<p>Copying the entire compiled graph into our own abstractions also caused a break in communication between the core runtime and PyTorch. PyTorch had no visibility as to whether tensors were resident to the device or to the host. As a consequence, PyTorch believed every tensor was a CPU tensor, even…+
<p>The new boundary keeps FX graphs in Inductor and prepares the compiled artifact once. After that, the runtime needs an ordered recipe of typed operations, not another model graph to reconstruct or traverse on every invocation.</p>+
<h2>Mapping 2 – device memory to the allocator</h2>+
<p>A PyTorch-native device cannot treat every operation as “copy inputs to the accelerator, run, and copy outputs back.” A tensor on <code>device="spyre"</code> needs real device storage whose lifetime PyTorch owns through its allocator and storage contracts. The earlier custom <code>tor…+
<p>Spyre adds a hardware constraint to that basic allocator contract. Framework allocators commonly present device memory as a flat pool addressed by pointers. Underneath that interface, Spyre manages memory in regions – each a contiguous chunk of device memory identified by a handle. A region…+
<p>Spyre supports PF mode, where a card is dedicated to one tenant, and VF mode, where several tenants share it. VF mode has the tighter handle budget. The allocator acquires a small number of large regions and sub-allocates aligned blocks within them, allowing many live tensors to share a handful o…+
<p>Every allocation uses the same kind of description in either mode: one or more region-and-offset pieces. Only the interpretation of the region identifier changes, so the layers above the allocator do not branch on the deployment mode. The allocator returns that description and a way to free it; P…+
<p>The fuller allocator interface adds PyTorch’s standard memory-management API, including usage statistics, cache controls, and stream association. It is useful, but it is not what makes tensors resident.</p>+
<p>An allocation here is described by more than an address: the basic addressable piece is a region identifier, an offset within that region, and a length. A tensor interleaved across memory domains is several such pieces, not one flat pointer, and PyTorch has room for that: an allocation can carry …+
<p><!-- IMAGE PLACEHOLDER: Figure 2 - "What a device allocation really is" (upload via WP media library) --></p>+
<p><em><img decoding="async" class="alignnone size-full wp-image-172575" src="https://pytorch.org/wp-content/uploads/2026/10/fig2-what-an-allocation-is-2.svg" alt="" />Figure 2: What a device allocation really is</em></p>+
<p><strong><em>An allocation is a description, not just an address.</em></strong> <em>A tensor bound to one memory domain is a single region-and-offset piece; a tensor interleaved across domains is several. PyTorch carries that description in an opaque context, and its reference counting tells the b…+
<p>The same representation can cover hardware with non-uniform memory. On such a device, memory is divided into domains and a core reaches some domains more efficiently than others. A tensor can be bound near the cores that use it or interleaved across domains to use their combined bandwidth. This i…+
<p>Placement must be explicit and binding because compiled code may depend on it. Every layer that carries an allocation therefore also needs to carry whether the allocation is bound to one domain or spread across several.</p>+
<p>Residency also makes layout an explicit contract. Model adapters place weights and KV caches in layouts compiled kernels expect, while compiler passes insert legal layout conversions for intermediate tensors. The work did not disappear; it moved from a hidden runtime graph into model preparation,…+
<p>Device placement and offload are separate concerns. Placement determines where a tensor lives within the accelerator. Offload moves state that is not immediately needed, such as KV-cache pages, to a secondary storage tier and restores it before use. That tier may be host memory or, at the serving…+
<h2>Mapping 3 – ahead-of-time programs to launch-time work</h2>+
<p>Every ahead-of-time accelerator has to solve this somewhere. The compiler knows the structure of the computation but not where the caller’s data will be, and the two have to be reconciled before the program runs. Spyre has used three approaches as the runtime has evolved. The runtime now us…+
<p><strong>The earlier graph-mediated path let the compiled job own the addresses.</strong> Inputs were copied into device buffers owned by the job, and results were copied back out. Those buffers could keep concrete addresses across launches because their lifetime and placement belonged to the job.…+
<p><strong>An interim runtime pointed the program’s address windows at tensors.</strong> The device translation, or xlat, table exposes a small set of address windows, so the runtime bound each window to one allocator-placed tensor. Nothing was patched and the tensors did not need to be copied…+
<p><strong>The current runtime patches the program at launch.</strong> Addresses remain symbolic until the allocator has placed the tensors. This avoids both repeated input and output copies and the window-count ceiling, making it the generic path for allocator-placed tensors. On Spyre the compiler …+
<ol>+
<li><strong>A host callback</strong> runs on the CPU and writes the resolved addresses into a pinned host buffer.</li>+
<li><strong>A transfer</strong> copies that buffer into the program’s own allocation, at an offset the compiler specified.</li>+
<li><strong>The computation</strong> runs, reads those values, and patches its operands.</li>+
</ol>+
<p>On one stream, completion ordering is enough to keep the three steps from racing. In the initial two-stream design described next, the host callback and transfer remain ordered on one preparation stream. The compute operation runs on the device stream and waits for an event recorded after the tra…+
<p>Queue ordering controls when operations are issued. It does not control hardware that runs ahead of them. Spyre’s program distribution does exactly that: the unit feeding instruction buffers fetches eagerly rather than waiting for the current program to finish, so it can pull the program th…+
<p><!-- IMAGE PLACEHOLDER: Figure 3 - "Launch-time patching as three ordered steps" (upload via WP media library) --></p>+
<p><em><img decoding="async" class="alignnone size-full wp-image-172577" src="https://pytorch.org/wp-content/uploads/2026/10/fig3-launch-time-patching-1.svg" alt="" />Figure 3: Launch-time patching as three ordered steps</em></p>+
<p><strong><em>Launch-time patching is three ordinary steps and two guarantees.</em></strong> <em>A host callback writes resolved addresses into a pinned buffer, a transfer moves them into the program’s allocation at a compiler-specified offset, and the device reads them and patches its operan…+
<p>The initial two-stream design can reuse one <strong>pinned host staging buffer</strong> because every write to that buffer and every transfer from it stays on the same preparation stream. Completion ordering prevents the next host callback from overwriting the host buffer until the previous trans…+
<p>The device correction area is separate. The transfer writes it and compute reads it, so an event orders those operations across streams. If the next launch reuses that same device area, its transfer must also wait until the previous compute has finished reading it. Host preparation can still over…+
<p>Launch-time patching is the current general solution. Its host preparation can overlap device compute using the multiple streams described next, and transfers can overlap when they target independent device storage. Compiler-owned addresses and address-window binding may still fit specialized cas…+
<p>Future support would allow the runtime to supply the actual sizes of variable tensor dimensions through the same launch-time correction mechanism used for addresses. When compiled with those dimensions left symbolic, one artifact could serve multiple input sizes within its supported bounds and la…+
<h2>Mapping 4 – independent pipelines to concurrent streams</h2>+
<p>Transfer and compute use separate hardware pipelines, but a single stream can still leave one waiting on the other. In the single-stream sequence, the host prepares correction data, the transfer moves it to the device, and compute consumes it before the next iteration begins. This is correct, but…+
<p>Within an iteration, the host callback must finish before the transfer reads the correction buffer, and the transfer must finish before compute uses the corrected program. That prevents the transfer from reading an incomplete buffer and compute from reading an unpatched program. Across iterations…+
<p>The runtime keeps the required ordering and moves independent work onto another stream. The preparation stream can produce correction data for iteration N+1 while the device stream computes iteration N. An event joins the streams at the actual dependency: compute N+1 waits for its transfer, but u…+
<p>Events connect the queues using a primitive a PyTorch reader already knows: a marker one stream records and another waits on. The preparation stream signals that a transfer is complete; the device stream waits for that signal before computing. Ordering stays total where it matters, expressed betw…+
<p>At the runtime layer, event and synchronization operations map closely to CUDA:</p>+
<table>+
<thead>+
<tr>+
<th>CUDA</th>+
<th>Here</th>+
</tr>+
</thead>+
<tbody>+
<tr>+
<td><code>cudaEventRecord(event, stream)</code></td>+
<td>enqueue an event-signal operation; the signal fires only once everything already queued on that stream has completed</td>+
</tr>+
<tr>+
<td><code>cudaStreamWaitEvent(stream, event)</code></td>+
<td>enqueue an event-wait operation; the queue stops at that wait until the event is signalled, and fails outright if the producing queue died first</td>+
</tr>+
<tr>+
<td><code>cudaStreamSynchronize()</code></td>+
<td>synchronize the queue: block until it drains, re-raising the first error recorded on it</td>+
</tr>+
</tbody>+
</table>+
<p>In CUDA the caller records and waits explicitly. Here the runtime can derive an edge for operations whose device-memory read and write footprints are available. The framework layer routes preparation and device work to the appropriate queues; the runtime inserts signal and wait operations when th…+
<p>The initial arrangement uses two queues in a producer/consumer relationship: a preparation queue carrying host work and its transfers, a device queue carrying compute. Nothing about the mechanism stops at two, and nothing forces the split either – a launch with no host work to do never need…+
<p>Host preparation becomes more visible as launch-time work grows and device compute gets faster. Resolving dimensions and addresses on one queue would serialize work that could run concurrently.</p>+
<p>Patching is not the only host work involved. Preparing a tensor in the device’s expected layout can also run alongside device compute and be connected to its consumer by an event.</p>+
<p>These events are software signals checked by the scheduler. They are single-use, so a fresh event backs each edge. The interface can later use hardware event support without changing the layers above it.</p>+
<h2>What the mapping buys, and what it costs</h2>+
<h3>User experience</h3>+
<p>The surface a user touches is the one they already know. <code>tensor.to("spyre")</code>, <code>torch.compile</code>, and <code>torch.Stream</code> use the same PyTorch concepts as other devices. That does not automatically provide operator coverage, profiling, distributed collectives, CUDA-speci…+
<p>For compile-backed eager operations, eager mode is not a second implementation. Torch-spyre registers a PrivateUse1 kernel around the same decomposition used during compilation. Inside a <code>torch.compile</code> trace, the wrapper calls that decomposition directly so PyTorch can capture it into…+
<p>That choice still exposes first-use compilation; unsupported operations can fail or fall back, and different tensor extents can require another compiled artifact. The benefit is one execution path whose fixes and regressions apply to both modes.</p>+
<p>A compile-backed eager operation is also the smallest possible compiled program: one operation, real tensors, real device memory, and a real launch. If a model fails but the corresponding eager operation passes, the fault is more likely in composition than in the operation. A test for a single op…+
<p>And because residency is generic rather than hand-written per case, tensors stay on the device without bespoke residency code for each kind of tensor. That is what makes eager mode practical.</p>+
<p>When the device behaves like a normal PyTorch device, a new model implementation is mostly just a model implementation. It can be picked up and run rather than ported, because its operations already dispatch and its tensors already live in the right place. The same holds one level up. Inference t…+
<p>The compiler makes the arithmetic fast, but much of an inference stack’s performance comes from how work is scheduled, batched, overlapped and speculated. A backend that maps cleanly onto the framework’s primitives leaves that space open to people who never touch the compiler.</p>+
<h3>Performance</h3>+
<p>The performance argument is about what leaves the launch path. Submitting typed operations to a queue, rather than constructing a graph to describe them, removes a disk read, a deserialize, a per-argument stitch, and a compile-and-parse from every single launch. A launch becomes building a few op…+
<p>On a real model the effect compounds, because those costs were being paid per launch. On <a href="https://huggingface.co/ibm-granite/granite-3.3-8b-instruct">Granite 3.3 8B</a> at batch size 1 and sequence length 1024, taking graph construction off the transfer path alone is worth 9.9% on prefill…+
<p>The concurrency described in Mapping 4 is a separate opportunity: every piece of host preparation moved off the compute path reduces device idle time.</p>+
<h3>What PyTorch does not provide by itself</h3>+
<p>PyTorch has a device-event abstraction, but the cross-stream ordering described above has to be enforced where the operations actually sit – in the runtime scheduler, which is the layer that sees their device-memory footprints. The runtime therefore builds software event primitives and deri…+
<h3>What we put back</h3>+
<p>Building on a framework’s extension points is not a one-way transaction. PyTorch’s core developers maintain <a href="https://pytorch.org/blog/openreg-a-self-contained-pytorch-accelerator-simulator/"><code>openreg</code></a>, the upstream reference backend and accelerator simulator tha…+
<h2>Conclusion</h2>+
<p>The win is that neither Spyre nor PyTorch has to pretend to be something else. Spyre keeps its dataflow execution model, while users get a native PyTorch device with familiar tensor placement, streams, and compilation.</p>+
<p>For Spyre, PyTorch’s interfaces provide generic tensor residency, one compile-and-launch path for eager and compiled execution, lower launch overhead, and a clear way to overlap host preparation with device work. They also leave less backend-specific machinery to maintain. Dispatch, storage…+
<p>For PyTorch, Spyre demonstrates that those extension points can support hardware that is not organized around a conventional GPU execution model. The lessons that flow back through <a href="https://pytorch.org/blog/openreg-a-self-contained-pytorch-accelerator-simulator/"><code>openreg</code></a> …+
<p>The post <a href="https://pytorch.org/blog/building-spyre-as-a-native-pytorch-device/">Building Spyre as a Native PyTorch Device</a> appeared first on <a href="https://pytorch.org">PyTorch</a>.</p>+
]]></content:encoded>+
+
+
+
</item>+
<item> <title>Modernizing Table Batched Embeddings with FBTriton</title> <link>https://pytorch.org/blog/modernizing-table-batched-embeddings-with-fbtriton/</link> Diff display stops at 400 lines. The line counts above are from the whole diff. 93 lines shown here cut at 300 characters. The raw artifact at this commit is linked above.