spdl.io.transfer_tensor_d2h

transfer_tensor_d2h(batch: T, /, *, device: TDevice | str | None = None, stream: torch.cuda.Stream | None = None) → T[source]

Transfer PyTorch CUDA tensors to CPU through a dedicated stream.

Added in version 0.7.0.

This function performs efficient GPU to CPU data transfer using page-locked (pinned) memory and a dedicated CUDA stream. The page-locked memory is cached and reused across calls.

The transfer process: 1. Gathers all tensors from the batch. 2. Allocates (or reuses cached) page-locked memory. 3. Asynchronously transfers data from GPU to page-locked memory. 4. Copies data from page-locked memory to new CPU tensors. 5. Rebuilds the batch structure with CPU tensors.

The copy stream waits for work already submitted to the caller’s current stream, and this function waits for the copy stream before returning. If a tensor was produced on another non-current stream, the caller must first establish an ordering dependency with the current stream. When called from a background CPU thread, the transfer can overlap with later GPU work submitted independently by a foreground thread. It is intended for offloading nested results before CPU post-processing or serialization.

If any transferred tensor requires gradients, the function uses a synchronous per-tensor transfer on the caller’s current stream to preserve autograd.

Example

import torch
from spdl.io import transfer_tensor_d2h

batch = {"scores": torch.randn(32, device="cuda:0")}
cpu_batch = transfer_tensor_d2h(batch, device="cuda:0")
Parameters:
  • batch – A torch.Tensor or a composition of tensors with container types such as list, tuple, dict and dataclass.

  • device –

    Optional CUDA device to transfer data from.

    If None the source device is determined by the LOCAL_RANK environment variable. If not set, cuda:0 is used.

  • stream –

    Optional Custom CUDA stream to use for the transfer. If None, a stream is created from the device argument, and cached to a thread-local storage for future reuse.

    When stream is not None, the device argument must be provided. The stream must be on the same device.

Returns:

An object of the same type as the input, but the PyTorch CUDA tensors on the specified device are transferred to CPU.

If there is no PyTorch tensor in the input, the input is returned as-is.

If there is no CUDA device available, the input is returned as-is.

Raises:
  • ValueError – If device is not a CUDA device, or a custom stream does not match the requested device.

  • RuntimeError – If the resolved CUDA device index is unavailable.