spdl.io.transfer_tensor_d2h¶
- transfer_tensor_d2h(batch: T, /, *, device: TDevice | str | None = None, stream: torch.cuda.Stream | None = None) T[source]¶
Transfer PyTorch CUDA tensors to CPU through a dedicated stream.
Added in version 0.7.0.
This function performs efficient GPU to CPU data transfer using page-locked (pinned) memory and a dedicated CUDA stream. The page-locked memory is cached and reused across calls.
The transfer process: 1. Gathers all tensors from the batch. 2. Allocates (or reuses cached) page-locked memory. 3. Asynchronously transfers data from GPU to page-locked memory. 4. Copies data from page-locked memory to new CPU tensors. 5. Rebuilds the batch structure with CPU tensors.
The copy stream waits for work already submitted to the caller’s current stream, and this function waits for the copy stream before returning. If a tensor was produced on another non-current stream, the caller must first establish an ordering dependency with the current stream. When called from a background CPU thread, the transfer can overlap with later GPU work submitted independently by a foreground thread. It is intended for offloading nested results before CPU post-processing or serialization.
If any transferred tensor requires gradients, the function uses a synchronous per-tensor transfer on the caller’s current stream to preserve autograd.
Example
import torch from spdl.io import transfer_tensor_d2h batch = {"scores": torch.randn(32, device="cuda:0")} cpu_batch = transfer_tensor_d2h(batch, device="cuda:0")
- Parameters:
batch – A
torch.Tensoror a composition of tensors with container types such aslist,tuple,dictanddataclass.device –
Optional CUDA device to transfer data from.
If
Nonethe source device is determined by theLOCAL_RANKenvironment variable. If not set,cuda:0is used.stream –
Optional Custom CUDA stream to use for the transfer. If
None, a stream is created from thedeviceargument, and cached to a thread-local storage for future reuse.When stream is not
None, thedeviceargument must be provided. The stream must be on the same device.
- Returns:
An object of the same type as the input, but the PyTorch CUDA tensors on the specified device are transferred to CPU.
If there is no PyTorch tensor in the input, the input is returned as-is.
If there is no CUDA device available, the input is returned as-is.
- Raises:
ValueError – If
deviceis not a CUDA device, or a custom stream does not match the requested device.RuntimeError – If the resolved CUDA device index is unavailable.