Asynchronous versus synchronous execution

PyTorch Developer Podcast

Nội dung được cung cấp bởi PyTorch, Edward Yang, and Team PyTorch. Tất cả nội dung podcast bao gồm các tập, đồ họa và mô tả podcast đều được PyTorch, Edward Yang, and Team PyTorch hoặc đối tác nền tảng podcast của họ tải lên và cung cấp trực tiếp. Nếu bạn cho rằng ai đó đang sử dụng tác phẩm có bản quyền của bạn mà không có sự cho phép của bạn, bạn có thể làm theo quy trình được nêu ở đây https://vi.player.fm/legal.

4+ y ago 15:03

MP3•Trang chủ episode

CUDA is asynchronous, CPU is synchronous. Making them play well together can be one of the more thorny and easy to get wrong aspects of the PyTorch API. I talk about why non_blocking is difficult to use correctly, a hypothetical "asynchronous CPU" device which would help smooth over some of the API problems and also why it used to be difficult to implement async CPU (but it's not hard anymore!) At the end, I also briefly talk about how async/sync impedance can also show up in unusual places, namely the CUDA caching allocator.

Further reading.

CUDA semantics which discuss non_blocking somewhat https://pytorch.org/docs/stable/notes/cuda.html
Issue requesting async cpu https://github.com/pytorch/pytorch/issues/44343

83 tập

#Tech #PyTorch #Edward Yang #Team PyTorch #Deep Learning #Machine Learning