CUDA has had managed memory that pages between VRAM and system RAM for a decade. Problem is doing so...

kcb • last Wednesday at 11:23 PM • 2 replies • view on HN

CUDA has had managed memory that pages between VRAM and system RAM for a decade. Problem is doing so is unusably slow for AI purposes. Seems like an unnecessary layer here.

Replies

hrmtst93837 • last Thursday at 8:22 AM

That slowness is almost useful. It makes the failure mode obvious instead of letting a 'transparent' layer hide it until some sloppy alloc or tensor blowup starts paging through system RAM or NVMe and the whole job turns into a smoke test for your storage stack.

For actual training, explicit sharding and RAM mapping are ugly, but at least you can see where the pressure is and reason about it. 'Transparent' often just means performance falls off a cliff and now debugging it sucks.

yjtpesesu2 • last Wednesday at 11:47 PM

[dead]

alt Hacker News

Replies