The Nvidia video presentation on DLSS 5 says that the model was only trained with various G-buffers as input (including the depth buffer) but during inference, the model only uses the rendered frame. As well as the previous rendered frame reprojected via motion vectors, if I understand correctly, likely to improve temporal stability.