logoalt Hacker News

__atx__today at 9:56 AM1 replyview on HN

Nice work! I have been messing with doing some radio DSP on the NPU in my laptop and found the Intel graph compiler to be rather unpredictable. Weird random mis-lowering to fp16 ops instead of int8, unpredictable throughput depending on tensor sizes, reshapes being sometimes extremely expensive and sometimes basically free etc.

Obviously I am mostly just driving the DPU, as the SHAVE cores are just not beefy enough, but maybe this will be useful at some point, at least to gain more insight into the architecture...


Replies

hsfzxjytoday at 10:31 AM

Thanks! I can observe similar behaviors on my laptop. FP32 operations are sometimes siliently lowered to a sequence of (FP32->FP16)->(FP16 OP)->(FP16->FP32), causing precision loss. It further frustrated me during npunlock development. Changing custom SHAVE OP from FP16 -> FP32 also changes the blob structure, and I have to explore different binary patching strategy.

As for the interesting cost variance for reshape you've mentioned, I guess surrounding context is the cause. The NPU compiler might adjust the data layout to match successive OPs' requirements. But, yes, that's annoying~

Studying DPU is not my current priority, but I may dig it up in the future to understand its invocation descriptor, and hopefully to discover more interesting stuff. I hope this project ends up being useful to you!