Nice work! I have been messing with doing some radio DSP on the NPU in my laptop and found the Intel graph compiler to be rather unpredictable. Weird random mis-lowering to fp16 ops instead of int8, unpredictable throughput depending on tensor sizes, reshapes being sometimes extremely expensive and sometimes basically free etc.
Obviously I am mostly just driving the DPU, as the SHAVE cores are just not beefy enough, but maybe this will be useful at some point, at least to gain more insight into the architecture...
Nice work! I have been messing with doing some radio DSP on the NPU in my laptop and found the Intel graph compiler to be rather unpredictable. Weird random mis-lowering to fp16 ops instead of int8, unpredictable throughput depending on tensor sizes, reshapes being sometimes extremely expensive and sometimes basically free etc.
Obviously I am mostly just driving the DPU, as the SHAVE cores are just not beefy enough, but maybe this will be useful at some point, at least to gain more insight into the architecture...