logoalt Hacker News

librasteve • today at 7:52 PM • 1 reply • view on HN

As the video points out, there is a hardware/software feedback loop at play here. Since the hardware is deeply pipelined SIMD FPU datapath, the software is hand tuned machine coded transformations. In addition to limiting flexibility (every variation has to be ground out in CUDA), this prevents sparse matrix type optimisations. I predict that a set of general purpose CPUs - non shared memory at this scale - would be a much better use of transistor/power. And you can code that at high level give a CSP style approach such as https://bil-lang.org


Replies

feffe • today at 10:11 PM

I think tenstorrent architecture is more like this. A grid of RISC-V cores with local SRAM and vector units.