logoalt Hacker News

volf_today at 9:22 PM0 repliesview on HN

I've got a working recipe to run this model on Dual DGX Spark: https://github.com/volfco/spark-vllm-docker/blob/main/recipe...

Averages ~25-35tok/s which isn't bad for a first attempt.