And the hidden, trial and error runs tweaking the hyperparams. guaranteed, it cost >$5000 in training runs alone.
I always wonder how often people in charge of massive compute tasks like LLM pre-training have messed up some detail that invalidates or fails to persist the results and only realise afterwards.
I always wonder how often people in charge of massive compute tasks like LLM pre-training have messed up some detail that invalidates or fails to persist the results and only realise afterwards.