logoalt Hacker News

lostmsutoday at 11:15 AM0 repliesview on HN

This is slop. 8M parameter dense model with context length 64 that you train on enwik9 in 2h will have 1.15 bpb. This model has 1.8 (bits per byte, lower is better).