logoalt Hacker News

stymaaryesterday at 8:45 PM1 replyview on HN

n-gram per-layer embeddings[1][2] might be it.

[1] https://sebastianraschka.com/llm-architecture-gallery/per-la...

[2]: See DS 4.1-Flash and Qwen-3.8-Next.


Replies

verdvermyesterday at 8:48 PM

this is to offload VRAM to DRAM (for GP comment), and makes no difference for URAM

show 4 replies