Naive question: would flash or some kind of write once memory be dense enough to replace a mask ROM? I'm wondering whether it is feasible to manufacture "blank" chips at the semiconductor factory that have a fixed architecture, but are model-agnostic. These chips would get the then-current weights burned in on first use. It's probably quite wasteful, but it could keep the same chip design alive for a longer time.
DRAM seems to sit in the economic sweet spot of density and bandwidth. I'm not sure if any other storage technology can provide non-volatile or one time programable storage with either the same density and bandwidth. Although I must say, it seems insanely wasteful to use RAM to store model weights, there must a better solution.
The closest anyone has gotten is Cerebras with their Wafer Scale Engine. It uses SRAM embedded with the compute. A single chip is an entire wafer, but the headline spec, how much ram, only 44GB, which is tiny for the silicon area used.
I understand traditional IC production workflows are ludicrously expensive, and glacially slow, but surely at some point the economics are going to tip in favour of mask rom.
Say you setup your foundry/packaging/ai chip facility. You come up with a new set of model weights. Run your CI/CD pipeline to produce a new mask output. The only thing you've changed are the assignments of the bits, this is extremely low risk change. The new masks should be completely interchangeable with the current process.
You produce the new masks, swap them into your foundry process, and all of a sudden your new chips have the latest version of the model. This would probably manifest itself in the form of inference providers having yearly / bi-yearly "updates" to their models as new hardware is brought online.
Tiered subscription levels would gate keep access to the latest and greatest model, cheaper subscriptions will be limited to older versions of the model, and so on, until running the hardware is no longer economically viable (no demand/running costs exceeding what the market is willing to pay).
I think Taalus is a bit reluctant to publish precisely how they do it but vaguely:
> We basically have an architecture where we are embedding the models, and we are hard coding the models and the weights into our what we call the mask ROM recall fabric, which is paired with an SRAM recall fabric. Together, they are able to store both the model as well as do all the computations of KV cache. We have adapters and customizations – we support all of that. This design allows us to be super-dense in terms of compute and in terms of storage, and we can do compute on that storage incredibly fast, which is what drives density up and cost down
Source: https://www.nextplatform.com/compute/2026/02/19/taalas-etche...