I think Taalus is a bit reluctant to publish precisely how they do it but vaguely:
> We basically have an architecture where we are embedding the models, and we are hard coding the models and the weights into our what we call the mask ROM recall fabric, which is paired with an SRAM recall fabric. Together, they are able to store both the model as well as do all the computations of KV cache. We have adapters and customizations – we support all of that. This design allows us to be super-dense in terms of compute and in terms of storage, and we can do compute on that storage incredibly fast, which is what drives density up and cost down
Source: https://www.nextplatform.com/compute/2026/02/19/taalas-etche...