Memory is not necessarily a good thing. What we need is a combination of sufficiently large context windows (10-100 million tokens) along with curate, bloat-free data.