> The mentioned approach is fundamentally flawed, since the inputs are used during pretraining constituting to a leakage, a universally recognized flaw of ML training.
I saw this on the community note for the last blog you wrote - anything to do here.
Not true. This is allowed in a metalearning context. Its called transductive learning and has existed since the 90s: https://en.wikipedia.org/wiki/Transduction_(machine_learning...
I address this in more detail in the blog