But something like TPC is diverse enough to show whole picture?.. Especially compared to popular clickbench.
It is absolutely without a doubt 100% not diverse enough.
I've spent years of my life running TPC benchmarks. They're useful guidelines, but they're simply not diverse enough to show the whole picture unless your picture is very simple. Certainly no single TPC benchmark in isolation. Maybe you can get a better idea by running all of them.
The largest gaps are around non-relational style data, JSON querying and the like. Nearly every organisation has something in that space now, and TPC-H/TPC-DS don't touch it at all.
I think the person you're replying to is more so saying that by choosing the right queries, machine, disk setup, caching, settings, thread count, RAM amount, etc, you can get quite different results. There is a reason everyone always wins their own benchmarks, and it doesn't even have to be dishonest - you optimize and iterate for your own benchmark whereas everyone else just gets one shot to do well out of the box.
I tried my best to be as transparent and fair as possible, running everyone with out-of-the-box settings on a third party's queries (DuckDB), a third party's data generator (tpcgen-rs, from the DataFusion guys), on a stock setup available to everyone (AWS machines).
The one exception is that we also ran Polars locked to 32 threads on the large machine (in addition to the out-of-the-box setup), which was to highlight we can do a lot better on small data. We still suffer from a relatively high constant overhead on very high CPU count machines if the data isn't large enough, but I'm working on fixing that. It's possible that DuckDB / DataFusion have similar scaling issues with high thread counts and would do better with 32 threads as well, I didn't test that.