My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.