The missing piece is whether “75%” has the same denominator in each chart. If three models were evaluated on the same benchmark set, against the same versions at the same time, with exactly one winner per benchmark, their shares of wins would add up to 100%. They couldn't each win 75% of that shared set. But different benchmark selections, older comparator versions, different reasoning or tool budgets, and later releases change the comparison. Counting ties as wins changes it too. So the marketing percentages alone don't tell you how many benchmarks exist or how much selection happened. To audit a particular claim, I'd want the full shared matrix, exact model versions, benchmark versions, test dates and evaluation settings, rather than just the highlighted cells. That also separates actual leapfrogging from a change in what was measured.
I maintain https://llmbenchmarks.io/analyses/#test-setups, which illustrates some of these differences using published measurements and their sources. Even results for the same model and benchmark can involve different setups. That isn't a verdict on this particular Gemini release; its specific claims need their own source and protocol checks.