Guided tour · Evidence · step 10 of 11
Two numbers are not a comparison
Most invalid scientific comparisons look exactly like valid ones until you check what was held fixed.
Now go and look
Look at how runs are grouped, and find a case the platform refuses to rank.
What you should see
Runs measured at different code distances, error rates, noise models or suite versions are never put in one table. A single merged leaderboard is the artifact that makes an invalid comparison look authoritative.