Instead of giving users a confidence interval, ask the user to specify a p value. Then run something like Welch’s t test.
Nah because we care about the size of the effect too.
We can get a p<0.0001 finding of an otherwise terrible, annoying change that saves 1%.
Nah because we care about the size of the effect too.
We can get a p<0.0001 finding of an otherwise terrible, annoying change that saves 1%.