The cost of peeking at test results
Sequential testing (checking results and stopping early if they look good) inflates false-positive rates dramatically. A test designed for 5 percent significance with a fixed sample size will deliver 5 percent false-positive rate. But if you peek at the results halfway through and stop if p-value looks good, your false-positive rate jumps to 15-20 percent. This happens because low-p-value flukes are more common in small samples; stopping on a fluke locks in a false positive.
The mechanism is simple: as you accumulate more samples, your statistical power to detect real effects increases, but it also increases your power to detect random noise. If you stop whenever you see a low p-value, you're guaranteed to stop on noise more often than if you'd waited for the full sample.
Always-valid p-values and online testing
Always-valid p-values fix sequential testing by adjusting the threshold dynamically as samples accumulate. Instead of stopping at p-value less than 0.05, always-valid methods raise the threshold in early peeks (e.g., p-value less than 0.001 at 25 percent of samples) and lower it toward the standard 0.05 at the full sample size. This preserves the 5 percent false-positive rate across all peek points. Most modern experimentation platforms (Optimizely, Statsig) use sequential-safe methods, making it safe to peek without penalty.