It has long been known that the frequentist one-sided p-value from Null-Hypothesis Statistical Testing can be interpreted as a poor person’s Bayesian posterior probability. Specifically, the one-sided p-value is usually a good approximation to the posterior probability that the effect is in the direction opposite to that stipulated by the alternative hypothesis (e.g., Casella & Berger, 1987; Jeffreys, 1939; Marsman & Wagenmakers, 2017, and references therein). The consequences of this are rather dramatic, so let’s make that concrete. Suppose we ask 100 people about their favorite orientation toilet paper, and 60 answer “over”.
First let’s compute a standard one-sided p-value for this scenario. The alternative hypothesis states that across the population, people prefer the “over” orientation. I had ChatGPT Pro create an html demo, “Two tails, two questions“, which allows users to compute both the one-sided p-value and the posterior probability so that they can be compared side-by-side. For the case of 60 successes out of 100 trials, the p-value equals 0.028444:
For the Bayesian analysis, I initially used a uniform beta(1,1) prior, which gives the following result:
The probability that the chance theta is lower than 1/2 is 0.023022, relatively close to the one-sided p-value! When ChatGPT created the app, it discovered that the correspondence between the one-sided p-value and the posterior probability can be made exact if the p-value is computed by taking half of the probability for the observed data. Maybe this can be seen as some sort of continuity correction. Anyway, here is the result for the “mid-p-value”:
Boom! This exact equality holds only under the uniform beta(1,1) prior. The demo allows users to specify a different prior. When this prior is highly peaked and the data are sparse, the p-value and the posterior probability can be very different. However, most analyses in practice use a relatively uninformative prior, in which case the relation is either exact or approximate. If you do not believe me, you can check out the demo yourself (it requires only html, so should run even on your cell phone).
Interpreted as a poor person’s posterior probability of the sign of the effect, this means that the one-sided p-value is not actually a hypothesis test at all. Instead, it is a method of estimation that seeks to quantify the intensity of our conviction that the effect is positive vs. negative, while completely ignoring the hypothesis that the effect could be exactly zero (cf. Jeffreys, 1939). In other words, a statistician who somehow wishes to make a Big Deal out of the metaphysical speculation that all point-null hypotheses are wrong should embrace the one-sided p-value as the inference tool of choice. In contrast, those who seek to test the null-hypothesis have to look elsewhere (and not at the two-sided p-value, since this can be interpreted as two one-sided p-values corrected for multiplicity).
I am not sure what is more shocking, the fact that (from a Bayesian perspective!) the p-value does not actually test the null-hypothesis but rather ignores it, or the ease with which LLMs can generate educational demo’s on statistical problems — complete with sliders, mathematical exposition, etc. And the demo I have was created based on a single, very general prompt:
Can you build an app that clarifies why the one-sided p-value and the posterior probability are sometimes the same? For my course I am teaching using the binomial, so a binomial test vs theta_0 = 1/2 would be ideal.
With a little extra effort, I would have had the LLM tidy up the equations etc., but I wanted to present the demo in its “raw” form, which is already better than I could have produced myself.
References
Casella, G., & Berger, R. L. (1987) Reconciling Bayesian and frequentist evidence in the one-sided testing problem. Journal of the American Statistical Association, 82, 106-111.
Jeffreys, H. (1939). Theory of Probability. Oxford: Oxford University Press.
Marsman, M., & Wagenmakers, E.–J. (2017). Three insights from a Bayesian interpretation of the one-sided P value. Educational and Psychological Measurement, 77, 529-539.
Two tails, two questions. Demo developed by ChatGPT Pro, September 2026.
Toilet roll image created by ChatGPT Pro “in the style of Buffon”.
Eric-Jan Wagenmakers
Eric-Jan (EJ) Wagenmakers is professor at the Psychological Methods Group at the University of Amsterdam.







