Recommendations for visual predictive checks in Bayesian workflow

Main Article Content

Teemu Säilynoja
https://orcid.org/0000-0002-5249-348X
Andrew Johnson
Osvaldo Martin
https://orcid.org/0000-0001-7419-8978
Aki Vehtari
https://orcid.org/0000-0003-2164-9469

Abstract


Introduction

A key step in the Bayesian workflow for model building is the graphical assessment of model predictions, whether these are drawn from the prior or posterior predictive distribution. The goal of these assessments is to identify whether the model is a reasonable (and ideally accurate) representation of the domain knowledge and/or observed data. There are many commonly used visual predictive checks which can be misleading if their implicit assumptions do not match the reality. Thus, there is a need for more guidance for selecting, interpreting, and diagnosing appropriate visualizations. As a visual predictive check itself can be viewed as a model fit to data, assessing when this model fails to represent the data is important for drawing well-informed conclusions.




Demonstration

We present recommendations for appropriate visual predictive checks for observations that are: continuous, discrete, or a mixture of the two. We also discuss diagnostics to aid in the selection of visual methods. Specifically, in the detection of an incorrect assumption of continuously-distributed data: identifying when data is likely to be discrete or contain discrete components, detecting and estimating possible bounds in data, and a diagnostic of the goodness-of-fit to data for density plots made through kernel density estimates.




Conclusion

We offer recommendations and diagnostic tools to mitigate ad-hoc decision-making in visual predictive checks. These contributions aim to improve the robustness and interpretability of Bayesian model criticism practices.




Materials

The source code implementing the visualization functions and probabilistic models, as well as extended case-studies and the data for the examples are provided online. The version of these materials used in this article is available in Zenodo, and the latest version at https://github.com/TeemuSailynoja/visual-predictive-checks/tree/main/code.


Article Details

How to Cite
[1]
Säilynoja, T. et al. 2026. Recommendations for visual predictive checks in Bayesian workflow. Journal of Visualization and Interaction. 1, 1 (Sep. 2026). DOI:https://doi.org/10.54337/jovi.v1i1.11478.
Section
Articles

References

Agresti, Alan. 2013. Categorical Data Analysis. Third edition. Wiley Series in Probability and Statistics. Wiley-Interscience.

Ayer, Miriam, H. D. Brunk, G. M. Ewing, W. T. Reid, and Edward Silverman. 1955. “An Empirical Distribution Function for Sampling with Incomplete Information.” The Annals of Mathematical Statistics 26 (4): 641–47. https://doi.org/10.1214/aoms/1177728423. DOI: https://doi.org/10.1214/aoms/1177728423

Box, George E. P. 1980. “Sampling and Bayes’ Inference in Scientific Modelling and Robustness.” Journal of the Royal Statistical Society. Series A (General) 143 (4): 383. https://doi.org/10.2307/2982063. DOI: https://doi.org/10.2307/2982063

Bürkner, Paul-Christian. 2017. “brms: An R Package for Bayesian Multilevel Models Using Stan.” Journal of Statistical Software 80 (1): 1–28. https://doi.org/10.18637/jss.v080.i01. DOI: https://doi.org/10.18637/jss.v080.i01

Bürkner, Paul-Christian. 2018. “Advanced Bayesian Multilevel Modeling with the R Package brms.” The R Journal 10 (1): 395–411. https://doi.org/10.32614/RJ-2018-017. DOI: https://doi.org/10.32614/RJ-2018-017

Bürkner, Paul-Christian. 2021. “Bayesian Item Response Modeling in R with brms and Stan.” Journal of Statistical Software 100 (5): 1–54. https://doi.org/10.18637/jss.v100.i05. DOI: https://doi.org/10.18637/jss.v100.i05

Czado, Claudia, Tilmann Gneiting, and Leonhard Held. 2009. “Predictive Model Assessment for Count Data.” Biometrics 65 (4): 1254–61. https://doi.org/10.1111/j.1541-0420.2009.01191.x. DOI: https://doi.org/10.1111/j.1541-0420.2009.01191.x

DeGroot, Morris H., and Stephen E. Fienberg. 1983. “The Comparison and Evaluation of Forecasters.” The Statistician 32 (1/2): 12. https://doi.org/10.2307/2987588. DOI: https://doi.org/10.2307/2987588

Dimitriadis, Timo, Tilmann Gneiting, and Alexander I. Jordan. 2021. “Stable Reliability Diagrams for Probabilistic Classifiers.” Proceedings of the National Academy of Sciences 118 (8): e2016191118. https://doi.org/10.1073/pnas.2016191118. DOI: https://doi.org/10.1073/pnas.2016191118

Freedman, David, and Persi Diaconis. 1981. “On the Histogram as a Density Estimator:L 2 Theory.” Zeitschrift Für Wahrscheinlichkeitstheorie Und Verwandte Gebiete 57 (4): 453–76. https://doi.org/10.1007/BF01025868. DOI: https://doi.org/10.1007/BF01025868

Früiiwirth-Schnatter, Sylvia. 1996. “Recursive Residuals and Model Diagnostics for Normal and Non-Normal State Space Models.” Environmental and Ecological Statistics 3 (4): 291–309. https://doi.org/10.1007/BF00539368. DOI: https://doi.org/10.1007/BF00539368

Gabry, Jonah, Rok Češnovar, Andrew Johnson, and Steve Bronder. 2024. cmdstanr: R Interface to ’CmdStan’. https://mc-stan.org/cmdstanr/.

Gabry, Jonah, and Tristan Mahr. 2024. bayesplot: Plotting for Bayesian Models. https://doi.org/10.32614/CRAN.package.bayesplot. DOI: https://doi.org/10.32614/CRAN.package.bayesplot

Gabry, Jonah, Daniel Simpson, Aki Vehtari, Michael Betancourt, and Andrew Gelman. 2019. “Visualization in Bayesian Workflow.” Journal of the Royal Statistical Society: Series A (Statistics in Society) 182 (2): 389–402. https://doi.org/10.1111/rssa.12378. DOI: https://doi.org/10.1111/rssa.12378

Gelman, Andrew, John B. Carlin, Hal S. Stern, David B. Dunson, Aki Vehtari, and Donald B. Rubin. 2013. Bayesian Data Analysis. 0th ed. Chapman; Hall/CRC. https://doi.org/10.1201/b16018. DOI: https://doi.org/10.1201/b16018

Gelman, Andrew, Jennifer Hill, and Aki Vehtari. 2020. Regression and Other Stories. 1st ed. Cambridge University Press. https://doi.org/10.1017/9781139161879. DOI: https://doi.org/10.1017/9781139161879

Gelman, Andrew, Aki Vehtari, Daniel Simpson, et al. 2020. “Bayesian Workflow.” arXiv:2011.01808 [Stat], November. http://arxiv.org/abs/2011.01808.

Gelman, Andrew, Meng Xiao-Li, and Hal S. Stern. 1996. “Posterior Predictive Assessment of Model Fitness via Realized Discrepancies.” Statistica Sinica 6 (4): 733–60. https://www.jstor.org/stable/24306036.

Goodrich, Ben, Jonah Gabry, Imad Ali, and Sam Brilleman. 2024. rstanarm: Bayesian Applied Regression Modeling via Stan. https://doi.org/10.32614/CRAN.package.rstanarm. DOI: https://doi.org/10.32614/CRAN.package.rstanarm

Guo, Ziyang, Alex Kale, Matthew Kay, and Jessica Hullman. 2025. “VMC: A Grammar for Visualizing Statistical Model Checks.” IEEE Transactions on Visualization and Computer Graphics (USA) 31 (1): 798–808. https://doi.org/10.1109/TVCG.2024.3456402. DOI: https://doi.org/10.1109/TVCG.2024.3456402

Kay, Matthew. 2023. ggdist: Visualizations of Distributions and Uncertainty. Zenodo. https://doi.org/10.5281/zenodo.3879620. DOI: https://doi.org/10.31219/osf.io/2gsz6

Kay, Matthew. 2024. “ggdist: Visualizations of Distributions and Uncertainty in the Grammar of Graphics.” IEEE Transactions on Visualization and Computer Graphics 30 (1): 414–24. https://doi.org/10.1109/TVCG.2023.3327195. DOI: https://doi.org/10.1109/TVCG.2023.3327195

Kay, Matthew, Tara Kola, Jessica R. Hullman, and Sean A. Munson. 2016. “When (Ish) Is My Bus?: User-Centered Visualizations of Uncertainty in Everyday, Mobile Predictive Systems.” Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (San Jose California USA), May, 5092–103. https://doi.org/10.1145/2858036.2858558. DOI: https://doi.org/10.1145/2858036.2858558

Kleiber, Christian, and Achim Zeileis. 2016. “Visualizing Count Data Regressions Using Rootograms.” The American Statistician 70 (3): 296–303. https://doi.org/10.1080/00031305.2016.1173590. DOI: https://doi.org/10.1080/00031305.2016.1173590

Kuhn, and Max. 2008. “Building Predictive Models in R Using the caret Package.” Journal of Statistical Software 28 (5): 1–26. https://doi.org/10.18637/jss.v028.i05. DOI: https://doi.org/10.18637/jss.v028.i05

Liu, Lei, Ya-Chen Tina Shih, Robert L. Strawderman, Daowen Zhang, Bankole A. Johnson, and Haitao Chai. 2019. “Statistical Analysis of Zero-Inflated Nonnegative Continuous Data: A Review.” Statistical Science 34 (2): 253–79. https://doi.org/10.1214/18-STS681. DOI: https://doi.org/10.1214/18-STS681

Martin, Osvaldo A., Oriol Abril-Pla, Jordan Deklerk, et al. 2026. “ArviZ: A Modular and Flexible Library for Exploratory Analysis of Bayesian Models.” Journal of Open Source Software 11 (119): 9889. https://doi.org/10.21105/joss.09889. DOI: https://doi.org/10.21105/joss.09889

Martin, Osvaldo A, Oriol Abril-Pla, and Jordan Deklerk. 2025. Exploratory Analysis of Bayesian Models. Version v0.3.0. Zenodo. https://doi.org/10.5281/zenodo.15127548.

Narasimhan, Balasubramanian, Steven G. Johnson, Thomas Hahn, Annie Bouvier, and Kiên Kiêu. 2024. cubature: Adaptive Multivariate Integration over Hypercubes. https://doi.org/10.32614/CRAN.package.cubature. DOI: https://doi.org/10.32614/CRAN.package.cubature

Niculescu-Mizil, Alexandru, and Rich Caruana. 2005. “Predicting Good Probabilities with Supervised Learning.” Proceedings of the 22nd International Conference on Machine Learning - ICML ’05 (Bonn, Germany), 625–32. https://doi.org/10.1145/1102351.1102430. DOI: https://doi.org/10.1145/1102351.1102430

R Core Team. 2021. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. https://www.R-project.org/.

Rubin, Donald B. 1984. “Bayesianly Justifiable and Relevant Frequency Calculations for the Applied Statistician.” The Annals of Statistics 12 (4). https://doi.org/10.1214/aos/1176346785. DOI: https://doi.org/10.1214/aos/1176346785

Säilynoja, Teemu, Paul-Christian Bürkner, and Aki Vehtari. 2022. “Graphical Test for Discrete Uniformity and Its Applications in Goodness-of-Fit Evaluation and Multiple Sample Comparison.” Statistics and Computing 32 (2): 32. https://doi.org/10.1007/s11222-022-10090-6. DOI: https://doi.org/10.1007/s11222-022-10090-6

Scott, David W. 1992. Multivariate Density Estimation: Theory, Practice, and Visualization. 1st ed. Wiley Series in Probability and Statistics. Wiley. https://doi.org/10.1002/9780470316849. DOI: https://doi.org/10.1002/9780470316849

Sheather, S. J., and M. C. Jones. 1991. “A Reliable Data-Based Bandwidth Selection Method for Kernel Density Estimation.” Journal of the Royal Statistical Society: Series B (Methodological) 53 (3): 683–90. https://doi.org/10.1111/j.2517-6161.1991.tb01857.x. DOI: https://doi.org/10.1111/j.2517-6161.1991.tb01857.x

Silverman, B. W. 1986. Density Estimation for Statistics and Data Analysis. Chapman and Hall/CRC Monographs on Statistics and Applied Probability, v.26. Routledge.

Štrumbelj, Erik, Alexandre Bouchard-Côté, Jukka Corander, et al. 2024. “Past, Present and Future of Software for Bayesian Inference.” Statistical Science 39 (1): 46–61. https://doi.org/10.1214/23-STS907. DOI: https://doi.org/10.1214/23-STS907

Team, Stan Development. 2018. The Stan Core Library. http://mc-stan.org/.

Tesso, Herman, and Aki Vehtari. 2026. “LOO-PIT Predictive Model Checking.” arXiv Preprint arXiv:2603.02928.

Tukey, John Wilder. 1972. “Some Graphic and Semi-Graphic Displays.” In Statistical Papers in Honor of George W. Snedecor, edited by T. A. Bancroft and S. A. Brown. Iowa State University Press.

Turner, Rolf. 2023. Iso: Functions to Perform Isotonic Regression. https://doi.org/10.32614/CRAN.package.Iso. DOI: https://doi.org/10.32614/CRAN.package.Iso

Van Zwet, Erik W., and Eric A. Cator. 2021. “The Significance Filter, the Winner’s Curse and the Need to Shrink.” Statistica Neerlandica 75 (4): 437–52. https://doi.org/10.1111/stan.12241. DOI: https://doi.org/10.1111/stan.12241

Vehtari, Aki, Daniel Simpson, Andrew Gelman, Yuling Yao, and Jonah Gabry. 2024. “Pareto Smoothed Importance Sampling.” Journal of Machine Learning Research 25 (72): 1–58. http://jmlr.org/papers/v25/19-556.html.

Wesner, Jeff S., and Justin P. F. Pomeranz. 2021. “Choosing Priors in Bayesian Ecological Models by Simulating from the Prior Predictive Distribution.” Ecosphere 12 (9): e03739. https://doi.org/10.1002/ecs2.3739. DOI: https://doi.org/10.1002/ecs2.3739

Wickham, Hadley. 2016. ggplot2: Elegant Graphics for Data Analysis. Springer-Verlag New York. https://doi.org/10.1007/978-3-319-24277-4. DOI: https://doi.org/10.1007/978-3-319-24277-4

Wilkinson, Leland. 1999. “Dot Plots.” The American Statistician 53 (3): 276. https://doi.org/10.2307/2686111. DOI: https://doi.org/10.1080/00031305.1999.10474474