by the guest editors Peter Grünwald (CWI and Leiden University, Wouter Koolen (CWI and University of Twente) and Johanna Ziegel (ETH Zurich)
As new measurements become available over time, we face the classic problem of updating our information state. In science, this typically means refining our view of hypotheses based on experimental outcomes – either determining if the data allow us to reject a null hypothesis, or estimating which parameter values remain statistically plausible. Anytime-valid methods allow us to reliably refine these assessments sequentially while guaranteeing at most a controlled fraction of mistakes.
by Glenn Shafer
Data analysis requires principles as well as mathematics. Traditionally, we have relied on Cournot’s principle when we use probability theory for data analysis. But when we test by betting instead of relying on small probabilities, we can formulate principles that dig deeper into statistical practice and apply more broadly.
by Martin Larsson (Carnegie Mellon University), Aaditya Ramdas (Carnegie Mellon University), and Johannes Ruf (London School of Economics)
Under what conditions do optimal bets against a given probabilistic hypothesis exist? Answer: they always do!
by Eugenio Clerico (University of Oxford)
While e-values and p-values are often presented as competitors, they share deep structural connections. We highlight a functional perspective linking the two, and suggest how it might lead to new ways of translating p-value methods into the e-value framework.
by Peter Grünwald (CWI and Leiden University)
A major criticism of p-values and standard confidence intervals, first coined around 1960, is their sensitivity to counterfactuals: their validity depends on how data would have been collected in situations that never occurred, which is often unknown or even unknowable. The fact that e-based methods remain valid under optional continuation implies that they do not suffer from this problem…or does it?
by Vaidehi Dixit (University of Nottingham) and Ryan Martin (North Carolina State University)
Good e-processes can be constructed for testing a specific null hypothesis against a specific alternative, but general inference need not have a specific alternative in mind. In such cases, one might seek an alternative hypothesis-agnostic e-process with fast growth rate under a wide range of alternatives. Our predictive recursion-based e-process construction offers just that, along with some deeper insights related to “objective” empirical probability.
by Etienne Gauthier (Inria, Ecole Normale Supérieure, PSL Research University)
How can we trust a model’s predictions in the presence of uncertainty? Conformal prediction provides a principled framework for attaching reliable confidence guarantees to machine learning outputs. By incorporating e-values, this framework moves beyond rigid, pre-specified guarantees. It enables finer control over the predictions in a dynamic way, adapting the reliability of AI systems to the constraints of real-world applications.
by Rianne de Heide (University of Twente and CWI)
Modern data analysis can test thousands of scientific questions at once, from genes in cancer studies to voxels in brain scans. A new general principle called e-closure gives researchers more freedom to explore these results after seeing the data, while keeping false discoveries under control.
by Zhimei Ren (University of Pennsylvania)
How can we combine discoveries from multiple FDR-controlling rejection sets without losing statistical validity? In joint work with Rina Foygel Barber, we show that the knockoff procedure can be represented through e-values and e-BH, allowing rejection sets from multiple randomized runs to be aggregated by averaging e-values while preserving FDR control. More broadly, this e-value perspective provides a general framework for merging FDR-controlling procedures and opens new directions for understanding, improving, and aggregating large-scale testing methods.
by Gianna Serafina Monti (University of Milano-Bicocca), and Peter Filzmoser (TU Wien)
E-values provide a principled foundation for false discovery rate control in high-dimensional microbiome data analysis. We show how their aggregation properties enable a derandomized, robust knockoff filter that outperforms classical approaches in stability and reproducibility.
by Yo Joong Choe (INSEAD) and Sebastian Arnold (CWI)
New research develops novel statistical methodology, based on e-values, for monitoring whether one uncertain prospect (say, a new investment option) has an upside over another (say, the current investment). These e-values can flexibly and effectively test for multiple notions of upside over time, as defined by the decision-maker’s preferences, and they come with a direct monetary interpretation that guides the decision (say, whether to invest in the new option).
by Wouter M. Koolen (CWI & University of Twente), Shubhada Agrawal (IISc Bangalore) and Martin Larsson (CMU)
Could our data be sub-Gaussian noise? We explore rejecting that null hypothesis with the help of e-variables. We map the landscape of optimal e-variables against two-point alternatives.
by Thorsten Dickhaus (University of Bremen), Francesca Giuffrida (Leiden University and IMT School for Advanced Studies Lucca) and Yonqqi Wang (CWI)
We present powerful and easy-to-compute e-values for the classical statistical task of testing associations between two binary traits based on contingency table data. Genetic case-control association studies are our main intended use case.
by Yongxi Long and Erik van Zwet (Leiden University Medical Centre)
Anytime valid testing with e-values offers great flexibility allowing both optional stopping (peeking) and optional continuation (collecting more data). The price to pay is a reduction of statistical power. Using data from more than 20,000 randomized trials, we evaluate how e-values compare with classical p-values in balancing flexibility and efficiency.
by Sebastian Arias , Alexander Ly, Michele Meziu (CWI) and Angel Reyero Lobo (CWI and Inria)
Modern science generates data continuously, but the statistical methods that still dominate many fields generally require data collection to end before reliable meta-analysis can begin. New research on e-values offers a way to analyse evidence in real time, without sacrificing statistical reliability. The approach could make science not just more robust to modern research practices, but significantly more efficient.
by Stan Koobs and Nick W. Koning (Erasmus University Rotterdam)
A classical statistical problem is to assess whether an unknown quantity is negligible: that is, practically equivalent to zero. It is standard to define ‘negligible’ as being smaller in magnitude than some threshold margin. Specifying this margin has plagued statisticians for decades: if it is set too large, then one can hardly speak of negligibility, but if the margin is set too small, then one may need an enormous amount of data to statistically establish negligibility. In recent work, we study this problem in depth and show how e-values can be used to bypass it, by enabling one to select the margin post-hoc: after seeing the data.
by Adrienne Tuynman and Timothée Mathieu (Univ. Lille, Inria, CNRS, Centrale Lille, UMR 9189 – CRIStAL)
Political polls before elections are useful to identify promising candidates, and to allow parties to make compromises or build alliances. We are interested in conducting polls sequentially, so that one can stop acquiring data as soon as possible while safely yielding statistically significant results.
by Guneet Singh Dhillon (University of Oxford), Teodora Pandeva (Microsoft Research), and Alicia Curth (Microsoft Research)
Generative AI systems are becoming ubiquitous, but their outputs can still be inaccurate or misleading. Using e-values, the e-scores framework provides a statistically rigorous assessment of AI-generated responses while accommodating the adaptive and post-hoc nature of human-AI interactions.
by Stephan Bongers (CWI)
Standard statistical guarantees fail when analysts repeatedly check incoming data. By integrating anytime-valid inference with reinforcement learning, this work enables safe policy evaluation under continuous monitoring.
by Ruodu Wang (University of Waterloo)
A model-free method lets regulators and financial institutions continuously monitor tail-risk forecasts using e-values that remain valid whenever checked.
by Michael Scott Lindon (Netflix)
The statistical guarantees designed to protect against human failures in sequential experimentation turn out to be exactly what is needed to govern autonomous AI agents conducting experiments.