Meta-Analytic Pooling of CLES Estimates: A Comparison of Transformation Approaches

Bethany Hamilton Bhat and Wolfgang Viechtbauer

Department of Psychiatry and Neuropsychology, Faculty of Health, Medicine and Life Sciences, Maastricht University

8 June 2026

Motivation

Common measures for two-group comparisons on quantitative outcomes:

  • Mean differences
  • Standardized mean differences (\(d\))
  • Log ratio of means (lnRR)
  • Measures of stochastic superiority (\(\theta\))

Stochastic Superiority (\(\theta\))

The probability that a randomly selected person from one group has a higher outcome than a randomly selected person from another group

\[\theta = P(Y>X)+\frac{1}{2}P(Y=X)\]

  • \(\frac{1}{2}P(Y=X)\) term handles ties
  • Several parametric and non-parametric estimators:
    • CLES, AUC, Wilcoxon U, Concordance probability, Cliff’s \(\delta\)

McGraw and Wong (1992);Hanley and McNeil (1982); DeLong et al. (1988);Newson (2002);Cliff (1993)

Estimating \(\theta\): CLES

  • Proposed by McGraw and Wong (1992) under normality and homoscedasticity
  • Under these assumptions, \(d\) completely determines any measure of overlap between two distributions

\[\hat{\theta} = \Phi\left(\frac{d}{\sqrt{2}}\right), \quad \mathrm{var}(\hat{\theta}) = \frac{e^{-d^2 / 2}}{4\pi} \left(\frac{1}{n_1} + \frac{1}{n_2} + \frac{d^2}{2(n_1+n_2)} \right)\]

  • There is no comprehensive evaluation on how to best pool CLES estimates

Meta-analyzing CLES Estimates

  • CLES:
    • \(\hat{\theta} \in [0, 1]\)
    • Sampling distribution is skewed near 0 and 1
  • Three transformations of \(\hat{\theta}_i\):
    • logit, arcsine, and probit/SMD
  • Two back-transformations to obtain \(\hat{\mu}_{\theta}\):
    • Standard
    • Integral

Back-transformation Methods for \(\hat{\mu}_{\theta}\)

Two approaches to estimate \(\mu_{\theta}\) after pooling on the transformed scale:

  • Standard:
    • Returns the median of the distribution of true effects, not the mean (Jensen’s inequality)
  • Integral:
    • Integrates \(g^{-1}\) over the normal random-effects distribution
    • Recovers the mean of the distribution of true effects
    • See Igelmann et al. (2026)

Data Generation

  1. Sample \(k\) true CLES values (\(\theta_i\)) from \(\text{Beta}(\alpha, \beta)\) with moments \((\mu_\theta, \tau^2_\theta)\)

  2. Sample study sample sizes from right-skewed distribution (\(\gamma_1 \approx 1.45\)):

\[n_i \sim (\overline{n}/20) \times \chi_{df=4}^2 + (3\overline{n}/10),\]

where \(\overline{n}\) is the average sample size for a primary study and assuming equal group sizes Marín-Martínez and Sánchez-Meca (1998); Sanchez-Meca and Marin-Martinez (1998)

  1. Generated observed CLES estimates (\(\hat{\theta}_i\)) implied by the model that
  • Assumed normally distributed outcomes with equal population variances
  • \(\sigma_X = \sigma_Y = 1\) and \(\mu_X = 0\)
  • \(\mu_Y\) chosen so that \(P(Y>X)=\theta_i\)

Simulation Design

Factor Values
CLES parameter (\(\mu_{\theta}\)) 0.5, 0.555, 0.635, 0.710, 0.776, 0.833
Number of studies (\(k\)) 10, 20, 40, 80, 160
Avg. participants per study (\(\bar{n}\)) 30, 50, 80, 100
Between-study heterogeneity (\(\tau_{\theta}\)) 0, 0.012, 0.048, 0.094, 0.138
  • 600 conditions \(\times\) 10,000 replications each
  • \(\mu_{\theta}\) and \(\tau_{\theta}\) choosen to correspond to commonly observed SMD values
  • 7 pooling methods; REML estimation (DL when convergence failed); Knapp-Hartung test
  • Performance criteria: bias and RMSE of \(\hat{\mu}_{\theta}\); 95% CI coverage; rejection rates and power for test of \(H_0: \mu_{\theta} = 0.5\) and Q-test of \(H_0: \tau^2 = 0\); bias in \(\hat{\tau}_{\theta}\)

Results: Parameter Bias

Results: RMSE

Results: Coverage

Results: Coverage – Integral Methods Only

Conclusions and Discussion

  • Standard back-transformations are acceptable at low heterogeneity, but integral back-transformation is recommended when \(\tau_\theta\) is large
  • Under normality and homoscedasticity, SMD with integral back-transformation is the recommended
    • arcsine and logit transformations with the integral back-transformation were comparable across most conditions (except for coverage for moderate \(\tau_{\theta}\) and large \(k\))
    • Direct pooling should be avoided
  • Next Study:
    • Extend to non-normal outcomes and unequal population variances and evaluate other estimates (and corresponding sampling variances) of stochastic superiority

Thank you! Questions?

References

Cliff, Norman. 1993. “Dominance Statistics: Ordinal Analyses to Answer Ordinal Questions.” Psychological Bulletin 114 (3): 494–509. https://doi.org/10.1037/0033-2909.114.3.494.
DeLong, Elizabeth R., David M. DeLong, and Daniel L. Clarke-Pearson. 1988. “Comparing the Areas Under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach.” Biometrics 44 (3): 837–45. https://doi.org/10.2307/2531595.
Hanley, J. A., and B. J. McNeil. 1982. “The Meaning and Use of the Area Under a Receiver Operating Characteristic (ROC) Curve.” Radiology 143 (1): 29–36. https://doi.org/doi: 10.1148/radiology.143.1.7063747. PMID: 7063747.
Igelmann, Jan-Bernd, Markus Pauly, and Wolfgang Viechtbauer. 2026. “Back-Transformations in Random-Effects Meta-Analysis—Impact and Interpretation.” Research Synthesis Methods, May, 1–14. https://doi.org/10.1017/rsm.2026.10096.
Marín-Martínez, Fulgencio, and Julio Sánchez-Meca. 1998. “Testing for Dichotomous Moderators in Meta-Analysis.” The Journal of Experimental Education 67 (1): 69–81. https://www.jstor.org/stable/20152582.
McGraw, Kenneth O., and S. P. Wong. 1992. “A Common Language Effect Size Statistic.” Psychological Bulletin (US) 111 (2): 361–65. https://doi.org/10.1037/0033-2909.111.2.361.
Newson, Roger. 2002. “Parameters Behind ‘Nonparametric’ Statistics: Kendall’s Tau, Somers’ D and Median Differences.” The Stata Journal 2 (1): 45–64. https://doi.org/10.1177/1536867X0200200103.
Sanchez-Meca, Julio, and Fulgencio Marin-Martinez. 1998. “Weighting by Inverse Variance or by Sample Size in Meta-Analysis: A Simulation Study.” Educational and Psychological Measurement 58 (2): 211–19.

Extra Slides

Transformations of CLES Estimates

Direct Transformations, Back-transformations, and Sampling Variances
Transformation \(g(\hat{\theta})\) \(g^{-1}(t)\) \(\mathrm{var}(g(\hat{\theta}))\)
Logit \(\ln\left(\dfrac{\hat{\theta}}{1-\hat{\theta}}\right)\) \(\dfrac{1}{1+e^{-t}}\) \(\dfrac{1}{\hat{\theta}^2(1-\hat{\theta})^2} \times \mathrm{var}(\hat{\theta})\)
Arcsine \(\arcsin\left(\sqrt{\hat{\theta}}\right)\) \(\sin^2(t)\) \(\dfrac{1}{4\hat{\theta}(1-\hat{\theta})} \times \mathrm{var}(\hat{\theta})\)
Probit \(\Phi^{-1}(\hat{\theta})\) \(\Phi(t)\) \(\dfrac{1}{\left[\phi(\Phi^{-1}(\hat{\theta}))\right]^2} \times \mathrm{var}(\hat{\theta})\)
SMD \(\Phi^{-1}(\hat{\theta})\sqrt{2}\) \(\Phi\left(\dfrac{t}{\sqrt{2}}\right)\) \(\dfrac{1}{n_1} + \dfrac{1}{n_2} + \dfrac{d^2}{2(n_1+n_2)}\)

Back-transformation Methods

Integral Back-Transformations for Each Method
Transformation Integral
Logit \(\displaystyle\int_{-\infty}^{\infty} \dfrac{1}{1+e^{-t}} \, \phi(t \mid \mu, \tau^2) \, dt\)
Arcsine \(\displaystyle\int_{0}^{\pi/2} \sin^2(t) \, \dfrac{\phi(t \mid \mu, \tau^2)}{\Phi\left(\frac{\pi/2-\mu}{\tau}\right) - \Phi\left(\frac{-\mu}{\tau}\right)} \, dt\)
Probit \(\displaystyle\int_{-\infty}^{\infty} \Phi(t) \, \phi(t \mid \mu, \tau^2) \, dt\)
SMD \(\displaystyle\int_{-\infty}^{\infty} \Phi\left(\dfrac{t}{\sqrt{2}}\right) \phi(t \mid \mu, \tau^2) \, dt\)

Further Details on Between-study Heterogeneity

  • Specified \(\mu_{\theta}\) and \(\tau_{\theta}^2\) values as the moments of the beta distribution

  • Found the \(\alpha\) and \(\beta\) parameters of the beta distribution by: \(\alpha = \mu_{\theta} \times \nu\) and \(\beta = (1 - \mu_{\theta})\nu\) where \(\nu = (\alpha + \beta) =\frac{\mu_{\theta}(1 - \mu_{\theta})}{\tau^2_{\theta}} - 1\).

    • Note. values must meet the following condition: \(\nu = (\alpha + \beta) > 0\), so therefore \(\tau^2_{\theta} < \mu_{\theta}(1 - \mu_{\theta})\)
  • For the specified \(\mu_{\theta}\) and \(\tau_{\theta}^2\), we derived them from SMD parameters through numerical integration

    • specified all 20 pairwise combinations of the following: \(\mu_d = \{0, 0.2, 0.5, 0.8, 1.1, 1.4\}\) and \(\tau_d = \{0, 0.05, 0.20, 0.40, 0.60\}\)

Beta Distributions

Obtaining CLES Estimates

  1. After obtaining true CLES values, then transform to true SMD values:

\[ g_{\text{d}}(\theta_i) = \Phi^{-1}(\theta_i)\sqrt{2} \]

  1. Generate observed SMD estimates from a scaled non-central \(t\)-distribution assuming homoscedastic variances:

\[g_{\text{d}}(\hat{\theta_i}) \sim N\left(g_{\text{d}}(\theta_i), \sqrt{\frac{1}{n_1} + \frac{1}{n_2}}\right) \bigg/ \sqrt{\frac{\chi^2(df= n_1+n_2 -2)}{n_1 + n_2 -2}}\]