Loading Events

« All Events

Hybrid Event
  • This event has passed.

Fontana, J. (STAT) – When We’re Always Wrong: Scalable Variable Selection in M-Open Settings

July 24 @ 2:00 pm5:00 pm
Hybrid Event
Abstract digital illustration featuring gears and interconnected technology elements.

A ubiquitous task in statistical practice is that of variable selection, identifying which of a large set of features are the relevant ones. As data sets with a large number of observations have become increasingly common, new theoretical and computational challenges for model selection have emerged. We consider the variable selection problem for linear models in the M-open setting, where the data generating process is outside the model space. We focus on the novel problem of “model superinduction”, which refers to the tendency of model selection procedures to select larger models at an exponential rate as the sample size grows, resulting in overparametrized models which collapse model interpretability and induce severe computational difficulties. For a set of popular frequentist information criteria and the Bayesian case of mixtures of g-priors, we prove that when comparing nested models, the larger model will always be asymptotically selected. We seek to minimize this effect for large n while preserving variable selection consistency.We propose utilizing a mixture of g-priors, where the hyper-prior on g has hyper-parameters chosen to result in a slowly diminishing rate of prior influence on the posterior, which favors simpler models while preserving consistency. We also propose a model space prior which induces stronger model complexity penalization for large sample sizes. The posterior model probabilities under our prior choices further provide an alternative information criterion that is resistant to the effects of model superinduction.

Next, we extend our results to other classes of popular variable selection priors, the family of spike and slab priors, the non-local priors, and selection procedures that correspond to posterior modes such as the LASSO. We show that these procedures are all afflicted with model superinduction, except for the continuous spike and slab priors when a Student-t distribution is used for both the spike and the slab.

Finally, we address existing bottlenecks in the computation efficacy of spike and slab based variable selection. We demonstrate that Gibbs samplers scale poorly to large sample sizes, and the proposed alternatives in the literature result in posterior surrogates that are afflicted with superinduction. Instead, we propose a search strategy based off easy to compute approximations of the posterior model probabilities. We show this procedure, Fast Approximate Stochastic Search (FASS), coupled with post-hoc inference on parameters via either a Block-Variational Bayes approach or an Expectation Propagation approach, results in competitive performance. We demonstrate the aforementioned phenomena, and the efficacy of our proposed solutions, via synthetic data examples and case studies using albedo data from GOES satellites and LCA application data.

Event Host: Jacob Fontana, Ph.D. Candidate, Statistical Science

Advisor: Bruno Sansó

Zoom: https://ucsc.zoom.us/j/99965109575?pwd=uGkOwWM3Rl3zP66aL1RecoOfB8Yat0.1

Passcode: 879019

Details

Other

Room Number
E2-215

Venue