2023-2024
Fall 2023
Statistical Optimality and Computational Tractability of ICA
Friday, September 29, 2023
Time: 2:00 p.m. to 3:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Ming Yuan, Professor, Department of Statistics and Associate Director, Data Science Institute, Columbia University
Abstract: Independent component analysis (ICA) is a powerful and general data analysis tool. Yet there is an increasing amount of empirical evidence that the classical methods for ICA are not well suited for modern applications, both computationally and statistically, where the effect of dimensionality is not negligible. We will investigate the optimal sample complexity and statistical performance for ICA, and how considerations of computational tractability may affect them. We will also introduce estimating procedures for ICA that are both statistically efficient and computationally tractable. Our development exploits the close connection between ICA and moment estimation and reveals a number of new insights for both problems.
On Fine-Tuning Large Language Models with Less Labeling Cost
Friday, October 13, 2023
Time: 2:00 p.m. to 3:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Tuo Zhao, Assistant Professor, H. Milton Stewart School of Industrial and Systems Engineering, Georgia Tech
Abstract: Labeled data is critical to the success of deep learning across various applications, including natural language processing, computer vision, and computational biology. While recent advances like pre-training have reduced the need for labeled data in these domains, increasing the availability of labeled data remains the most effective way to improve model performance. However, human labeling of data continues to be expensive, even when leveraging cost-effective crowd-sourced labeling services. Further, in many domains, labeling requires specialized expertise, which adds to the difficulty of acquiring labeled data.
In this talk, we demonstrate how to utilize weak supervision together with efficient computational algorithms to reduce data labeling costs. Specifically, we investigate various forms of weak supervision, including external knowledge bases, auxiliary computational tools, and heuristic rule-based labeling. We showcase the application of weak supervision to both supervised learning and reinforcement learning across various tasks, including natural language understanding, molecular dynamics simulation, and code generation.
Gaussian random field approximation for wide neural networks
Friday, October 27, 2023
Time: 2:00 p.m. to 3:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Nathan Ross, Associate Professor, School of Mathematics and Statistics, University of Melbourne
Abstract: It has been observed that wide neural networks (NNs) with randomly initialized weights may be well-approximated by Gaussian fields indexed by the input space of the NN, and taking values in the output space. There has been a flurry of recent work making this observation precise, since it sheds light on regimes where neural networks can perform effectively. In this talk, I will discuss recent work where we derive bounds on Gaussian random field approximation of wide random neural networks of any depth, assuming Lipschitz activation functions. The bounds are on a Wasserstein transport distance in function space equipped with a strong (supremum) metric, and are explicit in the widths of the layers and natural parameters such as moments of the weights. The result follows from a general approximation result using Stein's method, combined with a novel Gaussian smoothing technique for random fields, which I will also describe. The talk covers joint works with Krishnakumar Balasubramanian, Larry Goldstein, and Adil Salim; and A.D. Barbour and Guangqu Zheng.
A Single level Deep Learning Approach to Solve Stackelberg Mean Field Game Problems
Friday, November 3, 2023
Time: 2:00 p.m. to 3:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Gökçe Dayanıklı, Assistant Professor, Department of Statistics, University of Illinois Urbana-Champaign
Abstract: In many real-life policy making applications, the principal (i.e., governor or regulator) wants to find the optimal policies for a large population of interacting agents who optimize their own objectives in a game theoretical framework. However, it is well known that finding an equilibrium in a game with a large number of agents is a challenging problem because of the increasing number of interactions among agents. In this talk, we introduce the Stackelberg mean field game problem to approximate the game between a principal and a large number of agents. Then, we discuss how to rewrite this bi-level problem as a single-level problem to propose an efficient numerical solution. In the model, the agents in the population play a non-cooperative game and choose their controls to optimize their individual objectives by interacting with the principal and other agents in the society through the population distribution. The principal can influence the resulting mean field game Nash equilibrium through incentives to optimize her own objective. After analyzing this game by using a probabilistic approach, we rewrite this bi-level problem as a single-level problem and propose a deep learning approach to solve the Stackelberg mean field game. We look at different applications such as the systemic risk model for a regulator and many banks and an optimal contract problem between a project manager and a large number of employees.
Sparse topic modeling via spectral decomposition and thresholding
Friday, November 10, 2023
Time: 2:00 p.m. to 3:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Claire Donnat, Assistant Professor, Department of Statistics, University of Chicago
Abstract: By modeling documents as mixtures of topics, Topic Modeling allows the discovery of latent thematic structures within large text corpora, and has played an important role in natural language processing over the past decades. Beyond text data, topic modeling has proven itself central to the analysis of microbiome data, population genetics, or, more recently, single-cell spatial transcriptomics. Given the model’s extensive use, the development of estimators — particularly those capable of leveraging known structure in the data — presents a compelling challenge. In this talk, we focus more specifically on the probabilistic Latent Semantic Indexing model, which assumes that the expectation of the corpus matrix is low-rank and can be written as the product of a topic-word matrix and a word-document matrix. Although various estimators of the topic matrix have recently been proposed, their error bounds highlight a number of data regimes in which the error can grow substantially — particularly in the case where the size of the dictionary p is large. In this talk, we propose studying the estimation of the topic-word matrix under the assumption that the ordered entries of its columns rapidly decay to zero. This sparsity assumption is motivated by the empirical observation that the word frequencies in a text often adhere to Zipf’s law. We introduce a new spectral procedure for estimating the topic-word matrix that thresholds words based on their corpus frequencies, and show that its ℓ1-error rate under our sparsity assumption depends on the vocabulary size p only via a logarithmic term. Our error bound is valid for all parameter regimes and in particular for the setting where p is extremely large; Our procedure also empirically performs well relative to well-established methods when applied to a large corpus of research paper abstracts, as well as the analysis of single-cell and microbiome data where the same statistical model is relevant but the parameter regimes are vastly different.
The unreasonable effectiveness of negative association
Friday, November 17, 2023
Time: 2:00 p.m. to 3:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Subhro Ghosh, Assistant Professor, Department of Mathematics, Dept of Statistics and Data Science, faculty affiliate Institute of Data Science, National University of Singapore
Abstract: In 1960, Wigner published an article famously titled "The Unreasonable Effectiveness of Mathematics in the Natural Sciences”. In this talk we will, in a small way, follow the spirit of Wigner’s coinage, and explore the unreasonable effectiveness of negatively associated (i.e., self-repelling) stochastic systems far beyond their context of origin. As a particular class of such models, determinantal processes (a.k.a. DPPs) originated in quantum and statistical physics, but have emerged in recent years to be a powerful toolbox for many fundamental learning problems. In this talk, we aim to explore the breadth and depth of these applications. On one hand, we will explore a class of Gaussian DPPs and the novel stochastic geometry of their parameter modulation, and their applications to the study of directionality in data and dimension reduction. At the other end, we will consider the fundamental paradigm of stochastic gradient descent, where we leverage connections with orthogonal polynomials to design a minibatch sampling technique based on data-sensitive DPPs; with provable guarantees for a faster convergence exponent compared to traditional sampling. Principally based on the following works [1] Gaussian determinantal processes: A new model for directionality in data, with P. Rigollet, Proceedings of the National Academy of Sciences, vol. 117, no. 24 (2020), pp. 13207--13213 (PNAS Direct Submission) [2] Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGD, with R. Bardenet and M. Lin Advances in Neural Information Processing Systems 34 (Spotlight Paper at NeurIPS 2021)
Winter 2024
Sharper Risk Bounds for Statistical Aggregation
Friday, January 12, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Nikita Zhivotovskiy, Assistant Professor, Department of Statistics, University of California Berkeley
Abstract: In this talk, we revisit classical results in the theory of statistical aggregation, focusing on the transition from global complexity to a more manageable local one. The goal of aggregation is to combine several base predictors to achieve a prediction nearly as accurate as the best one, without assumptions on the class structure or target. Though studied in both sequential and statistical settings, they traditionally use the same “global” complexity measure. We highlight the lesser-known PAC-Bayes localization enabling us to prove a localized bound for the exponential weights estimator, and a deviation-optimal localized bound for Q-aggregation. Finally, we demonstrate that our improvements allow us to obtain bounds based on the number of near-optimal functions in the class, and achieve polynomial improvements in sample size in certain nonparametric situations. This is contrary to the common belief that localization doesn’t benefit nonparametric classes. Joint work with Jaouad Mourtada and Tomas Vaškevičius.
Phylogenomics: Some Identifiability Results
Friday, January 19, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Sebastien Roch, Professor, Department of Mathematics, University of Wisconsin-Madison
Abstract: The estimation of species phylogenies from genome-scale data is an important step in modern evolutionary studies. This estimation is complicated by the fact that genes evolve under biological processes that produce discordant trees. Such processes include horizontal gene transfer (HGT), gene duplication and loss (GDL), and incomplete lineage sorting (ILS), all of which can be modeled using random tree distributions. I will discuss recent results on the identifiability of these complex probabilistic models.
I will focus in particular on theoretical results for probabilistic models of HGT. Prior work has suggested the possibility of a “phase transition”, whereby reconstruction of the species tree may become significantly harder when the rate of transfer is high enough. I will report on recent work showing that, in fact, the species tree is identifiable for any rate of transfer, answering an open question in this area. Time permitting, I will also discuss the case of GDL.
No biology background will be assumed.
From 20GB to 100TB: a journey on the metagenomic road
Friday, February 9, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Zhong Wang, Computational Biologist, Genome Analysis Group Lead, Lawrence Berkeley National Lab
Abstract: Metagenomics has revolutionized our understanding of microbial functions, ecology, and evolution. Unraveling the complexity of environmental microbial communities demands numerous gigabases, or even terabases, of sequence data, thereby posing extraordinary computational challenges associated with data analysis. In this talk, I will chronicle our journey, navigating through the challenges presented by initially modest 20GB datasets to our current capability of handling substantial 100TB experiments. This journey entailed more than just augmenting storage or computational power; it also involved innovative thinking, experimentation with hardware scaling solutions, and the development of scalable software tools designed for immense datasets. By sharing our experiences, successes, and failures, this presentation aims to offer insights and strategies to fellow biologists and bioinformaticians navigating the rapidly expanding sea of metagenomic data.
BET and BELIEF
Friday, February 23, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Kai Zhang, Associate Professor, Department of Statistics and Operations Research, University of North Carolina, Chapel Hill
Abstract: We study the problem of distribution-free dependence detection and modeling through the new framework of binary expansion statistics (BEStat). The binary expansion testing (BET) avoids the problem of non-uniform consistency and improves upon a wide class of commonly used methods (a) by achieving the minimax rate in sample size requirement for reliable power and (b) by providing clear interpretations of global relationships upon rejection of independence. The binary expansion approach also connects the symmetry statistics with the current computing system to facilitate efficient bitwise implementation. Modeling with the binary expansion linear effect (BELIEF) is motivated by the fact that two linearly uncorrelated binary variables must be also independent. Inferences from BELIEF are easily interpretable because they describe the association of binary variables in the language of linear models, yielding convenient theoretical insight and striking parallels with the Gaussian world. With BELIEF, one may study generalized linear models (GLM) through transparent linear models, providing insight into how modeling is affected by the choice of link. We explore these phenomena and provide a host of related theoretical results. This is joint work with Benjamin Brown and Xiao-Li Meng.
Approximate Co-sufficient Sampling with Regularization
Friday, March 1, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Wanrong Zhu, Final-year PhD student, Department of Statistics, University of Chicago
Abstract: Goodness-of-fit (GoF) testing is ubiquitous in statistics and is applicable in many areas, for example, conditional independence testing, model selection, multiple testing, etc. We consider the problem of GoF testing for parametric models – testing whether observed data comes from some parametric null model. This testing problem involves a composite null hypothesis, due to the unknown values of the model parameters. In some special cases, co-sufficient sampling (CSS) can remove the influence of these unknown parameters via conditioning on a sufficient statistic—often, the maximum likelihood estimator (MLE) of the unknown parameters. However, many common parametric settings (including logistic regression) do not permit this approach, since conditioning on a sufficient statistic leads to a powerless test. The recent approximate co-sufficient sampling (aCSS) framework offers an alternative, replacing sufficiency with an approximately sufficient statistic (namely, a noisy version of the MLE). This approach recovers power in a range of settings where CSS cannot be applied, but can only be applied in settings where the unconstrained MLE is well-defined and well-behaved, which implicitly assumes a low-dimensional regime. In this talk, we extend aCSS to the setting of constrained and penalized maximum likelihood estimation, so that more complex estimation problems can now be handled within the aCSS framework, including examples such as mixtures-of-Gaussians (where the unconstrained MLE is not well-defined due to degeneracy) and high-dimensional Gaussian linear models (where the MLE can perform well under regularization, such as an ℓ1 penalty or a shape constraint).
Spring 2024
Large Language Models to understand biomedical text
Friday, April 5, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Online - talk will be presented on Zoom, registration is required to receive the link
Speaker: Yuan Luo, Director, Institute for Artificial Intelligence in Medicine - Center for Collaborative AI in Healthcare; Associate Professor of Preventive Medicine (Health and Biomedical Informatics), McCormick School of Engineering and Pediatrics
Abstract: Large Language Models such as transformer-based models have been wildly successful in setting state-of-the-art benchmarks on a broad range of natural language processing (NLP) tasks, including question answering (QA), document classification, machine translation, text summarization, and others. Recently, the release of OpenAI’s free tool ChatGPT demonstrated the ability of large language models to generate content, with anticipations on its possible uses and potential controversies. The ethical and acceptable boundaries of ChatGPT’s use in scientific writing remain unclear. I will talk about our research on exploring large language models, e.g., long-sequence transformers and GPT style models, in the clinical and biomedical domains. Our work examines the adaptability of these large language models to a series of clinical NLP tasks including clinical inferencing, biomedical named entity recognition, EHR based question answering, interoperability etc.
Surprises in binary linear classification
Friday, April 12, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Andrea Montanari, John D. and Sigrid Banks Professor in Statistics and Mathematics, Stanford University
Abstract: Machine learning calls into question our understanding of statistical methodology, both because of the new classes of statistical models used, and because of the new regimes and use cases. I will focus on the latter aspect by considering the (supposedly) well understood case of binary linear classification. I will discuss a certain number of phenomena that are not captured by classical statistical theory: interpolation, universality, data subsampling, tractability. High-dimensional asymptotics will be used to shed light on these behaviors.
t-SNE and Local 1D Structures
Friday, April 26, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Anna Ma, Assistant Professor, UC Irvine, Department of Mathematics.
Abstract: Data visualization is a vital task in data exploration, especially in the presence of large-scale data sets. Rudimentary approaches for data visualization, such as scatter plots, histograms, and pie charts, can only represent a small number (typically, 1-2) of features at a time. Furthermore, such methods often lack the sophistication to capture higher dimensional structures in their representations. Fortunately, new approaches to high-dimensional data visualization, such as the t-distributed stochastic neighbor embedding (t-SNE) algorithm, have been proposed in recent years. One of t-SNE’s more interesting properties is its tendency to preserve local linear data structures while successfully representing clusterable data. Despite its wide success, there is limited mathematical understanding of the algorithm. In this talk, we will discuss the t-SNE algorithm and present theoretical guarantees for t-SNE’s output to answer the question: does t-SNE preserve 1-dimensional curves?
The work presented is joint with Kat Dover and Roman Vershynin.
Audience Choice: Bayesian Workflow / Causal Generalization / Modeling of Sampling Weights
Friday, May 3, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Andrew Gelman, Professor of Statistics and Political Science, Columbia University
The audience is invited to choose among three possible talks:
Bayesian Workflow: The workflow of applied Bayesian statistics includes not just inference but also building, checking, and understanding fitted models. We discuss various live issues including prior distributions, data models, and computation, in the context of ideas such as the Fail Fast Principle and the Folk Theorem of Statistical Computing. We also consider some examples of Bayesian models that give bad answers and see if we can develop a workflow that catches such problems. For background, see here: http://www.stat.columbia.edu/~gelman/research/unpublished/Bayesian_Workflow_article.pdf
Causal Generalization: In causal inference, we generalize from sample to population, from treatment to control group, and from observed measurements to underlying constructs of interest. The challenge is that models for varying effects can be difficult to estimate from available data. We discuss limitations of existing approaches to causal generalization and how it might be possible to do better using Bayesian multilevel models. For background, see here: http://www.stat.columbia.edu/~gelman/research/published/KennedyGelman_manuscript.pdf and here: http://www.stat.columbia.edu/~gelman/research/published/causalreview4.pdf and here: http://www.stat.columbia.edu/~gelman/research/unpublished/causal_quartets.pdf
Modeling of Sampling Weights: A well-known rule in practical survey research is to include weights when estimating a population average but not to use weights when fitting a regression model—as long as the regression includes as predictors all the information that went into the sampling weights. But what if you don’t know where the weights came from? We propose a quasi-Bayesian approach using a joint regression of the outcome and the sampling weight, followed by poststratifcation on the two variables, thus using design information within a model-based context to obtain inferences for small-area estimates, regressions, and other population quantities of interest. For background, see here: http://www.stat.columbia.edu/~gelman/research/unpublished/weight_regression.pdf
Topic will be chosen live by the audience attending the talk.
T-Stochastic Graphs
Friday, May 10, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Karl Rohe, Professor of Statistics, University of Wisconsin–Madison
Abstract: Previous statistical approaches to hierarchical clustering for social network analysis all construct an "ultrametric" hierarchy. While the assumption of ultrametricity has been discussed and studied in the phylogenetics literature, it has not yet been acknowledged in the social network literature. We show that "non-ultrametric structure" in the network introduces significant instabilities in the existing top-down recovery algorithms. To address this issue, we introduce an instability diagnostic plot and use it to examine a collection of empirical networks. These networks appear to violate the "ultrametric" assumption. We propose a deceptively simple class of probabilistic models called T-Stochastic Graphs which impose no topological restrictions on the latent hierarchy. Perhaps surprisingly, this model generalizes the previous models. To illustrate this model, we propose six alternative forms of hierarchical network models and then show that all six are equivalent to the T-Stochastic Graph model. These alternative models motivate a novel approach to hierarchical clustering that combines spectral techniques with the well-known Neighbor-Joining algorithm from phylogenetic reconstruction. We prove this spectral approach is statistically consistent.
An Automatic Finite-Sample Robustness Check: Can Dropping a Little Data Change Conclusions?
Friday, May 17, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Tamara Broderick, Associate Professor, Department of Electrical Engineering and Computer Science, MIT
Abstract: Practitioners will often analyze a data sample with the goal of applying any conclusions to a new population. For instance, if economists conclude microcredit is effective at alleviating poverty based on observed data, policymakers might decide to distribute microcredit in other locations or future years. Typically, the original data is not a perfect random sample from the population where policy is applied -- but researchers might feel comfortable generalizing anyway so long as deviations from random sampling are small, and the corresponding impact on conclusions is small as well. Conversely, researchers might worry if a very small proportion of the data sample was instrumental to the original conclusion. So we propose a method to assess the sensitivity of statistical conclusions to the removal of a very small fraction of the data set. Manually checking all small data subsets is computationally infeasible, so we propose an approximation based on the classical influence function. Our method is automatically computable for common estimators. We provide finite-sample error bounds on approximation performance and a low-cost exact lower bound on sensitivity. We find that sensitivity is driven by a signal-to-noise ratio in the inference problem, does not disappear asymptotically, and is not decided by misspecification. Empirically we find that many data analyses are robust, but the conclusions of several influential economics papers can be changed by removing (much) less than 1% of the data.