2024-2025
Fall 2024
Mathematics in Scientific Machine Learning
Friday, October 4, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Rebecca Willett, Professor of Statistics and Computer Science & the Faculty Director of AI at the Data Science Institute, University of Chicago
Abstract: Artificial intelligence (AI) and machine learning (ML) are poised to revolutionize the pace and nature of scientific discovery. The widespread adoption of AI in the sciences has the potential to integrate scientific inquiry with modes of hypothesis generation, data analysis, experimental design, and simulation, transforming our capacity to address scientific problems that currently seem insurmountable. The mathematical foundations of AI and ML are crucial for high-quality, reproducible, AI-enabled scientific research. However, blindly applying AI and ML poses significant risks, such as the rapid acceleration of the “reproducibility crisis” in science. In this talk, I will discuss fundamental machine learning challenges and opportunities that are particularly relevant to scientific discovery, such as emulators, generative models, and inverse problems. These problems underscore the importance of incorporating mathematical and physical models as well as numerical algorithms into ML frameworks, highlighting exciting directions for future work.
This talk co-sponsored by the Department of Engineering Sciences and Applied Mathematics in McCormick
How to Detect Out-of-Distribution Data in the Wild?
Friday, October 11, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Sharon Y. Li, Assistant Professor in the Department of Computer Sciences, University of Wisconsin Madison
Abstract: When deploying machine learning models in the open and non-stationary world, their reliability is often challenged by the presence of out-of-distribution (OOD) samples. Since data shifts happen prevalently in the real world, identifying OOD inputs has become an important problem in machine learning. In this talk, I will discuss challenges, research progress, and opportunities in OOD detection. Our work is motivated by the insufficiency of existing learning objectives such as ERM --- which focuses on minimizing error only on the in-distribution (ID) data, but does not explicitly account for the uncertainty that arises outside ID data. To mitigate the fundamental limitation, I will introduce a new algorithmic framework, which jointly optimizes for both accurate classification of ID samples and reliable detection of OOD data. The learning framework integrates distributional uncertainty as a first-class construct in the learning process, thus enabling both accuracy and safety guarantees.
A holistic and critical look at language agents
Friday, October 18, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Yu Su, Assistant Professor, Department of Computer Science and Engineering, The Ohio State University
Abstract: How are the contemporary AI agents powered by LLMs different from those of the earlier generations? I argue that their most distinct trait is a new capability of using language as a vehicle of both 'thought' and communication, and therefore they are best called "language agents." I will describe a conceptual framework for these language agents, followed by a more in-depth discussion on several core competencies, including memory, planning, and tool use. I will conclude the talk with interesting future directions.
Subsampling for Big Data Regression with Measurement Constraints
Friday, October 25, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Lin Wang, Assistant Professor of Statistics, Purdue University
Abstract: Despite the availability of extensive data sets, it is often impractical to observe the labels for all data points due to various measurement constraints in many applications. To address this challenge, subsampling approaches can be employed to select a subset of design points from a large pool for observation, resulting in substantial savings in labeling costs. In this presentation, I will introduce our recent research on computationally feasible subsampling techniques. Our primary focus is on regression with labeled data, which includes linear regression, ridge regression, and nonparametric additive regression. For these regression tasks, we have developed sampling approaches that aim to minimize the mean squared error in estimations and predictions. We will demonstrate the effectiveness of our proposed approaches through theoretical analysis and extensive numerical results.
Autonomous Learning: Unifying OOD Detection and Continual Learning
Friday, November 1, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Bing Liu, Distinguished Professor and Peter L. and Deborah K. Wexler Professor of Computing at the University of Illinois Chicago
Abstract: Continual learning (CL) focuses on incrementally learning a sequence of tasks, with class incremental learning (CIL) being one of the most challenging settings. This talk begins by presenting a theoretical study of the CIL problem. The key result is that the necessary and sufficient conditions for effective CIL are strong within-task prediction and reliable out-of-distribution (OOD) detection. The theory unifies CIL and OOD detection, which are regarded as two completely different problems. Building on the theory, new CIL methods have been developed, which significantly outperform existing baselines. However, traditional CIL operates in a closed-world context. We then extend the theory to the open world—where unknown and out-of-distribution objects are encountered—leading to the learning paradigm of open-world CIL, or open-world continual learning (OWCL), enabling autonomous learning. In the last part of the talk, I will discuss challenges in OWCL and present a prototype system that learns on the fly continually and autonomously after deployment.
Quantum Computation and Statistics
Friday, November 15, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Yazhen Wang, Department of Statistics, University of Wisconsin
Abstract: Quantum computation and quantum information are of great current interest across various fields, including computer science, mathematics and statistics, physical sciences and engineering. As the theory of quantum physics is fundamentally stochastic, quantum computation and quantum information are inherently infused with elements of randomness and uncertainty. Consequently, quantum algorithms are random in nature. This highlights the important role for statistics to play in the realm of quantum computation, which in turn offers great potential to revolutionize computational statistics. In this talk, I will provide an overview of quantum computation and statistics, covering the fundamental concepts and exploring quantum advantage along with the role of statistics and the implications for statistics.
Structure-driven design of reinforcement learning algorithms: a tale of two estimators
Friday, November 22, 2024
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Wenlong Mou, Assistant Professor of Statistical Sciences, University of Toronto
Abstract: Reinforcement learning (RL) offers a flexible framework for sequential decision-making in uncertain environments, and its success heavily depends on efficiently learning value functions. Over the years, a diverse range of RL algorithms has been proposed, but at their core, two foundational principles stand out: to solve the Bellman fixed-point equations (known as ``bootstrapping methods''), or to average the rollout rewards. Despite their success, finding the optimal trade-off between these principles in practical applications remains elusive. Current theoretical guarantees -- either worst-case or asymptotic -- often fall short of providing actionable insights.
In this talk, I will discuss recent advances in methods that optimally reconcile bootstrapping and rollout for policy evaluation. The bulk of this talk will focus on a new class of estimators that strikes an optimal balance between temporal difference learning and Monte Carlo methods. Through the statistical lens, I will highlight why the local structure of the underlying Markov chain determines the fundamental complexity for estimation, and how our estimator adapts to these structures. Extending this perspective to continuous-time RL, I will also explore how the elliptic structure of diffusion processes provides key insights for making algorithmic choices
Winter 2025
Leveraging multi-study, multi-outcome data to improve external validity and efficiency of clinical trials for managing schizophrenia
Friday, January 17, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Caleb H. Miles, Assistant Professor of Biostatistics, Columbia University Mailman School of Public Health
Abstract: As data sources have become more plentiful and readily accessible, the practice of data fusion has become increasingly ubiquitous. However, when the focus is on a causal effect on a particular outcome, a major limitation is that this outcome may not be available in all data sources. In fact, different randomized experiments or observational studies of a common exposure will often focus on potentially related, yet distinct outcomes. One such example is the Database of Cognitive Training and Remediation Studies (DoCTRS), which consists of several randomized trials of the effect of cognitive remediation therapy on various outcomes among patients with schizophrenia. We develop causally principled methodology for fusing data sets when multiple outcomes are observed across studies, which leverages outcomes of secondary interest as informative proxies for the missing outcome of primary interest, thereby maximizing power and efficiency by making full use of the available data. As this methodology relies on a key transportability assumption, we also develop methods to assess the degree of sensitivity to violations of this assumption. We apply this methodology to data from the DoCTRS trials to make improved causal inferences about the effectiveness of cognitive remediation therapy on cognition among patients with schizophrenia.
The Role of AI in Scientific Discovery: Opportunities and Limitations
Friday, January 24, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Xiangliang Zhang, Leonard C. Bettex Collegiate Professor of Computer Science, University of Notre Dame
Abstract: Artificial Intelligence (AI) is reshaping the landscape of scientific discovery, enabling breakthroughs across diverse fields. However, when these AI tools are applied to scientific problems, gaps and mismatches often arise. The inherent uncertainty in scientific phenomena, coupled with issues like data quality, biases, and interpretability, poses significant challenges. This talk will discuss the transformative potential of AI in scientific discovery, focusing on its applications in predictive modeling, generative tasks, optimization strategies, and literature analysis. Examples will include AI models ranging from traditional neural networks to large language models (LLMs). At the same time, their limitations will be critically examined, calling for collaboration between the AI and scientific communities to address these challenges and unlock AI’s full potential in advancing scientific discovery.
Tensor Time Series: Factor Modeling and Deep Neural Networks
Friday, January 31, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Yuefeng Han, Assistant Professor, Department of Applied and Computational Mathematics and Statistics, University of Notre Dame
Abstract: The analysis of tensors (multi-dimensional arrays) has become a vital area in modern statistics and data science, driven by advancements in scientific research and data collection. High-dimensional tensor data arise in diverse applications such as economics, genetics, microbiome studies, brain imaging, and hyperspectral imaging. These tensors are often high-dimensional and high-order, yet key information typically resides in reduced-dimensional subspaces governed by structural properties. This talk explores novel methodologies and theories for tensor time series analysis.
The presentation consists of two parts. The first part introduces a factor modeling framework for high-dimensional tensor time series, leveraging a structure similar to CP tensor decomposition. We propose a computationally efficient estimation procedure incorporating a warm-start initialization and an iterative simultaneous orthogonalization scheme. The algorithm achieves $\epsilon$-accuracy within $\log\log(1/\epsilon)$ iterations. Additionally, we establish inferential results, demonstrating consistency and asymptotic normality under relaxed assumptions. The second part integrates tensor factor models with deep neural networks. Specifically, a Tucker-type low-rank tensor structure is employed as a tensor-augmentation module in neural networks. Extensive experiments demonstrate the integration of this module into transformers and temporal neural networks for tensor time series prediction and tensor-on-tensor regression. The results highlight significant performance improvements, underscoring its potential for advancing time series forecasting.
AI for Nature: From Science to Impact
Friday, February 7, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Tanya Berger-Wolf, Professor, Computer Science and Engineering and Director, Translational Data Analytics Institute
Abstract: Computation has fundamentally changed the way we study nature. New data collection technologies, such as GPS, high-definition cameras, autonomous vehicles under water, on the ground, and in the air, genotyping, acoustic sensors, and crowdsourcing, are generating data about life on the planet that are orders of magnitude richer than any previously collected. Yet, our ability to extract insight from these data lags substantially behind our ability to collect it.
The need for understanding is more urgent ever and the challenges are great. We are in the middle of the 6th extinction, losing the planet's biodiversity at an unprecedented rate and scale. In many cases, we do not even have the basic numbers of what species we are losing, which impacts our ability to understand biodiversity loss drivers, predict the impact on ecosystems, and implement policy.
The talk will discuss how AI can turn these data into high resolution information source about living organisms, enabling scientific inquiry, conservation, and policy decisions. It will introduce a new field of science, imageomics, and present a vision and examples of AI as a trustworthy partner both in science and biodiversity conservation, discussing opportunities and challenges.
Towards Data-efficient Training of Large Language Models (LLMs)
Friday, February 14, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Virtual talk, registration required
Speaker: Baharan Mirzasoleiman, Assistant Professor, Computer Science Department, UCLA
Abstract: High quality data is crucial for training LLMs with superior performance. In this talk, I will present two theoretically-rigorous approaches to find smaller subsets of examples that can improve the performance and efficiency of training LLMs. First, I will present a one-shot data selection method for supervised fine-tuning of LLMs. Then, I'll talk about an iterative data selection strategy to pretrain or fine-tune LLMs on imbalanced mixtures of language data. I'll conclude by showing empirical results confirming that the above data selection strategies can effectively improve the performance of various LLMs during fine-tuning and pretraining.
Knowledge-Guided Machine Learning for Scientific Discovery: Challenges and Opportunities
Friday, February 21, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Virtual talk, registration required
Speaker: Xiaowei Jia, Assistant Professor, Department of Computer Science, University of Pittsburgh
Abstract: Data science and machine learning (ML) models, which have found tremendous success in several commercial applications where large-scale data is available, e.g., computer vision and natural language processing, has met with limited success in scientific domains. Traditionally, physics-based models of dynamical systems are often used to study engineering and environmental systems. Despite their extensive use, these models have several well-known limitations due to incomplete or inaccurate representations of the physical processes being modeled. Given rapid data growth due to advances in sensor technologies, there is a tremendous opportunity to systematically advance modeling in these domains by using machine learning methods. However, capturing this opportunity is contingent on a paradigm shift in data-intensive scientific discovery since the “black box” use of ML often leads to serious false discoveries in scientific applications. Because the hypothesis space of scientific applications is often complex and exponentially large, an uninformed data-driven search can easily select a highly complex model that is neither generalizable nor physically interpretable, resulting in the discovery of spurious relationships, predictors, and patterns. This problem becomes worse when there is a scarcity of labeled samples, which is quite common in science and engineering domains.
My work aims to build the foundations of knowledge-guided machine learning (KGML) by exploring several ways of bringing scientific knowledge and machine learning models together. In particular, we discuss gaps and opportunities in scientific discovery and show the effectiveness of KGML in multiple applications of great societal and scientific relevance. My work also has the potential to greatly advance the pace of discovery in a number of scientific and engineering disciplines where physics-based models are used, e.g., hydrology, agriculture, climate science, materials science, power engineering and biomedicine.
Spring 2025
Foundation Models and Generative AI for Medical Imaging Segmentation in Ultra-Low Data Regimes
Friday, April 11, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Pengtao Xie, Assistant Professor, Electrical and Computer Engineering, University of California, San Diego
Abstract: Semantic segmentation of medical images is pivotal in disease diagnosis and treatment planning. While deep learning has excelled in automating this task, a major hurdle is the need for numerous annotated masks, which are resource-intensive to produce due to the required expertise and time. This scenario often leads to ultra-low data regimes where annotated images are scarce, challenging the generalization of deep learning models on test images. To address this, we introduce two complementary approaches. One involves developing foundation models. The other involves generating high-fidelity training data consisting of paired segmentation masks and medical images. In the former, we develop a bi-level optimization based method which can effectively adapt the general-domain Segment Anything Model (SAM) to the medical domain with just a few medical images. In the latter, we propose a multi-level optimization based method which can perform end-to-end generation of high-quality training data from a minimal number of real images. On eight segmentation tasks involving various diseases, organs, and imaging modalities, our methods demonstrate strong generalization performance in both in-domain and out-of-domain settings. Our methods require 8-12 times less training data than baselines to achieve comparable performance.
Neural Collapse in AI Training
Friday, April 25, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: X.Y. Han, Assistant Professor of Operations Management, Booth School of Business, University of Chicago
Abstract: In the performance-dominated landscape of AI development, systematic understanding is difficult: Any architecture is fair game as long as it can climb the leaderboard. Yet, even among the overwhelming multiformity of neural networks, one motif remains: data is transformed layer-by-layer and iteration-by-iteration into high-dimensional representations with which predictions are eventually made. Thus, examining the geometry of these representations offers a powerful lens for the analysis of AI. Neural Collapse is a striking example, evident late in the training of predictive AI: Originally observed in classification networks during a 'Terminal Phase of Training' (TPT) – where training error vanishes but loss continues decreasing – Neural Collapse reveals a fundamental inductive bias. It encompasses four interconnected geometric regularities in the final layer representations: (NC1) Within-class variability of activations collapses towards zero as they converge to their class means; (NC2) These class means arrange themselves into a maximally separated, symmetric structure (a Simplex Equiangular Tight Frame); (NC3) The classifier vectors align with these class means in a self-dual configuration; (NC4) The model's decision mechanism effectively simplifies to Nearest Class-Center rule. This emergent geometric simplicity is linked to improved generalization, robustness, and interpretability. I will discuss the core Neural Collapse phenomenon, its theoretical underpinnings, and recent updates and extensions since its discovery.
A Bayesian semi-parametric model for functional near-infrared spectroscopy data
Friday, May 2, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Timothy D. Johnson, Professor, Department of Biostatistics, School of Public Health, University of Michigan
Abstract: Functional near-infrared spectroscopy (fNIRS) is a relatively new neuroimaging technique. It is a low cost, portable, and non-invasive method to monitor brain activity. Similar to fMRI, it measures changes in the level of blood oxygen in the brain. Its time resolution is much finer than fMRI, however its spatial resolution is much courser—similar to EEG or MEG. fNIRS is finding widespread use on young children that have trouble staying still in the MRI magnet and it can be used in situations where fMRI is contraindicated—such as chochlear implant patients. In this talk, I propose a fully Bayesian semi-parametric model to analyze fNIRS data. The hemodynamic response function is modeled with the canonical HRF. The model error and the autoregressive process vary with time and are modeled in the dynamic linear model framework. The low frequency drift is modeled non-parameterically with a variable B-spline model (both locations and number of knots are allowed to vary). Although motion is not as big an issue as in fMRI, it can still cause huge inferential bias and poor statistical properties if not handled appropriately. The variable B-spline model not only models the low frequency drift, but will regress out motion artifacts as well. Most methods require motion to be removed prior to statistical analysis except one, which I refer to as the ARIRLS model. Via simulation studies, I show that this Bayesian model easily handles motion artifacts and results in better statistical properties than the AR-IRLS model. I then show its performance on real data.
Learning Multi-Index Models
Friday, May 9, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Ilias Diakonikolas, Sheldon B. Lubar Professor, Department of Computer Sciences, University of Wisconsin, Madison
Abstract: Multi-index models (MIMs) are functions that depend on the projection of the input onto a low-dimensional subspace. These models offer a powerful framework for studying various machine learning tasks, including multiclass linear classification, learning intersections of halfspaces, and more complex neural networks. Despite extensive investigation, there remains a vast gap in our understanding of the efficient learnability of MIMs.
In this talk, we will survey recent algorithmic developments on learning MIMs, focusing on methods with provable performance guarantees. In particular, we will present a robust learning algorithm for a broad class of well-behaved MIMs under the Gaussian distribution. A key feature of our algorithm is that its running time has fixed-degree polynomial dependence on the input dimension. We will also demonstrate how this framework leads to more efficient and noise-tolerant learners for multiclass linear classifiers and intersections of halfspaces.
Time permitting, we will highlight some of the many open problems in this area.
The main part of the talk is based on joint work with I. Iakovidis, D. Kane, and N. Zarifis.
Data Augmentation for Graph Regression
Friday, May 23, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Ruan Conference Room – lower level (Chambers Hall 600 Foster Street)
Speaker: Meng Jiang, Associate Professor, Department of Computer Science and Engineering, University of Notre Dame
Abstract: Graph regression plays a key role in materials discovery by enabling the prediction of numerical properties of molecules and polymers. However, graph regression models often rely on training sets with only a few hundred labeled examples, and these labels are typically imbalanced. While a large number of unlabeled examples are available, they are often drawn from diverse domains, making them less effective for improving target label predictions. In machine learning, data augmentation refers to techniques that increase the size of the training set by generating slightly modified or synthetic versions of existing data. These methods are simple yet effective. In this talk, I will introduce three graph data augmentation techniques tailored for supervised learning, imbalanced learning, and transfer learning in graph regression tasks. The first technique leverages the mutual enhancement between model rationalization and data augmentation, improving both accuracy and interpretability in molecular and polymer property prediction. This approach demonstrates that graph data augmentation can be effectively performed in latent spaces. The second technique generates representations of additional data points with underrepresented labels to balance the training set. The third technique introduces a graph diffusion transformer (Graph DiT) that facilitates data-centric transfer learning, addressing the limitations of self-supervised methods when dealing with unlabeled graph data. Graph DiT integrates multiple properties such as synthetic score and gas permeability as condition constraints into diffusion models for multi-conditional polymer generation. Lastly we will discuss foundation model approaches for materials discovery.
LLM-Enhanced, Theme-Focused Science Discovery: A Retrieval and Structuring Approach
Friday, May 30, 2025
Time: 11:00 a.m. to 12:00 p.m. central time
Location: Virtual talk, registration required (link below)
Speaker: Jiawei Han, Michael Aiken Chair Professor, Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign
Abstract: Large Language Models (LLMs) may bring unprecedent power in scientific discovery. However, current LLMs may still encounter major challenges for effective scientific exploration due to their lack of in-depth, theme-focused data and knowledge. Retrieval augmented generation (RAG) has recently become an interesting approach for augmenting LLMs with grounded, theme-specific datasets. We discuss the challenges of RAG and propose a retrieval and structuring (RAS) approach, which enhances RAG by improving retrieval quality and mining structures (e.g., extracting entities and relations and building knowledge graphs) to ensure its effective integration of theme-specific data with LLM. We show the promise of this approach at augmenting LLMs and discuss its potential power for LLM-enabled science exploration.