Kinesiophobia, defined as an excessive and irrational fear of physical movement and activity, could be a barrier to physical activity practice, especially in the ageing population and in people living with Parkinson’s disease. However, prevalence of kinesiophobia and its predictors remain to be characterized in these populations. For this project, a MITACS intern will work on the analysis of longitudinal surveys conducted in the lab about kinesiophobia in the elderly population and in Parkinson’s disease. The project will include data cleaning, descriptive analysis, data visualization, statistical analysis and results interpretation.
Research area, student roles & skills
Research area: Our research focusses on the biopsychosocial effects of healthy and pathological aging on locomotion and physical activity practice, as well as their involvement for the development of technology-based interventions aimed at promoting an active lifestyle. Areas of expertise include neuroscience, motor control, kinesiophobia, movement disorders, aging and physical activity, and techniques currently used include virtual reality, actigraphy, movement kinematics, exoskeletons and surveys. Our studies are conducted on a variety of populations, including healthy elderly and individuals living with Parkinson’s disease.
Student roles: Perform data cleaning & coding Perform descriptive analysis on the data Create charts, graphs and tables to visually represent the data Analyze statistically the data (examine relationships between variables, conduct hypothesis testing, correlations, regression analysis and other statistical models, as needed) Conduct a qualitative analysis as needed Interpret the results Contribute to the report of the results (publication, abstracts, presentations, etc.) Collaborate with the principal investigator and her team to complete the project. Opportunity to contribute to data collection of other undergoing projects.
Skills required: Interest in clinical research and health science. Experience in epidemiology, statistical analysis or survey analysis is an asset.
2. Analyzing the health risk from compound extreme weather by explainable AI
Supervisor: Mengxin Pan
University: Simon Fraser University (Burnaby campus)
Climate change is escalating the frequency and intensity of extreme weather events, exacerbating health risks worldwide. Yet, most existing research treats each hazard independently. Compound extreme events (CEEs) – combinations of multiple extremes occurring together or in sequence – represent a critical but under-studied phenomenon. Similar to co-morbid illnesses compounding stress on a patient, CEEs could also amplify health impacts beyond the sum of individual events. However, public health research has paid limited attention to these synergistic health impacts of extreme weather to date, let alone considering the influence of climate change.
In this project, we aim to study the impact of compound extreme weather events on human health under global warming via state-of-the-art interpretable machine learning models. A case study will be focused, i.e., the compound impact of heatwaves, droughts, and air pollution on children’s health. Specifically, we propose an interpretable machine learning-based methodology designed explicitly for climate-health analysis: 1) To identify Compound Extreme Event-Health Linkage after building up the global database with various types of CEEs and health outcomes. 2) To develop interpretable machine learning models to predict health outcomes by compounded environmental conditions and explain the underlying mechanism. 3) To assess how CEE-related health risks might evolve as the climate continues to warm via the future projection by climate models.
Research area, student roles & skills
Research area: My research combines advanced data mining techniques with climate dynamics to address real-world problems, with two interconnected objectives. First, advancing the fundamental insight of climate and weather extremes, with a specific interest in Atmospheric Rivers and their changes in a warming climate. Second, bridging the physical and social dimensions in facing climate change challenges. Specifically, I am interested in interdisciplinary collaborative studies about the climate impact in a warming world.
Student roles: 1) Preprocessing and visualization of various health and climate data; 2) Correlation analysis of health and climate data; 3) Machine learning modeling.
Skills required: Data analysis and visualization; programming; statistics
3. Assessing the accuracy of predictive density estimates
Reporting on the accuracy of estimators is a fundamental task of the statistician and of a research scientist who wishes to quantify the degree of uncertainty. From both a Bayesian and frequentist point of view, it is either necessary or beneficial to calibrate such a report of accuracy after having viewed the data. Manifestations of such reports include standard errors and posterior standard deviations (point estimation), as well as confidence level estimates (interval estimation). This research project will address predictive density estimates and reports on their accuracy based on the data. A predictive density estimate $\hat{q}$ is a density based on $X \sim p_{\theta}$ used to estimate of a density $q_{\theta}$ for a future or missing $Y$. Given a loss function based on a measure of divergence between $\hat{q}$ and $q_{\theta}$, such as Kullback-Leibler or Hellinger, one objective is to seek for optimality through a Bayesian approach or based on frequentist risk (e.g., minimaxity). The central question in this project will be to efficiently estimate the loss based on the data $X$. Bayesian solutions will be explored and evaluated based on frequentist risk.
Research area, student roles & skills
Research area: Statistical inference, Bayesian analysis, Predictive density estimation, Applied probability
Student roles: The student will allocate some time on background material and implement various choices of estimators in the described of loss estimation. Along with several other students and myself, we will try to obtain analytical comparisons and we will report on numerical evaluations illustrating the strengths and weakeness of various estimators.
Skills required: Good knowledge in mathematical analysis, Bayesian analysis, statistical inference and probability. Expertise in mathematical programming.
4. Bayesian Causal Inference Methods for Measurement Error in Electronic Health Records
Supervisor: Sumeet Kalia
University: University of Manitoba (Winnipeg campus)
Electronic health records (EHRs) contain rich longitudinal data on patient diagnoses, medications, and laboratory values that can inform clinical decision-making. However, clinical measurements recorded in these repositories are frequently error-prone. Blood glucose readings may be imprecise due to instrument calibration, disease diagnoses may be misclassified due to incomplete documentation, and recorded outcomes may depend systematically on treatment status. Ignoring such measurement error in causal analyses leads to biased treatment effect estimates, potentially resulting in flawed clinical recommendations.
This project develops a Bayesian imputation framework for estimating average treatment effects from error-prone observational health data. The proposed methodology addresses two intertwined challenges simultaneously: the fundamental missing data problem in causal inference, where only one of two potential outcomes is ever observed per patient, and measurement error in treatments, outcomes, and confounders. Under the potential outcomes framework, the unobserved counterfactual outcome is treated as a missing parameter and recovered from its posterior predictive distribution. The joint model specifies a flexible bivariate outcome model conditional on latent true covariate values, a classical measurement error model linking observed surrogate measurements to their true counterparts, and a differential misclassification component for outcomes that allows measurement error to depend on treatment status. MCMC sampling iterates between imputing missing potential outcomes, imputing latent true covariate values, and updating model parameters, with average treatment effects computed directly from posterior draws.
The project will include Monte Carlo simulation studies evaluating finite-sample bias, variance, and credible interval coverage across a range of measurement error mechanisms. Methods will be applied to Canadian primary care EHR data from the Canadian Primary Care Sentinel Surveillance Network to examine the effect of glucose-lowering medications on urinary tract infection risk among patients with type 2 diabetes. An open-source R package implementing the proposed methods will be developed and disseminated to applied researchers.
Research area, student roles & skills
Research area: My research focuses on developing novel biostatistical methods for epidemiology and health services research, with particular emphasis on causal inference, pharmacoepidemiology, machine learning, and electronic health records. As an Assistant Professor in the Department of Statistics at the University of Manitoba, I have co-authored more than 35 publications in collaboration with clinicians, biostatisticians, and epidemiologists, and have supervised undergraduate and graduate students across a range of statistical topics including causal inference, measurement error, experimental design, Bayesian methods, infectious disease modelling, epidemiological surveillance studies.
Student roles: The student will play an important role in all phases of the project. They will begin with a thorough literature review covering Bayesian causal inference, measurement error methods, and electronic health records applications. The student will then design and implement Monte Carlo simulation studies in R to evaluate the finite-sample properties of the proposed estimators under varying measurement error mechanisms. They will assist with the development and documentation of an open-source R package implementing the proposed methodology. Finally, the student will contribute to writing a manuscript suitable for submission to a peer-reviewed statistics or biostatistics journal, gaining experience in scientific communication and collaborative research.
Skills required: The ideal candidate has a strong background in mathematical statistics including likelihood theory, regression, and asymptotic methods. Familiarity with Bayesian concepts is important, including prior and posterior distributions, MCMC sampling, and posterior predictive inference. Prior exposure to causal inference methods such as propensity scores or the potential outcomes framework is an asset but not required. Strong programming skills in R are preferred, including experience with simulation studies and statistical modeling. The ability to work independently, communicate technical results clearly, and collaborate across disciplines are also valued for this position.
5. Bayesian Extreme Value Analysis of Sea Ice Concentration
Supervisor: Xiaoting Li
University: University of Manitoba (Winnipeg campus)
Sea ice concentration (SIC) is the percentage of the ocean surface covered by sea ice. Persistent and record-low SIC values provide important evidence of structural changes in the climate system and have substantial ecological and socio-economic consequences.
This project aims to develop a Bayesian spatio-temporal extreme-value model to analyze extreme low-SIC events. Rather than focusing only on average sea ice trends, the project will examine the lower tail of the SIC distribution, quantify the frequency and severity of unusually low sea ice conditions, and study how these extremes evolve across space and time.
The analysis will use the NOAA/NSIDC Climate Data Record of Passive Microwave Sea Ice Concentration, which provides a consistent satellite-based record of sea ice concentration on a regular spatial grid. This dataset is well suited for statistical modeling because it provides long-term, gridded observations of sea ice concentration at daily and monthly time scales. The project will focus on the Arctic region, with possible emphasis on selected subregions where extreme sea ice decline is especially relevant.
Research area, student roles & skills
Research area: My research expertise is in dependence modeling with copulas, extreme-value theory, and spatio-temporal statistics with applications to risk management in finance, insurance, and environmental science.
Student roles: The student will work with the NOAA/NSIDC sea ice concentration dataset and gain hands-on experience in managing, processing, and analyzing large environmental datasets. A key component of the project will be learning how to work with raster data, including reading gridded climate data, extracting spatial information, and organizing the data into a format suitable for statistical analysis.
The student will begin with exploratory data analysis to examine seasonal variation, long-term trends, spatial patterns, and the occurrence of extreme low-SIC events. This stage will help the student understand the main features of the data and make informed choices about subsequent statistical modeling.
Building on this foundation, the student will learn the basic ideas of extreme value analysis, including threshold exceedance methods, return levels, and uncertainty quantification. The student will then help implement a Bayesian hierarchical model in Stan, gaining experience with Bayesian computation and posterior inference.
Depending on progress, the project may be extended to spatial or spatio-temporal dependence modeling. In this case, the student may investigate how extreme low-SIC events are correlated across nearby grid cells or regions, and whether this dependence changes over time.
Skills required: Coursework in probability, statistical inference, regression modeling, and computational statistics is necessary. Prior experience with spatial data, Bayesian statistics, extreme value theory, or Stan is not required, but would be an asset.
6. Cancer Subtype Detection via Data Depth Methods
Supervisor: Xiaoping Shi
University: University of British Columbia (Kelowna campus)
Cancer subtype detection is essential for precision medicine but is challenged by high-dimensional, heterogeneous, and noisy biomedical data. This project develops a statistical learning framework based on data depth and related distributional methods to identify clinically meaningful subtypes. By measuring centrality and structural variation in the data, the approach aims to provide robust classification without relying on strong parametric assumptions. The method will be applied to genomic or clinical datasets to detect subgroup structure, improve interpretability, and enhance the reliability of subtype identification in complex cancer data.
Research area, student roles & skills
Research area: My current areas of expertise include data sharpening, data depth, Gaussian graph models, Energy-based models, change point analysis, saddlepoint approximation, inverse moment approximation, two-sample comparison, and clustering analysis. I view my research as having three major branches: Applied Statistics, Computational Statistics, and Theoretical Statistics. In Applied Statistics, I research methods to model real data. In Computational Statistics, I develop new methods and implement them in software. In Theoretical Statistics, I focus on studying and developing the mathematical foundations of statistics. These three branches can be mixed.
Student roles: Understand references, test the performance of current methods using simulated data and new data collected, and propose new methods when old ones fail. Attend weekly group meetings and report on progress. In the first month, test the performance of current methods. In the second month, propose new methods where possible and conduct simulation tests. In the third month, provide detailed data analysis. Present at group meetings and seminars/conferences when the opportunity arises and write drafts for submission to journals.
Skills required: Prerequisites include a solid grasp of relevant literature, statistical inference, and data analysis/simulation coding, alongside strong presentation skills.
7. Classification of Sleep Apnea from hidden Semi‑Markov models
This project aims to develop an accurate algorithm for automated sleep state classification and obstructive sleep apnea detection. Our approach is novel in treating polysomnography (PSG) signals as functional data and using switching state space models.
Obstructive sleep apnea (OSA) arises when the upper airway repeatedly narrows or collapses during sleep, leading to disrupted breathing, intermittent drops in oxygen saturation, and low quality sleep. OSA affects a substantial proportion of people and is linked to serious cardiovascular and cognitive consequences, so timely diagnosis is essential. Most of the sleep apnea studies were done for adults, but OSA can also affect children, so we analyze the data from a sleep study performed for 41 patients 2-12 years old.
Manual scoring of PSG data is reliable but remains slow, labor intensive, and subjective. This project focuses on developing an algorithm capable of accurately classifying sleep states and detecting episodes of OSA. To represent the time spent in each sleep state, we use a switching state-space model including hidden semi-Markov chains. We design an Expectation Maximization algorithm that include a forward-backward algorithm for hidden semi-Markov chains to estimate the model parameters. We identify the states of the hidden semi-Markov chain and based on them classify the sleep states and determine sleep microstructures used for OSA diagnostic.
Research area, student roles & skills
Research area: My area of research includes applications of stochastic processes in statistical machine learning. I am interested in finding new clustering and classification methods for functional data. I am considering both non-parametric approaches and mixture models.
Student roles: The student will do an extensive literature review about hidden Markov and semi-Markov chains. Then the student will participate in the design and the computer implementation of the methodology. The designing part includes constructing the models and developing the procedure for updating the parameter estimations in the EM algorithm. The computer implementation includes data cleaning, programming the classification algorithms, and interpretation of the results.
Skills required: Good programming skills in at least one of the following languages: R, Python, Matlab, or C. Statistical/Machine Learning background in: Probability theory, Statistical estimation: EM algorithm, Regression, Multivariate analysis or Machine learning/ Artificial Intelligence
8. Conformal selection with batches for AI uncertainty quantification
Many decision-making processes—such as hiring or drug discovery—begin with a critical screening step. Before investing in costly evaluations, organizations rely on machine learning models to identify a small set of promising candidates from a large pool. The challenge lies in selecting candidates whose true outcomes exceed a desired threshold, all while accounting for predictive uncertainty.
Conformal selection offers a principled solution, enabling candidate shortlisting with finite-sample guarantees on the false selection rate. However, in real-world applications, predictions are often made on batches of test points, where decisions depend on properties of the batch as a whole. This includes tasks like constructing simultaneous prediction sets with guaranteed coverage or selecting data points that meet a specified condition while controlling false discoveries.
Standard methods struggle in this batch setting. This project aims to develop a general framework for predictive inference on functions of batches of test points. Our approach, called batch predictive inference (batch PI), delivers distribution-free coverage guarantees under the assumption of exchangeability between calibration and test data.
This advancement promises more robust and flexible decision-making in settings where conclusions depend on statistics computed over multiple uncertain outcomes.
Research area, student roles & skills
Research area: Statistical machine learning; high-dimensional statistics; statistical computing; biomedical, biochemical and industrial data science; applications in drug discovery
Student roles: In this project, the student will (1) study the foundational concept of multiple testing, BH-procedure, PRDS, conformal prediction, weighted likelihood, conformal selection; (2) to develop the proposed method and show theoretical guarantee holds under mild exchangeability conditions; (3) to demonstrate the empirical performance of the method via simulations, and apply to drug discovery datasets.
Skills required: The student will deepen their expertise in statistical machine learning, focusing on conformal prediction, selection, multiple testing, and uncertainty quantification. This work will contribute to research output and build a strong theoretical foundation for future academic pursuits. Additionally, the student will apply the method to real-world drug discovery applications.
The student will require strong analytical, quantitative and programming skills. I expect background (i.e., at least one undergraduate course) in probability, statistics, machine learning, optimization, computing, algorithms. Knowledge of drug discovery is useful but not required.
9. Continuous-Time Causal Inference Methods for Electronic Health Records with Irregular Observation Patterns
Supervisor: Sumeet Kalia
University: University of Manitoba (Winnipeg campus)
Primary care electronic health records (EHRs) are rich but complex data sources for studying treatment effectiveness, and standard causal inference methods fail to adequately address three interconnected challenges inherent in these data: informative and irregular patient visits, measurement error in confounders and outcomes, and time-dependent confounding. This project develops a unified continuous-time causal inference framework that simultaneously corrects for all three challenges by extending Marginal Structural Models using inverse probability weighting with marked point processes, avoiding the bias introduced by discretizing follow-up into fixed intervals. Measurement error correction is addressed through a Bayesian imputation framework under the potential outcomes paradigm, accommodating both differential and non-differential misclassification, with computationally efficient case-base sampling incorporated to make weight estimation tractable in large EHR datasets. The project proceeds in three phases: theoretical development of identifiability conditions and asymptotic properties, comprehensive Monte Carlo simulation studies evaluating finite-sample performance under realistic EHR scenarios, and empirical validation using Canadian primary care data to examine the effect of glucose-lowering medications on urinary tract infection risk among patients with type 2 diabetes, benchmarking results against published randomized trial estimates. Expected outputs include peer-reviewed methodological publications and an open-source R package making these advanced methods accessible to applied researchers.
Research area, student roles & skills
Research area: My research focuses on developing novel biostatistical methods for epidemiology and health services research, with particular emphasis on causal inference, pharmacoepidemiology, machine learning, and electronic health records. As an Assistant Professor in the Department of Statistics at the University of Manitoba, I have co-authored more than 35 publications in collaboration with clinicians, biostatisticians, and epidemiologists, and have supervised undergraduate and graduate students across a range of statistical topics including causal inference, measurement error, experimental design, Bayesian methods, infectious disease modelling, epidemiological surveillance studies.
Student roles: The student will play a central role in all phases of the project. They will begin with a thorough literature review covering continuous-time causal inference, marginal structural models, and measurement error methods in longitudinal EHR data. The student will then design and implement Monte Carlo simulation studies in R to evaluate the finite-sample properties of the proposed estimators under realistic EHR data scenarios, including informative observation times and measurement error. They will assist with the development and documentation of an open-source R package implementing the proposed methodology. Finally, the student will contribute to writing a manuscript suitable for submission to a peer-reviewed statistics or biostatistics journal.
Skills required: The ideal candidate has a strong background in mathematical statistics including likelihood theory, regression, and asymptotic methods. Familiarity with survival analysis or longitudinal data methods is important, as the project involves continuous-time modeling of irregular observation processes. Some exposure to Bayesian concepts such as prior and posterior distributions and MCMC sampling is an asset. Prior exposure to causal inference methods such as propensity scores or marginal structural models is beneficial but not required. Strong programming skills in R are preferred, including experience with simulation studies and statistical modeling. The ability to work independently, communicate technical results clearly, and collaborate across
10. Developing Synthetic Data Pipelines for Pediatric Sepsis Research
Supervisor: Matthew Wiens
University: University of British Columbia (Vancouver campus)
Location: Vancouver, British Columbia
Start date: 2027-06-01 (flexible)
Disciplines: Statistics, Computer Science, Medical Sciences
High-quality clinical datasets from prospective studies in low- and middle-income countries remain underutilized because of privacy concerns, re-identification risks, and governance barriers. This limits opportunities for researchers, trainees, and data scientists to develop tools, test methods, and generate hypotheses using data that reflect real-world pediatric sepsis care
This project will develop additional synthetic datasets derived from existing high-value pediatric sepsis cohorts, including datasets related to the successful 2024 Pediatric Sepsis Data Challenge. The project will also initiate a reusable, modular pipeline for synthetic data generation that can be applied to current and future datasets shared through the Pediatric Sepsis Data CoLab network.
The pipeline will integrate proven statistical and machine learning approaches to produce realistic tabular clinical data that preserves important statistical properties, clinical logic, variable distributions, and relationships between predictors and outcomes while minimizing disclosure risk. It will include standardized preprocessing, implementation of multiple synthesis methods, and a multi-dimensional evaluation framework assessing fidelity, privacy, and broad utility through proxy tasks relevant to pediatric sepsis research.
The student will work with the research team to prototype the pipeline, compare candidate synthetic data methods, and generate example synthetic datasets. They will evaluate whether synthetic datasets preserve core characteristics of the original data, support meaningful exploratory and predictive analyses, and reduce privacy risks. The final outputs will include documented code, example synthetic datasets, evaluation summaries, and an open repository that can support future use by the Pediatric Sepsis Data CoLab.
Project objectives:
1. Develop a modular preprocessing and synthetic data generation pipeline for pediatric sepsis datasets.
2. Generate and evaluate example synthetic versions of high-value clinical datasets.
3. Create documentation and reusable code to support safe, scalable, and equitable data access for global pediatric sepsis research.
Research area, student roles & skills
Research area: Along with partners across sub-Saharan Africa, our team at the University of British Columbia (UBC) develops, evaluates and implements data-driven approaches to improve quality of care for children with severe infectious illnesses, including sepsis. A major focus is expanding safe and equitable access to high-quality clinical research data while protecting participant privacy and respecting data governance requirements. Through initiatives such as the Pediatric Sepsis Data CoLab and prior pediatric sepsis data challenges, we support researchers worldwide in using clinical datasets for model development, education, tool testing, and hypothesis generation to improve outcomes for children with sepsis.
Student roles: The student will support the development of a reusable synthetic data generation pipeline for pediatric sepsis research. Working with a multidisciplinary research team in Canada and international partner sites, they will gain experience in:
1. Privacy in global health data science: The student will learn how privacy, data governance, and re-identification risks affect the sharing and reuse of clinical datasets from low- and middle-income countries. They will work with the team to understand dataset structure, variable definitions, clinical logic, and governance considerations relevant to pediatric sepsis cohorts. They will participate in regular team meetings, provide progress updates, and contribute to discussions about safe and equitable data access through the Pediatric Sepsis Data CoLab.
2. Programming, statistical analysis and synthetic data evaluation: The student will review relevant literature and existing methods for generating synthetic tabular health data. They will assist in preparing datasets for synthesis, implementing multiple statistical or machine learning-based synthesis approaches, and documenting methodological choices. They will help evaluate synthetic datasets using measures of fidelity, privacy and identification risk, and utility, including comparisons of variable distributions, relationships between variables, missingness patterns, and performance on proxy analytic tasks relevant to pediatric sepsis research.
3. Documentation and open research: The student will help organize the project code into a modular, reusable pipeline with clear documentation so that the approach can be adapted for future datasets. They will contribute to an open repository, prepare example outputs, and summarize results for the research team. Depending on progress, they may also assist with a short technical report, conference abstract, manuscript, or training materials.
The student will be supported by the principal investigator, biostatistician, and report to the project manager. Through this project, they will gain practical experience in data-driven global health research, reproducible analytics, synthetic clinical data, and responsible data sharing in resource-limited settings.
Skills required: We are looking for candidates with an understanding of fundamental statistical concepts and methods and experience with standard statistical programming software, including Python, SAS, R or STATA, preferably R and/or Python. This student will be mentored and supervised by the biostatistician and principal investigator. Basic familiarity with concepts such as data cleaning, tabular datasets, predictive modelling, or machine learning is an asset. The student should have effective oral and written communication skills and be able to work as part of a team. The ability to maintain accuracy and attention to detail within a privacy sensitive environment is required.
11. Enhancing Insurance Rate Fairness and Profitability Using Lorenz Curves and Optimal-Transport Projected Loss Functions
The insurance industry has long relied on statistical modeling and machine learning for risk assessment and pricing. Traditionally, Lorenz curves have been used to analyze the fairness and profitability of insurance rate structures, often in conjunction with machine learning models for claims prediction. However, due to the highly skewed distribution of insurance claims—where most policyholders do not file claims while a small fraction incur significant losses—traditional methods face challenges such as complex loss distributions, unreasonable rate structures, and the absence of learnable loss functions. In existing approaches, pricing models are primarily based on predefined rules, rather than being directly optimized through learning objectives. This project aims to introduce a Lorenz curve and leverage learnable loss functions to optimize insurance pricing, thereby improving the pricing methods used in the insurance industry.
This research is highly relevant to the student’s academic background in statistics and economics and is closely related to actuarial science, statistics, and machine learning. The proposed optimization method for loss functions can significantly enhance the performance of machine learning models on insurance data and improve their ability to model right-skewed data distributions. Furthermore, this methodology can be widely applied to homeowners, auto, and health insurance rate optimization, helping insurance companies refine their rate structures, increase profitability, and ensure pricing fairness, while also considering market competition and economic factors influencing insurance pricing.
Research area, student roles & skills
Research area: Statistical machine learning; high-dimensional statistics; statistical computing; biomedical, biochemical and industrial data science; applications in drug discovery
Student roles: The student's role covers data analysis, model implementation and optimization, and Gini index estimation. The student will work with real-world insurance data, such as homeowners insurance datasets, to build models and analyze the performance of different pricing methods. Additionally, the student will implement a Lorenz curve as a learnable loss function and validate its performance in deep learning models. Furthermore, the student will integrate gradient boosting and deep learning techniques to model the right-skewed distribution of insurance claims and optimize rate-setting. Finally, the student will explore methods for estimating the Gini index when the exact CDF is unknown and evaluate its effectiveness in measuring the fairness and profitability of insurance pricing.
This project aims to help the student gain an in-depth understanding of actuarial science and core principles of insurance pricing, particularly Lorenz curves, Gini indices, and their applications in the insurance industry. The student will improve data modeling and analytical skills, with a focus on highly skewed insurance data and optimizing pricing models. Additionally, the student will learn machine learning techniques, including optimal transport algorithms and Gini index estimation, and explore their applications in insurance pricing. Finally, the student will study loss functions in pricing optimization and understand how to fine-tune pre-trained models to further improve pricing accuracy.
Skills required: The student will require strong analytical, quantitative and programming skills. I expect background (i.e., at least one undergraduate course) in probability, statistics, machine learning, optimization, computing, algorithms. Knowledge of actuarial science is useful but not required.
12. Estimating functional brain networks from fMRI data
Supervisor: Xiaoping Shi
University: University of British Columbia (Kelowna campus)
Estimating functional brain networks from fMRI data is often complicated by severe noise. While Gaussian Graphical Models (GGMs) are widely used to infer conditional dependence structures, conventional estimators are highly sensitive to data contamination, frequently producing false positive edges. To address this, we consider a robust GGM estimator grounded in Distributionally Robust Optimization (DRO). By accounting for worst-case distributions, the proposed method aims to preserve the true graph structure even with corrupted observations. Our primary focus is to apply this robust framework to high-dimensional fMRI datasets to estimate neural networks.
Research area, student roles & skills
Research area: My current areas of expertise include data sharpening, data depth, Gaussian graph models, Energy-based models, change point analysis, saddlepoint approximation, inverse moment approximation, two-sample comparison, and clustering analysis. I view my research as having three major branches: Applied Statistics, Computational Statistics, and Theoretical Statistics. In Applied Statistics, I research methods to model real data. In Computational Statistics, I develop new methods and implement them in software. In Theoretical Statistics, I focus on studying and developing the mathematical foundations of statistics. These three branches can be mixed.
Student roles: Understand references, test the performance of current methods using simulated data and new data collected, and propose new methods when old ones fail. Attend weekly group meetings and report on progress. In the first month, test the performance of current methods. In the second month, propose new methods where possible and conduct simulation tests. In the third month, provide detailed data analysis. Present at group meetings and seminars/conferences when the opportunity arises and write drafts for submission to journals.
Skills required: Prerequisites include a solid grasp of relevant literature, statistical inference, and data analysis/simulation coding, alongside strong presentation skills.
13. Estimation and inference of reproduction numbers from time series data
Reproduction numbers quantify the rate at which an infectious disease spreads through a population. Estimating these quantities from surveillance data routinely collected by public health authorities is challenging, as such data are often affected by numerous sources of bias. Furthermore, reproduction numbers can be difficult to define because both disease dynamics and surveillance systems evolve over time.
In this project, you will investigate a novel definition of the effective reproduction number and develop methods for estimating it from multivariate time series data. A key challenge is that surveillance data provide an imperfect representation of disease transmission. Changes in testing practices, reporting delays, under-reporting, and other biases can substantially affect estimates of disease spread. You will examine the performance of proposed methods through mathematical analysis, simulation studies, and applications to real-world infectious disease data.
The project will draw on ideas from time series analysis, computational statistics, epidemiological modelling, and statistical inference. Students will gain experience in both methodological research and applied data analysis, with the goal of developing tools that improve the interpretation of infectious disease surveillance data and support evidence-based public health decision-making.
Research area, student roles & skills
Research area: I develop statistical methods to solve problems in epidemiology. This primarily involves spatial statistics, time series, and Bayesian hierarchical modelling methods, and multiple data sources. My recent focus has been on infectious disease epidemiology, veterinary epidemiology, and indirect treatment comparisons. If any of these terms or concepts are new to you, please review them before applying! You can read more about my research on my website: justinslater.ca .
Student roles: The student will contribute to the development of statistical methodology for estimating effective reproduction numbers from multivariate infectious disease time series. This will include deriving and implementing new estimation procedures, conducting simulation studies to evaluate performance under realistic surveillance biases, and applying the methods to real-world public health data. The student will also implement their methods on real data, and conduct simulations in R or related software. Regular participation in research meetings is expected, along with contributions to written reports and dissemination materials. The role is research-intensive and designed to provide training in modern statistical methods for epidemiological applications.
Skills required: The student should be enrolled in a statistics, mathematics, or other highly quantitative program. Applicants should have a strong foundation in computational statistics (particularly Markov chain Monte Carlo methods), time series analysis, and linear algebra. Experience with statistical programming in R or a related language is required, demonstrated through coursework, research experience, or previous employment. Familiarity with Bayesian inference, time series analysis, or infectious disease modelling would be considered an asset.
14. Integrating Bioinformatics into Data Science: A Comprehensive Course for Modern Biological Data Analysis
Supervisor: Yue Zhang
University: Thompson Rivers University (Kamloops campus)
The project aims to develop a specialized course that merges bioinformatics with data science, addressing the increasing need for expertise in analyzing complex biological datasets. This course will provide students with a thorough understanding of both fields, beginning with an introduction to bioinformatics and its applications in genomics, proteomics, and personalized medicine.
The course will cover essential data science concepts such as data cleaning, visualization, and statistical analysis, with a focus on programming languages commonly used in bioinformatics, including Python and R. Students will explore key biological databases like GenBank and UniProt and learn to utilize bioinformatics tools such as BLAST and Clustal Omega.
A significant portion of the course is dedicated to genomic data analysis, teaching methods for sequence alignment, variant calling, and genome assembly, complemented by case studies demonstrating real-world applications. Proteomics and protein structure prediction are also key components, where students will learn to analyze protein expression and interactions and understand the importance of predicting protein structures.
Machine learning will be integrated into the course, with students applying algorithms to biological data for classification, clustering, and prediction tasks, including the use of advanced models like neural networks. The course emphasizes hands-on, project-based learning, with students engaging in group projects and peer reviews to foster collaboration and critical thinking.
By the end of the course, students will be proficient in applying data science techniques to bioinformatics problems, preparing them to contribute to advancements in biomedical research and personalized medicine. This course targets undergraduate and graduate students in biology, computer science, and data science, as well as professionals and researchers seeking to enhance their skills in these areas.
Research area, student roles & skills
Research area: I specialize in statistical mathematics, particularly in the application of advanced statistical methods and machine learning algorithms to solve complex problems in bioinformatics and computational biology. My research focuses on developing novel approaches to analyze large-scale biological data, such as genomic data, to gain insights into biological processes, disease mechanisms, and personalized medicine. I also work on developing computational tools and models to aid in the interpretation and prediction of biological phenomena, with the ultimate goal of advancing our understanding of life sciences.
Student roles: The student will assist in developing and implementing the "Integrating Bioinformatics into Data Science" course. Responsibilities include:
Curriculum Development: Collaborate on designing course modules, assignments, and projects. Content Creation: Develop instructional materials, including lectures, tutorials, and documentation. Data Analysis: Perform data cleaning , visualization, and analysis on biological datasets to create real-world examples and case studies.
Tool Implementation: Assist in setting up and demonstrating bioinformatics tools and software, such as BLAST and Clustal Omega. Machine Learning Integration: Apply machine learning algorithms to biological data, creating practical examples for students. Student Support: Provide guidance and support to students, answering questions and assisting with technical issues. Evaluation: Help in designing quizzes, assignments, and projects, and assist in grading and providing feedback.
Skills required: Ideal candidates are undergraduate or graduate students in bioinformatics, data science, computer science, or biology. They should possess strong programming skills in Python and R, experience with data analysis and visualization tools, and familiarity with bioinformatics databases and tools like BLAST and GenBank. Understanding of machine learning algorithms and their applications to biological data is essential. Strong analytical, problem-solving, and communication skills, along with the ability to work collaboratively in a team, are crucial. Passion for bioinformatics and adaptability to learn new techniques are highly valued.
15. Integrating Factor of Safety and Probability of Failure: A Streamlined Approach for Dam Safety Assessment
This research project focuses on improving how the safety of dams and other water-retaining structures is evaluated. Structures such as dams, spillways, and intake systems are essential for water management and hydroelectric power, but their safety is often assessed using simplified methods that do not fully account for uncertainty or the real consequences of poor performance. This project aims to develop better tools to help engineers make more informed decisions about safety, maintenance, and rehabilitation.
The project will explore how traditional safety measures, such as the factor of safety, can be better connected to probability of failure and risk. It will also use numerical modeling, data analysis, and machine learning to create faster and more efficient ways of evaluating structural performance under different loading conditions. A longer-term goal is to support a more performance-based approach, where engineers can better understand not only whether a structure is safe, but also how it is likely to behave under different scenarios and what the consequences could be.
Conducted in collaboration with Hydro-Québec, this project has strong practical relevance and offers an opportunity to contribute to research that supports safer and more resilient infrastructure.
Research area, student roles & skills
Research area: I specialize in performance-based safety assessment, vulnerability evaluation, and multi-hazard analysis of critical infrastructure. My research advances resilience to extreme events and climate change by integrating infrastructure equity, community resilience, and climate justice. I also use data science to bridge built, natural, and social systems, developing holistic approaches to infrastructure risk, adaptation, and environmental challenges.
Student roles: For an undergraduate student, the role will be to support research activities related to the safety assessment of dams and other water-retaining structures. The student will assist other graduate students working on the project with tasks such as organizing data, reviewing literature, preparing input files, helping with numerical and statistical analyses, and supporting the interpretation and visualization of results.
The role will be adapted to the level of an undergraduate intern. Depending on the student’s background, they may also contribute to MATLAB and/or Python scripts for data processing, reliability-related calculations, and simplified machine-learning-based or surrogate modeling tasks. The student may help with parts of the project related to linking factor of safety to probability of failure, supporting the development of simplified computational tools and figures, and organizing workflows for fragility and performance-based assessment.
The student may also assist with preparing figures, summaries, presentations, and technical documentation. This position offers hands-on experience in structural safety, uncertainty and risk analysis, and applied computational methods in civil engineering, while allowing the student to learn directly from other graduate students and the broader research team.
Skills required: I am looking for students interested in dams, structural analysis, reliability analysis, machine learning, and data analysis. Experience with Python and/or MATLAB is important for working with data, developing models, and creating visualizations. A good foundation in probability and statistics is expected, and familiarity with machine learning, classification, regression, and uncertainty analysis would be an asset. Students should be curious, motivated, and comfortable working with data to support practical engineering decision-making.
16. Interpretability of black-box models, theoretical and computational aspects
Black-box machine learning models achieve strong predictive performance but remain opaque: their internal logic is neither transparent nor auditable, leaving potential biases and unintended discrimination undetected. This internship aims to study post-hoc explainability methods, including feature attribution, surrogate models, and game-theoretic approaches, to rigorously characterize the theoretical foundations and computational trade-offs of existing techniques.
Research area, student roles & skills
Research area: I am interested in all aspects of machine learning interpretability and especially of non-linear models with dependent features.
Student roles: The student will have the opportunity to: (1) work on a real dataset to analyze a concrete problem; (2) implement well-documented algorithms to operationalize a methodology; (3) conduct a scientific literature review on a targeted research area.
Skills required: Strong proficiency in statistics and mathematics. Ability to code in R. A plus if familiar with parallelization techniques.
17. Is the interaction between depression and function moderated by depression treatment across cognitive statuses
Supervisor: Shanna Trenaman
University: Dalhousie University (Halifax campus)
Location: Halifax, Nova Scotia
Start date: 2027-05-10 (flexible)
Disciplines: Statistics, Pharmacy, Pharmacology, Medicine, Medical Sciences
The research project will use data from the COMPASS-ND longitudinal cohort. This cross-sectional study will assess whether antidepressant exposure is associated with differences in executive function among adults living with depression across varying levels of cognitive status, including normal cognition, subjective cognitive impairment, mild cognitive impairment, and dementia (across subtypes). Antidepressant use will be characterized based on exposure status and relevant medication features (e.g., class, anticholinergic burden), and executive function will be evaluated using standardized cognitive measures available within the dataset. The analysis will compare executive function outcomes between exposed and unexposed individuals within and across cognitive diagnostic groups, while accounting for relevant demographic and clinical covariates.
Research area, student roles & skills
Research area: I am a clinician-scientist who studies appropriate drug use in older adults from the perspectives of pharmacology, epidemiology, pharmacoepidemiology, and health services research. My research focuses on polypharmacy, deprescribing, medications used by older adults with dementia, anticholinergic medications and antidepressants, and sex- and gender-based differences in drug use. My work includes surveillance of medication use, particularly antipsychotics, with a focus on safer prescribing, deprescribing, and improving outcomes for older adults, especially those living with dementia.
Student roles: The student(s) will contribute to the design of a study assessing how antidepressant exposure affects functional outcomes in adults living with depression across a range of cognitive diagnoses. They will conduct literature reviews to develop a strong understanding of the area and assist in refining the study design. The student(s) will also learn to code in R statistical software and complete the data analysis. There will be opportunities to contribute to a manuscript and/or develop a poster to present the findings.
Skills required: The student should have basic statistical skills and foundational knowledge of geriatric medicine, medications, depression, and dementia. They should demonstrate an interest in learning to code in R statistical software. The student should be comfortable working both independently and as part of a team.
18. Lead-lag detection in time series models, with applications to financial data
The main goal of the project will be to study methods for lead-lag detection in financial time series data. We will take a Bayesian approach to the problem, which will allow us to quantify uncertainty in the resultant lead-lag estimates. The project will involve understanding and testing different factor models for lead-lag detection as well as experimenting with a variety of Monte Carlo methods to perform Bayesian inference for the lag values. We will apply the methods to financial time series data and test their performance in real-world scenarios, with a focus on the creation of lead-lag portfolios.
Research area, student roles & skills
Research area: I conduct research in Bayesian statistics, with a particular focus on the modelling of time-varying data. This includes the development of models that capture unobserved structure in time series data, such as the presence of lead-lag relationships where the dynamics of one time series lead another one. I am also interested in understanding how to detect structural changes in time series data where the underlying process transitions from one mode to another.
Student roles: The student will implement and test a variety of Bayesian lead-lag detection methods. The student will also have an opportunity to work on novel methods, depending on the progress of the project.
Skills required: The student should be familiar with statistics and probability, basic time series, Monte Carlo methods and programming in a statistical computing language such as R or Python.
19. Leveraging Statistical Learning to Improve Water Efficiency and Sustainability in Vineyard Production
Supervisor: Yue Zhang
University: Thompson Rivers University (Kamloops campus)
This research project will use existing datasets from two field-based horticultural studies to develop appropriate statistical analysis pathways, generate clear data visualizations, and interpret environmental and crop performance responses to different soil and water management practices.
Project 1 focuses on improving the sustainability of irrigated vineyards through reduced irrigation, water-retention technology, and organic soil amendments. These treatments were arranged in a split-plot factorial experimental design. Data were collected over multiple years, including vine performance, soil and plant responses, and continuously monitored variables such as soil moisture and soil temperature. The student will review the experimental design and dataset structure, identify suitable statistical models, and analyze treatment effects and interactions across time. The project may include mixed-effects modelling, repeated-measures analysis, time-series visualization, and exploratory machine learning approaches to identify environmental patterns and predictors of vineyard performance. The results will help clarify how irrigation reduction, soil amendments, and water-retention strategies influence soil conditions and vineyard sustainability.
Research area, student roles & skills
Research area: Dr.Mehdi Sharifi is a soil nutrient management research scientist at Agriculture and Agri-Food Canada. His research program is focused on soil nutrient management and cover crops in perennial horticultural systems.
Dr. Yue Zhang’s research focuses on probabilistic modelling, statistical learning, and AI-driven data analysis for biological and environmental systems.
Student roles: The student will assist with data collection, preprocessing, statistical analysis, and computational modelling related to biological and environmental datasets. Responsibilities may include implementing machine learning or statistical methods, conducting literature reviews, preparing figures and summaries, and supporting reproducible research workflows. The student may also help evaluate model performance, organize datasets, and participate in research discussions and team meetings. Opportunities to contribute to conference presentations, reports, or publications may be available depending on project progress and student interests.
Skills required: The student should have a background in mathematics, statistics, computer science, bioinformatics, or a related quantitative field. Experience with programming (e.g., Python or R), data analysis, and basic machine learning or statistical methods is preferred. Familiarity with biological or agricultural or environmental data is an asset. Strong problem-solving skills, attention to detail, and the ability to work independently and collaboratively are important.
20. Longitudinal Statistical and Machine Learning Methods in Basketball Analytics
Supervisor: Sean Hellingman
University: Thompson Rivers University (Kamloops campus)
The student will work on an applied statistical modelling project focused on the National Basketball Association (NBA) and the National Collegiate Athletic Association Division I (NCAA DI). A database containing twenty years of NCAA and NBA player data, including season-by-season outcomes, has been constructed. The student will apply statistical models and machine learning methods to draw meaningful conclusions about players, teams, and player development trajectories. Specific research objectives will be refined based on identified research gaps prior to May 2027.
The longitudinal nature of the data presents several interesting methodological challenges. In particular, relatively few NCAA DI players transition to the NBA, resulting in substantial sample imbalance. Additionally, the structure of the dataset allows for the investigation of changes in player and league outcomes over time, including potential effects of rule or league changes across the twenty-year period.
The ultimate goal of the project is for the student to contribute to a peer-reviewed publication in an applied statistics or sports analytics journal as a named co-author.
Research area, student roles & skills
Research area: My research focuses on statistical modelling with applications in sports analytics. My interests include multilevel models, statistical learning, stochastic modelling for sports tracking data, survival analysis, applied computer vision, machine learning, and machine learning applications in ecological monitoring.
Student roles: The student will assist with data cleaning, exploratory analysis, statistical modelling, diagnostic checks, and the interpretation and communication of results. Responsibilities may include developing reproducible analysis pipelines, implementing machine learning models, conducting literature reviews, and preparing figures, tables, and written summaries of findings.
The student will participate in regular research meetings, contribute to methodological discussions, and help identify meaningful research questions from the dataset. Depending on project progress and interests, the student may also contribute to manuscript preparation.
Skills required: The ideal student will have a working understanding of statistical modelling and its underlying assumptions, experience with applied machine learning, and proficiency in data cleaning and restructuring. Strong communication and technical writing skills are important. Preference will be given to students with an interest in sports analytics and familiarity with basketball.
Proficiency in R or Python is required, as all analyses will be conducted using reproducible workflows.
21. Novel statistical models and machine learning algorithms for next-generation drug discovery
The primary goal of this research proposal is to develop mathematically rigorous methods for quantifying uncertainty in machine learning models, emphasizing innovations like conformal prediction and nonparametric statistical techniques. Our aim is to enhance the reliability and interpretability of machine learning-driven decision making, making it more trustworthy for critical applications, particularly in drug discovery and precision medicine.
Decision-making in fields like drug development often requires resource-intensive evaluation processes and multiple rounds of expensive trials. Early-stage identification of viable drug compounds is essential for efficient resource allocation. Machine learning can significantly streamline the decision-making process by providing an initial screening that narrows down the candidate pool for more detailed evaluation in later stages.
In this proposal, we seek to further refine the conformal selection framework for machine learning-driven decision-making. This could be achieved through several perspectives: (i) We propose a multivariate conformal selection method that enables simultaneous candidate filtering across multiple targets, thereby preserving risk control and enhancing power. (ii) We propose that a more flexible approach would be to design a multi-category selection procedure, allowing candidates to be filtered into multiple zones that correspond to varying levels of confidence in their selection. This would provide greater adaptability in the decision-making process and better align with the practical considerations of the industry. (iii) In practice, incoming candidates for selection may differ from those used to train the machine learning prediction models, potentially violating key assumptions of our method. However, when this difference can be attributed to covariate shift, the weighted conformal selection method provides an effective solution. We will evaluate its performance in real-world drug discovery, where it could serve as a valuable extension for broader application scenarios.
Research area, student roles & skills
Research area: Statistical machine learning; high-dimensional statistics; statistical computing; biomedical, biochemical and industrial data science; applications in drug discovery
Student roles: Our proposed program offers research trainees (3 UG students) a comprehensive and enriched training experience, providing both technical and professional skill development opportunities. Through hands-on participation in computational projects, one UG student will strengthen his/her technical expertise, focusing on practical challenges related to drug discovery. Two UG student, on the other hand, will contribute directly to the development of new methodologies and theories, gaining advanced knowledge in statistical learning, high-dimensional data analysis, and machine learning, with a particular emphasis on conformal prediction in drug discovery and screening.
Skills required: The student will deepen their expertise in statistical machine learning, focusing on conformal prediction, selection, multiple testing, and uncertainty quantification. This work will contribute to research output and build a strong theoretical foundation for future academic pursuits. Additionally, the student will apply the method to real-world drug discovery applications.
The student will require strong analytical, quantitative and programming skills. I expect background (i.e., at least one undergraduate course) in probability, statistics, machine learning, optimization, computing, algorithms. Knowledge of drug discovery is useful but not required.
22. Optimal subsampling for big data regression
Supervisor: Po Yang
University: University of Manitoba (Winnipeg campus)
Modern technological advancements have resulted in a massive surge of data across various fields. Computational cost when analyzing the data is a crucial problem. Subsampling methods which select a subset of data from the full dataset have attracted significant attention in recent years. The goal of subsampling is to identify a subset that either comprises the most informative observations within the big data or captures the overall structure of the full data set. Optimality criteria, such as D- and A-criteria, for selecting optimal designs have been applied to choose the most informative subsamples for various models.
Most existing research in this area assumes that all variables are continuous. However, in practice, data often involve categorical variables. Although studies on this type of data have appeared in the literature, the existing research remains very limited. Most existing subsampling methods focus on continuous variables and overlook the prevalence of categorical data in many fields.
In Summer 2027, I plan to consider cases where some variables are categorical. I plan to develop a stratified subsampling method that enables accurate prediction in the presence of both continuous and categorical variables under a linear model. In my approach, observations corresponding to the same combination of categorical variable levels are first grouped into strata. Then, observations are selected proportionally from each stratum using simple random sampling. The combined observations form a subsample, which is evaluated using the I-criterion. We will develop an efficient algorithm that accommodates the stratified subsampling procedure. The selected subsample is expected to comprise the most informative observations with the reduced prediction variance. The results will provide guidance for data analysts to select informative subsamples.
Research area, student roles & skills
Research area: In the era of big data, computational cost when analyzing the data is a crucial problem. Strategically selecting informative subsets of data which can retain statistical efficiency while significantly reducing computational and financial costs has become increasingly critical. My primary research interests focus on the development of optimal subsampling methods for big data. I investigate new theory, methodology, and computational algorithms that can be used to select informative subsamples from big data.
Student roles: In my research work, a lot of problems involve the developments of computer programs that are used to search for informative subsamples. I also need to compare my results with others. Learning and developing computer programs will provide high quality training for the students. During the first three weeks, students will learn some basic programming skill and necessary theoretical background related to my project. After that, students will start to work on the computer programs and analyze results. They will work with my current research team including PhD and MSc students, as well as possible one or two undergraduate summer research students.
Skills required: Students have taken some basic mathematics and statistics courses. Experience of using statistical software, such as R, is preferred but not required.
23. Quantum-Enhanced Fusion of Heterogeneous Geospatial Data for Advanced Forest Inventories
Supervisor: Ahmed Ragab
University: École Polytechnique de Montréal
Location: Montreal, Québec
Start date: 2027-05-03 (flexible)
Disciplines: Statistics, Science and Technology, Industrial Design and Technology, Geomatics, Forestry, Environmental Studies, Engineering, Engg-Systems and Technology, Engg-Software
Accurate forest inventories require the integration of large and diverse geospatial datasets. This project will investigate quantum-enhanced and quantum-inspired machine learning approaches for fusing LiDAR, satellite, and field data to improve forest attribute estimation. The student will explore the use of quantum optimization and hybrid quantum-classical learning methods using platforms such as Qiskit to accelerate model training and feature selection. The project aims to assess the potential benefits of emerging quantum technologies for improving prediction accuracy and computational efficiency in forest inventory applications.
Research area, student roles & skills
Research area: My research focuses on artificial intelligence, geospatial analytics, and emerging quantum computing technologies for environmental monitoring and natural resource management. We develop advanced methods that integrate satellite imagery, airborne LiDAR, drone observations, and field measurements to improve forest inventories and support sustainable forest management. Our work combines machine learning, data fusion, uncertainty quantification, and quantum-inspired optimization techniques to address computational challenges associated with large-scale geospatial datasets and predictive modeling.
Student roles: The student will contribute to the development and evaluation of machine learning and quantum-inspired approaches for integrating heterogeneous geospatial datasets. Activities include data preprocessing, implementation of AI models, experimentation with Qiskit-based algorithms, performance benchmarking, and analysis of computational efficiency. The student will collaborate with researchers from AI, forestry, and data science disciplines and contribute to technical reports, presentations, and potential scientific publications.
Skills required: Students should have a background in computer science, engineering, forestry, geomatics, environmental sciences, physics, or a related field. Experience with Python programming, machine learning, statistics, or data analytics is desirable. Familiarity with remote sensing, GIS, deep learning, or quantum computing concepts is an asset but not required. Students with an interest in artificial intelligence, quantum technologies, and environmental applications are encouraged to apply.
24. Re-analysis of COVID-19 incidence models with up-to-date data
Supervisor: Devan Becker
University: Wilfrid Laurier University (Waterloo campus)
In the early stages of the SARS-CoV-2 pandemic there was a flurry of academic activity to characterize the spread of the disease and inform governments and the public. However, it is difficult to know how soon a question can be reliably answered from limited data. How much data do you need before you can make a conclusion about the relationship between the climate and the spread of COVID-19? This re-analysis will leverage snapshots from the COVID-19 data hub (https://covid19datahub.io/) to replicate the analysis of published papers as if they were performed at various stages of the pandemic.
Focus will be on the papers that have graciously published their code and any extra data used, as identified on the COVID-19 data hub (I am not against GenAI assistance to help create code where it is not available). For each identified paper, we will attempt to replicate the results using data that were different stages of the pandemic that demonstrates the effect of new data as the pandemic carries on. The results will be evaluated over time, especially to see how early we could have reached similar conclusions and whether the waves of the pandemic would have changed the conclusions in the paper.
This analysis will provide insights into the epistemic variability of an ongoing global health crisis, and will allow us to know when models are appropriate in potential future pathogen outbreaks.
Research area, student roles & skills
Research area: My research pertains to modelling of time series data for infectious diseases, primarily from wastewater. I emphasize interpretable and explainable models with proper quantification of uncertainty so that the results and limitations of statistical modelling can be easily understood by a broader audience.
Student roles: 1. Identify a paper from a given collection with interesting results and available data. 2. Reproduce the paper's results as closely as is feasible. 3. Re-run the analysis at various points in the pandemic, quantifying the results over time with models and/or visualizations. 4. Interpret and communicate the results.
Skills required: - Programming skills in R (preferred) or Python (GenAI coding assistance is allowed if the student can independently verify the code quality). - Some prior training in statistics, probability, and regression. Time series analysis is a bonus. - Data visualization skills. - Ability to read and understand academic publications. - Ability to communicate technical results to a broader audience.
25. Recovering Structure under Measurement Error
Supervisor: Xiaoping Shi
University: University of British Columbia (Kelowna campus)
Measurement error can distort observed data and obscure important underlying structures, leading to biased inference and reduced model performance. This project develops a data sharpening approach to adjust noisy observations so that they better reflect latent patterns. The method works by modifying the distribution of covariates and combining it with smoothing techniques to reduce the flattening effect caused by noise. The goal is to improve the accuracy of model estimation and decision-making by recovering latent patterns that are otherwise hidden by noise.
Research area, student roles & skills
Research area: My current areas of expertise include data sharpening, data depth, Gaussian graph models, Energy-based models, change point analysis, saddlepoint approximation, inverse moment approximation, two-sample comparison, and clustering analysis. I view my research as having three major branches: Applied Statistics, Computational Statistics, and Theoretical Statistics. In Applied Statistics, I research methods to model real data. In Computational Statistics, I develop new methods and implement them in software. In Theoretical Statistics, I focus on studying and developing the mathematical foundations of statistics. These three branches can be mixed.
Student roles: Understand references, test the performance of current methods using simulated data and new data collected, and propose new methods when old ones fail. Attend weekly group meetings and report on progress. In the first month, test the performance of current methods. In the second month, propose new methods where possible and conduct simulation tests. In the third month, provide detailed data analysis. Present at group meetings and seminars/conferences when the opportunity arises and write drafts for submission to journals.
Skills required: Prerequisites include a solid grasp of relevant literature, statistical inference, and data analysis/simulation coding, alongside strong presentation skills.
26. Sparse Multiblock Methods for the Analysis of Mixed Data
Modern datasets often consist of multiple blocks of data collected on the same individuals, such as demographic characteristics, clinical measurements, biomarkers, genetic information, and questionnaire responses. These datasets frequently contain both quantitative and qualitative variables and may involve a large number of variables, making interpretation challenging.
The objective of this project is to develop and evaluate sparse multiblock statistical methods for the analysis of mixed data. Sparse methods perform variable selection while preserving the multiblock structure of the data, helping identify the most relevant variables and improving interpretability. The project will focus on extending existing multivariate approaches to jointly analyze multiple blocks containing both quantitative and qualitative variables.
The student will contribute to simulation studies designed to assess the performance of existing and newly developed methods under a variety of scenarios. The project will involve statistical programming in R, method evaluation, and the analysis of real datasets from health, epidemiology, genetics, or sensory science. The student will participate in the development and assessment of new statistical approaches and help interpret and communicate the results.
Research area, student roles & skills
Research area: My research focuses on the development of multivariate statistical methods for the analysis of complex datasets containing multiple blocks of quantitative and qualitative variables. Particular interests include dimension reduction, variable selection, multiblock analysis, sparse methods, and the integration of heterogeneous data sources. These methods have applications in health research, epidemiology, genetics, and other fields involving high-dimensional and complex data.
Student roles: The student will participate in the development and evaluation of statistical methods for the analysis of mixed data. Responsibilities will include conducting simulation studies, analyzing simulated and real datasets, implementing methods in R, preparing reports, and summarizing results. The student will also participate in regular research meetings and discussions. Depending on progress and interest, opportunities may exist to contribute to scientific publications and conference presentations.
Skills required: Candidates are expected to have a background in statistics, biostatistics, or data science. Experience with R programming is required. Knowledge of multivariate analysis is an asset but is not required. Interest in statistical methodology, simulation studies, and data analysis is desirable. Good oral and written communication skills are also expected.
27. Spatial non-negative matrix factorization for deconvoluting transcriptomics data
The traditional non-negative matrix factorization model overlooks the spatial correlation structure among different locations, which is often present in various applications. Latent features in neighboring locations tend to exhibit greater similarity compared to those in distant locations, leading to spatial correlation among rows in the factorization model. This project aims to devise a spatially constrained non-negative matrix factorization approach to leverage spatial information across locations for improved estimation of latent features. Additionally, we intend to apply this method to facilitate cell-type deconvolution in bulk RNA-seq studies.
Research area, student roles & skills
Research area: Statistical machine learning; high-dimensional statistics; statistical computing; biomedical, biochemical and industrial data science; applications in drug discovery
Student roles: This student will:
Conduct thorough literature reviews.
Establish the theoretical framework.
Actively engage in implementing the proposed method.
Interpret the results and contribute to writing research reports.
The student's involvement is expected to enhance his research skills on statistical learning, particularly on dimension reduction, matrix factorization and spatial statistics, leading to meaningful contributions to research output and providing valuable research experience for future academic pursuits.
Skills required: The student will require strong analytical, quantitative and programming skills. I expect background (i.e., at least one undergraduate course) in probability, statistics, machine learning, optimization, computing, algorithms. Knowledge of biology is useful but not required.
28. Statistical Methods for Identifying X-Chromosome Inactivation Patterns
The X chromosome plays an important role in many complex diseases, yet it remains underrepresented in genetic association studies due to its unique biological characteristics, including X-chromosome inactivation (XCI). XCI is a process by which one of the two X chromosomes in females is partially or completely silenced, and different XCI patterns can influence genetic association results.
The objective of this project is to develop and evaluate statistical methods for identifying XCI patterns and incorporating them into genetic association analyses. The student will contribute to simulation studies designed to assess the performance of existing and newly developed methods under a variety of genetic and biological scenarios. The project will involve statistical programming, method evaluation, and data analysis using both simulated and real genetic datasets.
The student will participate in the development and assessment of new statistical approaches, compare their performance with existing methods, and help interpret and communicate the results. The findings from this project may contribute to improving statistical tools for X-chromosome association studies and advancing our understanding of the role of XCI in complex diseases.
Research area, student roles & skills
Research area: My research focuses on the development of statistical methods for genetic and genomic studies, with particular expertise in X-chromosome association analysis. I develop and evaluate statistical approaches for identifying and modeling X-chromosome inactivation (XCI) patterns and their impact on complex diseases. My work combines biostatistics, statistical genetics, simulation studies, and data analysis to improve our understanding of genetic mechanisms and to enhance the accuracy of association studies.
Student roles: The student will participate in the development and evaluation of statistical methods related to X-chromosome inactivation. Responsibilities will include conducting simulation studies, generating and analyzing simulated datasets, implementing statistical methods in R, summarizing results, preparing reports, and helping compare the performance of competing approaches. The student may also contribute to the analysis of real genetic datasets and participate in regular research meetings. Depending on progress and interest, opportunities may exist to contribute to scientific publications and conference presentations.
Skills required: Candidates are expected to have a background in statistics, biostatistics, bioinformatics, computer science, or a related field. Experience with R programming is required. Knowledge of genetics is an asset but is not required. Interest in statistical methodology, simulation studies, and data analysis is desirable. Good oral and written communication skills are also expected.
29. Statistical Modeling and Data Visualization of Ground Cover Effects on Sweet Cherry Orchard Performance
Supervisor: Yue Zhang
University: Thompson Rivers University (Kamloops campus)
This research project will use existing datasets from two field-based horticultural studies to develop appropriate statistical analysis pathways, generate clear data visualizations, and interpret environmental and crop performance responses to different soil and water management practices.
Project 2 focuses on the effect of orchard ground cover management on sweet cherry production. Three ground-cover treatments, including an untreated control, bark mulch, and woven black plastic, were applied beneath cherry trees. Soil characteristics, tree growth, productivity, and related orchard performance parameters were measured. The student will determine the most appropriate statistical methods for this experimental dataset, analyze treatment effects, and interpret how different ground covers influence soil conditions, tree development, and yield-related outcomes.
Research area, student roles & skills
Research area: Dr.Mehdi Sharifi is a soil nutrient management research scientist at Agriculture and Agri-Food Canada. His research program is focused on soil nutrient management and cover crops in perennial horticultural systems.
Dr. Yue Zhang’s research focuses on probabilistic modelling, statistical learning, and AI-driven data analysis for biological and environmental systems.
Student roles: The student will assist with data collection, preprocessing, statistical analysis, and computational modelling related to biological and environmental datasets. Responsibilities may include implementing machine learning or statistical methods, conducting literature reviews, preparing figures and summaries, and supporting reproducible research workflows. The student may also help evaluate model performance, organize datasets, and participate in research discussions and team meetings. Opportunities to contribute to conference presentations, reports, or publications may be available depending on project progress and student interests.
Skills required: The student should have a background in mathematics, statistics, computer science, bioinformatics, or a related quantitative field. Experience with programming (e.g., Python or R), data analysis, and basic machine learning or statistical methods is preferred. Familiarity with biological or agricultural or environmental data is an asset. Strong problem-solving skills, attention to detail, and the ability to work independently and collaboratively are important.
30. Statistical and Machine Deep Learning Strategies with Applications
Supervisor: S. Ejaz Ahmed
University: Brock University (St. Catherines campus)
The proposed research project aims to develop and evaluate advanced statistical and machine learning methods for analyzing high-dimensional data, where the number of variables can greatly exceed the sample size. Traditional inference methods often rely on strong sparsity assumptions, which may not hold in many modern applications, such as genomics, neuroscience, and other complex data-rich domains. This project seeks to address these limitations by developing post-shrinkage strategies for regression parameters across a variety of statistical and generative models. These strategies are designed to improve the accuracy, interpretability, and robustness of parameter estimates in settings characterized by weak or approximate sparsity.
A central focus of the project will be the theoretical investigation of these methods, including their statistical properties, consistency, and performance across diverse sparsity regimes. Furthermore, the methods will be applied to real-world datasets, such as genomics and other high-dimensional biomedical data, to demonstrate their practical utility and ability to generate meaningful scientific insights.
Another important aspect of the project is to provide trustworthy inference in situations where datasets are extremely large and cannot be stored or processed on a single machine. To address this challenge, we will develop shrinkage-based methods combined with distributed learning techniques, enabling scalable and efficient analysis while maintaining statistical accuracy. These approaches aim to produce reliable and interpretable results even when working with massive datasets, ensuring that insights drawn from big data remain valid and meaningful.
The ultimate goal of this research is to provide a suite of flexible, interpretable, and widely applicable tools for modern data analysis, enabling researchers to draw reliable inferences from increasingly complex and high-dimensional datasets.
Refernce:
Post-Shrinkage Strategies in Statistical and Machine Learning for High Dimensional Data, By S. Ejaz Ahmed, F. Ahmed, B. Yüzbaşı, Published May 25, 2023 by Chapman & Hall.
Research area, student roles & skills
Research area: My research focuses on developing statistical methodology for high-dimensional data, where the number of variables may greatly exceed the sample size. In particular, I work on penalized regression methods, weak sparsity frameworks, and mixed linear models. A central theme of my work is the development of post-shrinkage estimation strategies that improve inference beyond standard regularization techniques. I study the theoretical properties of these methods across a broad parameter space and evaluate their performance through simulations and applications to complex data, including genomics. My goal is to provide robust and interpretable tools for modern data analysis in scientifically relevant settings.
Student roles: The role will be tailored to each student’s background and experience, ensuring an appropriate balance of challenge and learning. Based on prior experience, the work will be primarily numerical rather than theoretical, focusing on conducting simulations, analyzing data examples, and implementing statistical and machine learning methods. In addition, students will gain experience in scientific communication through report writing, summarizing results clearly and concisely, and presenting findings in a structured, professional manner. This combination of hands-on computational work and effective reporting is designed to strengthen both technical and analytical skills.
Skills required: Intermediate proficiency in probability, statistics, and mathematical modeling, with practical skills in computer modeling and coding for implementing statistical and machine learning methods in real-world data analysis.
31. Understanding Representation Learning via Contrastive Principal Component Analysis
Recent empirical studies have effectively utilized unlabeled data to develop feature representations that prove beneficial for subsequent classification tasks. Building upon this line of research, Abid (2018) introduced a novel approach known as contrastive principal component analysis (cPCA). Unlike traditional PCA, cPCA is tailored to identify low-dimensional structures unique to a dataset or enriched in one dataset compared to others. It serves as a generalization of standard PCA, specifically applicable when multiple datasets are available, such as a treatment group versus a control group or a mixed versus a homogeneous population. The primary objective is to uncover patterns that are specific to one of the datasets. However, the theoretical understanding of this technique remains limited, with previous studies predominantly focusing on computational aspects. In this project, we aim to address this gap by leveraging theoretical frameworks from learning theory to establish generalization bounds. Our objective is to demonstrate that accuracy guarantees hold when optimizing the contrastive PCA objective.
Research area, student roles & skills
Research area: Statistical machine learning; high-dimensional statistics; statistical computing; biomedical, biochemical and industrial data science; applications in drug discovery
Student roles: In this student project, the student will actively participate in implementing contrastive principal component analysis (cPCA) and its application to real-world datasets. The student will conduct literature reviews on cPCA under supervision and will collaborate with the supervisor to develop the theoretical framework, interpret the findings, and contribute to writing research reports.
Skills required: The student's involvement is expected to enhance his research skills on statistical learning, particularly contrastive learning and self-supervised learning, leading to meaningful contributions to research output and providing a strong theoretical foundation for future academic pursuits.
The student will require strong analytical, quantitative and programming skills. I expect background (i.e., at least one undergraduate course) in probability, statistics, machine learning, optimization, computing, algorithms.