About: Informal use and much anecdotal evidence suggests that the most recent LLMs, accessed via agentic AI coding systems, have reached a stage where they are very capable of exploring large datasets under supervision and with human guidance. Both exploratory and confirmatory analysis appears to be possible with results presented for verification by the practitioner. The A in AStats could stand for autonomous, augmented, automatic, applied, etc.
The project will explore and define good practices for robust workflows that incorporate agentic AI into statical exploration and practice. Practitioners already often use recipe-driven methods (e.g. JASP, Jamovi) to guide their use of statistical tools in familiar contexts. A major focus will be on the automatic exploration of large datasets, as well as the possibility of fine-tuning workflows or even models and using open-weight models to reduce cost and customize usage and make workflows more predictable.
Research area, student roles & skills
Research area: Large language models, AI, Machine learning, Statistics, AI for Science, Agentic AI, Data Science
Student roles: The student will take the lead and will be mentored on a specific aspect of the project that is of interest and that fits their skills.
Skills required: Familiarity with the use of agentic AI workflows and the use of LLMs. Familiarity with statistical practice at a moderately advanced level is a plus. Familiarity with setting up and using open-weight LLMs and with fine-tuning LLMs is a plus.
2. Application of Large Language Models for Sustainability Assessment
The assessment of sustainability of undertakings or policies is done based on analysis of reporting of attainment of qualitative and quantitative targets which are presented in textual documents. Growing number of such documents must be processed and comprehended to evaluate the current state of attainment of SDGs to make future decisions. This pressing need can be addressed by automated processing of textual documents and unstructured data utilizing advancements of Natural Language Processing (NLP) techniques. The project is aimed at developing efficient framework for identification and processing publicly available documents to extract information on attainment of Sustainable Development Goals. Implementation of the project requires building corpora of documents, fine-tuning pre-trained Large Language Models (LLMs) for the downstream tasks and evaluation of document relevance to a target text. Document selection is based on application of semantic similarity metrics. The results of the project will be a part of framework for automated sustainability appraisal.
Research area, student roles & skills
Research area: Professor Erechtchoukova's research is in the fields of data-driven and model-driven decision support in environmental sustainability, application of Natural Language Processing (NLP) techniques to document categorization and developing question answering systems in low resource domains, application of machine learning techniques to water resource management, simulation modeling of complex systems, data modeling for semistructured data, and optimization of scheduling systems. Current research projects are devoted to developing hybrid predictive modeling frameworks for integrated hydrology, application of artificial intelligence for education and sustainability assessment, optimization of healthcare scheduling systems, sustainability appraisal, and optimization of environmental monitoring
Student roles: An intern will become a team member working on extraction of documents relevant to a selected problem domain and available online. They will implement text pre-processing and embedding using recommended software. The Intern will analyze and select language models, fine-tune them, and evaluate model performance following the provided methodology. They will apply pre-trained and fine-tuned LLMs to generate summaries of collected documents, and implement similarity analysis, participate in planning and preparation of computational experiments, and will be responsible for their execution, and analysis of the results.
Skills required: Interest in Big data analytics, Artificial Intelligence and Natural Language Processing is a must; desire to learn and explore new computational tools is a must; ability to understand software manuals and use the software to run computational experiments; ability to work in cloud-based environment (e.g. Google Colab); willingness to run large number of computational experiments; ability to write computer code in Python or R, understanding basic statistics; ability to implement repetitive tasks. Familiarity with text preprocessing and various embedding techniques is an asset.
3. Breeding Better Plant-based Proteins: Unlocking the Nutritional Potential of Pulses
Plant-based proteins are increasingly recognized as essential components of sustainable food systems. However, not all proteins are nutritionally equivalent. Protein quality depends on factors such as amino acid composition, and the presence of antinutritional compounds, all of which influence the body's ability to utilize dietary protein.
Despite significant advances in pulse breeding, most breeding programs have traditionally focused on agronomic traits such as yield, disease resistance, and environmental adaptation. Comparatively little attention has been given to understanding the genetic factors that govern protein nutritional quality. As a result, there remains a critical knowledge gap linking plant genetics to nutritionally relevant protein traits.
The objective of this project is to identify and characterize the factors that contribute to protein nutritional quality in pulse crops. Using a diverse collection of pulse genotypes, the project will investigate variation in protein content, amino acid composition, digestibility, and antinutritional factors. The student will explore how these quality traits are associated with genetic and phenotypic variation and identify promising targets for future breeding programs.
The project will combine laboratory-based nutritional analyses with data-driven approaches to better understand the relationships among genetics, composition, and protein quality. Ultimately, this work aims to support the development of pulse varieties with enhanced nutritional value, contributing to healthier diets, more sustainable protein production, and future crop improvement strategies.
Research area, student roles & skills
Research area: I am an Associate Professor in Food Science at the University of Guelph and lead the Food Processing, Structure and Quality Lab. My research focuses on understanding how genetics, growing conditions, and food processing influence the nutritional, functional, and sensory quality of plant-based foods. I integrate food science, nutrition, analytical chemistry, digestion studies, and data science to develop innovative approaches for improving food quality. A major area of interest is identifying the biological and environmental factors that determine protein quality in pulses and other plant-based foods.
Student roles: he student will participate in a research project aimed at understanding the biological and genetic factors that influence protein nutritional quality in pulse crops.
The project will begin with a literature review to identify current knowledge, emerging technologies, and key research gaps related to protein quality, digestibility, amino acid composition, and pulse breeding. The student will assist in developing experimental plans and identifying priority traits for investigation.
The student will participate in laboratory analyses to evaluate protein nutritional quality, including measurements of protein content, amino acid composition, digestibility, and selected antinutritional factors. The student will assist with sample preparation, data collection, quality control procedures, and interpretation of results.
A significant component of the project will involve data analysis. The student will work with phenotypic, nutritional, and compositional datasets to explore relationships among protein quality traits. Depending on the student's background and interests, opportunities may also be available to gain experience with statistical analysis, multivariate analysis, and machine-learning approaches used to identify patterns and predictors of protein quality.
The student will work closely with graduate students and researchers and participate in regular laboratory meetings, research discussions, and project presentations. Through this experience, the student will gain training in food and nutritional analysis, crop quality evaluation, scientific communication, and data-driven research methods relevant to future careers in food science, nutrition, agriculture, and biotechnology.
Skills required: The ideal candidate is an undergraduate student with a background in Food Science, Nutrition, Plant Science, Biology, Biochemistry, Genetics, Agriculture, or a related discipline. Students should have an interest in plant-based foods, nutrition, crop improvement, and data analysis. Previous laboratory experience is beneficial but not required. The successful candidate should be motivated, detail-oriented, and enthusiastic about working at the interface of food science, nutrition, and plant breeding.
4. Financial Planning and Portfolio Optimization with Data Analytics and Artificial Intelligence
Supervisor: Oleksandr Romanko
University: University of Toronto
Location: Toronto, Ontario
Start date: 2027-05-10 (flexible)
Disciplines: Actuarial Science, Business, Computer Science, Econometrics, Engg-Computer, Engg-Industrial, Engg-Software, Engg-Systems and Technology, Engineering, Finance, Management Information Systems, Mathematics, Science and Technology, Statistics, Studies Science and Technology, Sociology, Quantitative Surveying, Banking, Economics, Information Studies
The recent global financial crisis has made risk analytics and financial stability a foremost concern of investors and corporations worldwide. Given the rapidly expanding scope and complexity of risk-aware management and finance, mathematical and software innovation is central to the field. Monte Carlo simulation techniques and numerical optimization, which computes portfolio risks and automates the construction of portfolios that best meet specified requirements, is finding novel uses in the field of finance, investments and risk management. Monte Carlo simulation techniques allow modeling future uncertainty of portfolio value and computing financial risks under different scenarios of future events, and optimization techniques can serve as one of the tools for individual investors, wealth managers and companies to find better solutions and improve decision-making.
While portfolio optimization algorithms are often applied to solve portfolio construction problems for wealth management and pension planning, translating individual preferences of investors expressed in natural language into optimization problem objective functions and constrains remain a challenge. Based on available data, individual goals and personalized use cases we plan to develop an artificial intelligence software tool that helps formulating optimization problems and guiding users through portfolio selection process. Combining visualizations, natural language processing, machine learning, generative AI and quantitative algorithms the tool would interactively guide user throughout their portfolio selection journey allowing them to compare portfolios until they arrive at their "ideal" portfolio choice. It is planned that developed software tool would be deployed on a cloud.
Research area, student roles & skills
Research area: This research area is at the intersection of data science, analytics, machine learning, quantitative modeling and artificial intelligence with finance, wealth management and investment industries. We are using both quantitative and cognitive algorithms to improve decision-making in finance, risk and wealth management.
Student roles: A student will be analyzing data and applying algorithms to develop a cognitive software tool for portfolio construction, optimization and visualization. Writing code, preferably in Python, to apply quantitative, machine learning and artificial intelligence algorithms will be integral part of the student role.
Skills required: Student background would preferably be in a quantitative field such as data science, information technology, computer science, engineering, mathematics, statistics, economics or finance, but students with other backgrounds that have basic knowledge of data analytics are welcome to apply as well. Required skills are basic mathematics and statistics, understanding of optimization and algorithms, and ability to program in Python, R or Matlab (Python preferred). Familiarity with visualization tools and basic knowledge of finance would be an asset, but are not strictly required.
Hydrological risk assessment is more and more important, such as the case of floods. In addition, hydrological events are complex where they can be affected by climate change and hence represent nonstationarity. They can also be represented by more than one variable (multivariate). Some of these issues have been treated in the literature separately. The aim of this project is to provide an overview of these issues in an integrated way as well as providing practical solutions especially in terms of programing.
Research area, student roles & skills
Research area: Data Science in environnemental applications, including machine learning algorithms and statistical approaches to deal with a large number of applications such as hydrological risk, water demand forecasting, climate effects on health.
Student roles: - Coding new methods - Building R packages based on previous codes - Testing the codes on real data sets - Writing documentation, tutorials and guidelines
Skills required: - Programing in R - Statistical methods
6. Machine Learning Techniques for Analyzing Automobile Statistical Plan Data
Machine learning (ML) algorithms have been powerful tools in making decisions and predictions in real-world complex systems, e.g., insurance pricing, medical diagnosis and imaging, the spread of infectious diseases, consumer buying behaviours, banking fraud detection and prevention. However, in most cases, these algorithms still need to sufficiently explain how they reached the results. The research goal is the development of novel statistical approaches for making machine learning techniques more interpretable for analyzing Automobile Statistical Plan data. This research aims to develop novel methods to make machine learning models more interpretable, particularly complex neural networks. The interns will be involved in the development of new interpretable machine-learning techniques at all stages of research. They will be responsible for the implementation of techniques in either R or Python. By conducting this research, the interns will gain experience in analyzing insurance loss data and using machine learning Techniques for solving real-world problems.
Research area, student roles & skills
Research area: Statistical Machine Learning, Explainable Data Analytics, Risk Modeling, Auto Insurance Pricing, Multivariate Statistical Methods, Time Series Analysis, Predictive Analytics, and Health Informatics.
Student roles: The primary role of the student will be working as a data analyst, who is responsible for data manipulation, and statistical analysis and modeling. The students will also conduct a literature review for the research project and write a report at the end of this study.
More specifically, the student is expected to do the following:
A. Literature review: Published works related to applications of modern statistical models to auto insurance loss data will be reviewed and surveyed by the interns.
B. Data collection: It is necessary to use reliable data to develop robust and precise predictive models. The required data for modelling and analysis will be obtained from the Insurance Bureau of Canada. The students will also be expected to take the initiative in searching the relevant information or data from available sources related to this research for a potential comparative study, particularly the data associated with UBI.
C. Data preprocessing: Before introducing the datasets into the modelling algorithms, data preprocessing (data cleaning, data normalization, data visualization, etc.) will be performed by the student to ensure that the datasets are consistent and usable.
D. Model development: The student will assist in constructing predictive models and implementing the statistical models using R software.
E. Model assessment: The student will apply the cross-validation methods (e.g., k-fold validation, re-sample approach, etc.) to assess the reliability of the models. After model validation, the student will conduct a comparative study using the data from different reporting years.
F. Scientific Writing: The student will materialize the obtained results and their analysis. Bi-weekly reports and a final project report are two key deliverable items for the research interns.
Skills required: Students should have knowledge and skills in organizing data and conducting statistical analysis and modelling using both statistical and computational approaches. Students are expected to know statistical tools such as R, and basic knowledge in statistical machine learning. Students are also expected to be able to write a scientific report that summarizes the main findings of the study. Advanced statistical and actuarial background will be considered an asset.
7. Machine Learning for National Collision Data from Canada
Car accidents are a pressing public health concern and are considered one of the leading causes of death and disability. The World Health Organization's Global Status Report on Road Safety 2018 emphasized the urgent need to improve road safety and reduce the number of fatalities and injuries. Therefore, analyzing and predicting the fatal events of car accidents is important. It enables health organizations to allocate their resources effectively, ultimately reducing the fatalities resulting from car accidents. As a result, various data management and reporting systems have been developed to collect data related to car accidents. The National Collision Database (NCDB), an open database in Canada, contains data on police-reported motor vehicle collisions on public roads. These databases contain various variables and cover fatal and injury crashes from 1999 to the most recent available data, i.e. 2021. These invaluable resources allow for examining the underlying risk factors and their interrelationship with the fatality rate, enabling a better understanding of the driving forces behind auto collision fatalities. By leveraging these databases, researchers can study and identify potential strategies to mitigate risks and improve road safety, thereby reducing car accident deaths.
Research area, student roles & skills
Research area: I am conducting research in Auto Insurance Pricing and Regulation, Computational Intelligence, and Machine Learning with applications. My main interest is in developing interpretable statistical machine learning techniques that have a wide impact on real-world applications.
Student roles: Students are expected to conduct research by developing and implementing methodologies or algorithms with minimal supervision. Students will report the findings in writing and communicate with the supervisor or other research team members.
Skills required: Statistical machine learning techniques, including regression, clustering and classification, are highly desired. Experience and the Ability to materialize the results into a research article are expected. Experience in using R or other statistical packages is required.
8. Mathematical modelling of climate-related health and insurance risks
Supervisor: Jérémie Boudreault
University: Université Laval (Québec campus)
Location: Québec, Québec
Start date: 2027-05-17 (flexible)
Disciplines: Actuarial Science, Mathematics, Statistics, Environmental Studies, Computer Science, Public Health
In this project, we aim to develop mathematical models of the effects of climate variables (such as temperature, precipitation, wildfire, etc.) for applications on both human health and insured assets (e.g., homes, cars). To this end, we consider various existing techniques including statistical, epidemiological, geospatial and machine/deep learning models. In a second stage of the project, we apply the models to various outcomes such as mortality, morbidity, home damage or car accidents. The expected results should be helpful to insurers, pensions funds and public health authorities to better understand climate-related risks.
Research area, student roles & skills
Research area: I am an assistant professor in data science and climate risk at Université Laval. My research focuses on modelling the health and economic impacts of climate hazards such as extreme temperatures, wildfires and flooding. To that end, I develop novel statistical models and AI/machine learning approaches.
Student roles: The trainee will be part of a multidisciplinary team. Their role will be: - Reading scientific articles and discussing findings during lab meetings - Developing mathematical models (e.g., probabilistic, dependence, regression, etc.) - Analyzing different types of data (e.g., climate, health and claims) with R or Python - Producing interactives reports (e.g., Rmarkdown, Leaflet, Shiny) - Writing parts of reports and articles, and scientific presentations - Collaborating with MSc and PhD students
Skills required: Skills (at least three) : - Knowledge in mathematical or statistical modelling - Knowledge in data science, machine or deep learning models - Previous work with time series or geospatial data - Programming language such as R or Python - Interest for actuarial, health or environnemental research
The aim is to develop predictive models (data science models) for drinking water consumption. These models aim to regulate flow and pressures in distribution networks, proactive management of water needs, energy optimization of pumps and detection of major leaks. Climate change, population growth and sprawl urban is putting increasing pressure on water services and resources. Several approaches have been developed to secure and optimize the supply of drinking water. These applications could perform better based on reliable predictions of water demand.
Several of the existing models for predicting water demand have many limitations, in the face of trends (population growth, global warming), has some irregularities in cycles and aberrant values (due to climate, socio-cultural events, pipe breaks), heterogeneity (industrial vs. residential) and lack of data (case of small municipalities). With the proposed models, forecasts can be made for a large number of municipalities. We propose methodologies with practical benefits directly usable, but also economical for municipalities in Canada.
Research area, student roles & skills
Research area: Data Science in environnemental applications, including machine learning algorithms and statistical approaches to deal with a large number of applications such as hydrological risk, water demand forecasting, climate effects on health.
Student roles: - Coding new methods - Building Python codes based on previous codes from the team - Testing the codes on real data sets - Writing documentation, tutorials and guidelines
Skills required: - Programing in Python - Statistical methods - Machine learning - Data science