Sandrine Dudoit
high-dimensional statistical learning, statistical computing, computational biology and genomics, precision health and medicine
My research and teaching activities concern the development and application of statistical learning methods and software for the analysis of high-throughput -omic data in both basic biology and precision health and medicine.
Statistical methodology. My methodological research interests regard high-dimensional statistical learning and include exploratory data analysis (EDA), unsupervised learning (e.g., cluster analysis, dimensionality reduction), loss-based estimation with cross-validation (e.g., in density estimation, classification, regression, model selection), and causal inference.
Applications to biomedical and genomic research. My methodological work is motivated in large part by statistical learning questions arising in biological and medical research and, in particular, high-throughput gene expression studies at single-cell and spatial resolution. My contributions span a broad range of questions throughout the data science pipeline, of both practical relevance and theoretical interest: experimental design, EDA, normalization, expression quantitation, differential expression analysis, biomarker and treatment effect modifier discovery, class discovery, class prediction, inference of cell lineages, and integration of biological annotation metadata (e.g., Gene Ontology annotation, gene regulatory networks).
Statistical computing. I am also interested in statistical computing and, in particular, reproducible research. I am a founding core developer of the Bioconductor Project (http://www.bioconductor.org), an international collaborative effort for the design and deployment of an open-source and open-development software ecosystem for the analysis of biomedical and -omic data.