Artificial intelligence · Optimization · Data science
Turning complex data into useful information and decisions
I develop machine-learning and optimization methods for discovering structure in complex data and making decisions under uncertainty. My work connects methodological research with reproducible open-source software and applications in healthcare, bioinformatics, earth science, and engineering.
Research themes
Discovering structure
Pattern discovery and machine learning
Methods for association-rule and sequence mining, data-stream clustering, recommender systems, density-based clustering, and interpretable data visualization.
Choosing actions
Decision-making and optimization
Models and algorithms for reinforcement learning, Markov and partially observable Markov decision processes, optimal ordering, scheduling, seriation, and routing.
Putting methods to work
Applied data science
Collaborative applications in healthcare analytics, bioinformatics, quantitative marketing, earth science, manufacturing, and engineering.
Featured recent work
Healthcare analytics · 2025
Improving access to kidney care
Analytics and simulation-optimization models reveal barriers in emergent dialysis and study how transplant oversight affects waitlist management.
Artificial intelligence · 2025
Attributing AI-generated content
Direct-origin detection tests whether transformer-based language models can distinguish their own output from human-written text.
Decision-making · 2024
POMDP infrastructure for R
A computational environment for defining, solving, simulating, and analyzing partially observable Markov decision processes.
Browse all publications Google Scholar
Selected funded projects
2021–2022 · NSF
Data Science Supplemental
Data-science methods for evaluating the liquefaction potential of saturated granular soils under partial drainage conditions; supplement to CMMI-1728612 with Usama El Shamy.
2017–2020 · NIST
SAFE-NET
An integrated connected-vehicle and computing platform for public-safety applications (60NANB17D180).
2011–2014 · NIH/NHGRI
QuasiAlign
Position-sensitive p-mer frequency clustering for efficient, alignment-free classification and differentiation of biological sequences (R21HG005912).
2009–2013 · NSF
TRACDS
Temporal relationships among clusters in data streams, including models for tracking evolving populations and predicting hurricane intensity (IIS-0948893).
Patent
2014 · US20140344195A1
System and method for machine learning and classifying data
Introduces a method for clasifying large scale sequebnce data using MinHash and MapReduce.
Support
This research has received support from the following organizations.