Artificial intelligence · Data mining · Optimization
Research contributions from methods to applications
My publications span artificial intelligence, machine learning, data mining, optimization, and applied data science, with an emphasis on reproducible methods and open-source software.
Machine Learning and AI
| [1] |
Antônio Junior Alves Caiado and Michael Hahsler.
Dropout robustness and cognitive profiling of transformer models via stochastic inference.
2603.17811 [cs.AI], 2026.
[ DOI |
at the publisher ]
Transformer-based language models are widely deployed for reasoning, yet their behavior under inference-time stochasticity remains underexplored. While dropout is common during training, its inference-time effects via Monte Carlo sampling lack systematic evaluation across architectures, limiting understanding of model reliability in uncertainty-aware applications. This work analyzes dropout-induced variability across 19 transformer models using MC Dropout with 100 stochastic forward passes per sample. Dropout robustness is defined as maintaining high accuracy and stable predictions under stochastic inference, measured by standard deviation of per-run accuracies. A cognitive decomposition framework disentangles performance into memory and reasoning components. Experiments span five dropout configurations yielding 95 unique evaluations on 1,000 samples. Results reveal substantial architectural variation. Smaller models demonstrate perfect prediction stability while medium-sized models exhibit notable volatility. Mid-sized models achieve the best overall performance; larger models excel at memory tasks. Critically, 53% of models suffer severe accuracy degradation under baseline MC Dropout, with task-specialized models losing up to 24 percentage points, indicating unsuitability for uncertainty quantification in these architectures. Asymmetric effects emerge: high dropout reduces memory accuracy by 27 percentage points while reasoning degrades only 1 point, suggesting memory tasks rely on stable representations that dropout disrupts. 84% of models demonstrate memory-biased performance. This provides the first comprehensive MC Dropout benchmark for transformers, revealing dropout robustness is architecture-dependent and uncorrelated with scale. The cognitive profiling framework offers actionable guidance for model selection in uncertainty-aware applications. |
| [2] |
Zerui Ma, Michael Hahsler, and Peter Moore.
A recommender system architecture for university curriculum advising.
Proceedings of the AAAI Symposium Series, 5(1):235--241, May 2025.
[ DOI ]
Effective academic advising plays a crucial role in student success, yet universities face challenges in optimizing advising processes and course enrollment. This task is complicated by the fact that several graduation requirements have to be met while also taking the students’ interests into account. Academic advising has historically been performed by a skilled human adviser. Universities can optimize course planning and help students make informed decisions about their academic path with recommender systems. This case study develops a goal-based agent recommender system based on a large language model tailored to undergraduate students, depending on curriculum requirements, prerequisite dependencies, and student preferences. The developed recommendation system helps universities increase student advising efficiency and create more intuitive and student-centric curricula. We show how to structure and process complex curriculum data to create an algorithm-ready environment, simplifying the relationships between degree requirements and course offerings. This study evaluates multiple algorithms based on recommendation accuracy, computational efficiency, and their ability to meet degree requirements while fostering academic engagement. By streamlining course selection and exploring possible degree paths, the system may also help students graduate on time and navigate complex curricula. This system also collects important metrics to accurately predict student enrollment for classes, enabling college departments to plan their course offerings better. The system poses a significant benefit to university advising offices by reducing advisor workloads and encouraging student engagement, advancing the academic achievement of the entire student body. |
| [3] |
Michael Hahsler and Anthony R. Cassandra.
Pomdp: A computational infrastructure for partially observable Markov decision processes.
R Journal, 16:116--133, 2025.
[ DOI ]
Many important problems involve decision-making under uncertainty. For example, a medical professional needs to make decisions about the best treatment option based on limited information about the current state of the patient and uncertainty about outcomes. Different approaches have been developed by the applied mathematics, operations research, and artificial intelligence communities to address this difficult class of decision-making problems. This paper presents the pomdp package, which provides a computational infrastructure for an approach called the partially observable Markov decision process (POMDP), which models the problem as a discrete-time stochastic control process. The package lets the user specify POMDPs using familiar R syntax, apply state-of-the-art POMDP solvers, and then take full advantage of R's range of capabilities, including statistical analysis, simulation, and visualization, to work with the resulting models. |
| [4] |
Antônio Junior Alves Caiado and Michael Hahsler.
AI content self-detection for transformer-based large language models.
In Kohei Arai, editor, Intelligent Systems and Applications, volume 1554 of Lecture Notes in Networks and Systems, pages 147--168, Cham, 2025. Springer Nature Switzerland.
[ DOI |
preprint (PDF) ]
The usage of generative artificial intelligence (AI) tools based on large language models, including ChatGPT, Bard, and Claude, for text generation has many exciting applications with the potential for phenomenal productivity gains. One issue is authorship attribution when using AI tools. This is especially important in an academic setting, where the inappropriate use of generative AI tools may hinder student learning or stifle research by creating a large amount of automatically generated derivative work. Existing plagiarism detection systems can trace the source of submitted text but are not yet equipped with methods to accurately detect AI-generated text. This study introduces the idea of direct origin detection and evaluates whether generative AI systems can recognize their output and distinguish it from human-written texts. We argue, why current transformer-based models may be able to self-detect their own generated text and perform a small empirical study using zero-shot learning to investigate if that is the case. Results reveal varying capabilities of AI systems to identify their generated text. Google's Bard model exhibits the largest capability of self-detection with an accuracy of 94 On the other hand, Anthropic's Claude model seems to be not able to self-detect. |
| [5] |
Michael Hahsler.
An R companion for introduction to data mining.
Journal of Open Source Education, 7(82):223, 2024.
[ DOI |
at the publisher ]
An R Companion for Introduction to Data Mining is an open-source learning and teaching resource that covers how to implement data mining concepts using R. It is designed to accompany the popular data mining textbook Introduction to Data Mining (Tan et al., 2017) to study the implementation of the basic data mining concepts including data preparation, classification, clustering, and association analysis. The resource uses complete, annotated examples to demonstrate how data mining concepts can be translated into R code. |
| [6] |
Farzad Kamalzadeh, Vishal Ahuja, Michael Hahsler, and Michael E. Bowen.
An analytics-driven approach for optimal individualized diabetes screening.
Production and Operations Management, 30(9):3161--3191, September 2021.
[ DOI |
preprint (PDF) |
at the publisher ]
Type 2 diabetes is a chronic disease that affects millions of Americans and puts a significant burden on the healthcare system. The medical community sees screening patients to identify and treat prediabetes and diabetes early as an important goal; however, universal population screening is operationally not feasible, and screening policies need to take characteristics of the patient population into account. For instance, the screening policy for a population in an affluent neighborhood may differ from that of a safety-net hospital. The problem of optimal diabetes screening—whom to screen and when to screen—is clearly important, and small improvements could have an enormous impact. However, the problem is typically only discussed from a practical viewpoint in the medical literature; a thorough theoretical framework from an operational viewpoint is largely missing. In this study, we propose an approach that builds on multiple methods—partially observable Markov decision process (POMDP), hidden Markov model (HMM), and predictive risk modeling (PRM). It uses available clinical information, in the form of electronic health records (EHRs), on specific patient populations to derive an optimal policy, which is used to generate screening decisions, individualized for each patient. The POMDP model, used for determining optimal decisions, lies at the core of our approach. We use HMM to estimate the cohort-specific progression of diabetes (i.e., transition probability matrix) and the emission matrix. We use PRM to generate observations—in the form of individualized risk scores—for the POMDP. Both HMM and PRM are learned from EHR data. Our approach is unique because (i) it introduces a novel way of incorporating predictive modeling into a formal decision framework to derive an optimal screening policy; and (ii) it is based on real clinical data. We fit our model using data on a cohort of more than 60,000 patients over 5 years from a large safety-net health system and then demonstrate the model’s utility by conducting a simulation study. The results indicate that our proposed screening policy outperforms existing guidelines widely used in clinical practice. Our estimates suggest that implementing our policy for the studied cohort would add one quality-adjusted life year for every patient, and at a cost that is 35% lower, compared with existing guidelines. Our proposed framework is generalizable to other chronic diseases, such as cancer and HIV. |
| [7] |
Xinyi Ding, Zohreh Raziei, Eric C. Larson, Eli V. Olinick, Paul Krueger, and Michael Hahsler.
Swapped face detection using deep learning and subjective assessment.
EURASIP Journal on Information Security, 2020(6):1--12, May 2020.
[ DOI ]
The tremendous success of deep learning for imaging applications has resulted in numerous beneficial advances. Unfortunately, this success has also been a catalyst for malicious uses such as photo-realistic face swapping of parties without consent. In this study, we use deep transfer learning for face swapping detection, showing true positive rates greater than 96% with very few false alarms. Distinguished from existing methods that only provide detection accuracy, we also provide uncertainty for each prediction, which is critical for trust in the deployment of such detection systems. Moreover, we provide a comparison to human subjects. To capture human recognition performance, we build a website to collect pairwise comparisons of images from human subjects. Based on these comparisons, we infer a consensus ranking from the image perceived as most real to the image perceived as most fake. Overall, the results show the effectiveness of our method. As part of this study, we create a novel dataset that is, to the best of our knowledge, the largest swapped face dataset created using still images. This dataset will be available for academic research use per request. Our goal of this study is to inspire more research in the field of image forensics through the creation of a dataset and initial analysis. |
| [8] |
Paul S. Krueger, Michael Hahsler, Eli V. Olinick, Sheila H. Williams, and Mohammadreza Zharfa.
Quantitative classification of vortical flows based on topological features using graph matching.
Proceedings of the Royal Society A, 475(2228):1--16, August 2019.
[ DOI ]
Vortical flow patterns generated by swimming animals or flow separation (e.g. behind bluff objects such as cylinders) provide important insight to global flow behaviour such as fluid dynamic drag or propulsive performance. The present work introduces a new method for quantitatively comparing and classifying flow fields using a novel graph-theoretic concept, called a weighted Gabriel graph, that employs critical points of the velocity vector field, which identify key flow features such as vortices, as graph vertices. The edges (connections between vertices) and edge weights of the weighted Gabriel graph encode local geometric structure. The resulting graph exhibits robustness to minor changes in the flow fields. Dissimilarity between flow fields is quantified by finding the best match (minimum difference) in weights of matched graph edges under relevant constraints on the properties of the edge vertices, and flows are classified using hierarchical clustering based on computed dissimilarity. Application of this approach to a set of artificially generated, periodic vortical flows demonstrates high classification accuracy, even for large perturbations, and insensitivity to scale variations and number of periods in the periodic flow pattern. The generality of the approach allows for comparison of flows generated by very different means (e.g. different animal species). |
| [9] |
Usama El Shamy and Michael Hahsler.
Data analytics applied to a microscale simulation model of soil liquefaction.
In Geotechnical Earthquake Engineering and Soil Dynamics V. ASCE, June 2018.
[ DOI ]
Recent computational models create a large amount of data, which can be hard to analyze. In this paper, we demonstrate the power of employing data analytics techniques to characterize soil behavior during liquefaction. We used simple simulation output aggregated over several locations along the depth of the deposit. Available were five simulated quantities (shear strain, acceleration, coordination number, pore pressure, change in volume) and we added four change features (direction and magnitude of change of the quantity). We then performed data mining using conventional analytics methods (clustering using k-means with k = 5 and principal components analysis). The clustering visualization showed that a visible break started to propagate downward on the onset of liquefaction through all depth locations. Using this scheme, we were able to build cluster profiles that generate new insights by visualizing details of the transformation process of the soil from a solid state to a liquefied state. |
| [10] |
Jake Drew, Michael Hahsler, and Tyler Moore.
Polymorphic malware detection using sequence classification methods.
EURASIP Journal on Information Security, 2017(1):1--12, January 2017.
[ DOI ]
Identifying malicious software executables is made difficult by the constant adaptations introduced by miscreants in order to evade detection by antivirus software. Such changes are akin to mutations in biological sequences. Recently, high-throughput methods for gene sequence classification have been developed by the bioinformatics and computational biology communities. In this paper, we apply methods designed for gene sequencing to detect malware in a manner robust to attacker adaptations. Whereas most gene classification tools are optimized for and restricted to an alphabet of four letters (nucleic acids), we have selected the Strand gene sequence classifier for malware classification. Strand's design can easily accommodate unstructured data with any alphabet, including source code or compiled machine code. To demonstrate that gene sequence classification tools are suitable for classifying malware, we apply Strand to approximately 500GB of malware data provided by the Kaggle Microsoft Malware Classification Challenge (BIG 2015) used for predicting 9 classes of polymorphic malware. Experiments show that, with minimal adaptation, the method achieves accuracy levels well above 95% requiring only a fraction of the training times used by the winning team's method. |
| [11] |
Jake Drew, Michael Hahsler, and Tyler Moore.
Polymorphic malware detection using sequence classification methods.
In International Workshop on Bio-inspired Security, Trust, Assurance and Resilience (BioSTAR 2016), May 2016.
[ preprint (PDF) ]
Polymorphic malware detection is challenging due to the continual mutations miscreants introduce to successive instances of a particular virus. Such changes are akin to mutations in biological sequences. Recently, high-throughput methods for gene sequence classification have been developed by the bioinformatics and computational biology communities. In this paper, we argue that these methods can be usefully applied to malware detection. Unfortunately, gene classification tools are usually optimized for and restricted to an alphabet of four letters (nucleic acids). Consequently, we have selected the Strand gene sequence classifier, which offers a robust classification strategy that can easily accommodate unstructured data with any alphabet including source code or compiled machine code. To demonstrate Stand's suitability for classifying malware, we execute it on approximately 500GB of malware data provided by the Kaggle Microsoft Malware Classification Challenge (BIG 2015) used for predicting 9 classes of polymorphic malware. Experiments show that, with minimal adaptation, the method achieves accuracy levels well above 95% requiring only a fraction of the training times used by the winning team's method. |
| [12] |
Sudheer Chelluboina and Michael Hahsler.
Trajectory segmentation using oblique envelopes.
In 2015 IEEE International Conference on Information Reuse and Integration (IRI), pages 470--475. IEEE, August 2015.
[ DOI ]
Trajectory segmentation, i.e., breaking the trajectory into sub-trajectories, is a fundamental task needed for many applications dealing with moving objects. Several methods for trajectory segmentation, e.g., based on minimum description length (MDL), have been proposed. In this paper, we develop a novel technique for trajectory segmentation which created a series of oblique envelopes to partition the trajectory into sub-trajectories. Experiments with the new algorithm on hurricane trajectory data, taxi GPS data and simulated data of tracks show that oblique envelopes out-perform MDL-based trajectory segmentation. |