Artificial intelligence · Data mining · Optimization

Research contributions from methods to applications

My publications span artificial intelligence, machine learning, data mining, optimization, and applied data science, with an emphasis on reproducible methods and open-source software.

Earth Science Data Research

[1] Usama El Shamy and Michael Hahsler. Data analytics applied to a microscale simulation model of soil liquefaction. In Geotechnical Earthquake Engineering and Soil Dynamics V. ASCE, June 2018. [ DOI ]
Recent computational models create a large amount of data, which can be hard to analyze. In this paper, we demonstrate the power of employing data analytics techniques to characterize soil behavior during liquefaction. We used simple simulation output aggregated over several locations along the depth of the deposit. Available were five simulated quantities (shear strain, acceleration, coordination number, pore pressure, change in volume) and we added four change features (direction and magnitude of change of the quantity). We then performed data mining using conventional analytics methods (clustering using k-means with k = 5 and principal components analysis). The clustering visualization showed that a visible break started to propagate downward on the onset of liquefaction through all depth locations. Using this scheme, we were able to build cluster profiles that generate new insights by visualizing details of the transformation process of the soil from a solid state to a liquefied state.
[2] Shaiba Hadil and Michael Hahsler. A comparison of machine learning methods for predicting tropical cyclone rapid intensification events. Research Journal of Applied Sciences, Engineering and Technology, 13(8):638--651, 2016. [ DOI ]
The aim of this study is to improve the intensity prediction of hurricanes by accounting for rapid intensification (RI) events. Modern machine learning methods offer much promise for predicting meteorological events. One application is providing timely and accurate predictions of tropical cyclone (TC) behavior, which is crucial for saving lives and reducing damage to property. Current TC track prediction models perform much better than intensity (wind speed) models. This is partially due to the existence of RI events. An RI event is defined as a sudden change in the maximum sustained wind speed of 30 knots or greater within 24 hours. Forecasting RI events is so important that it has been put on the National Hurricane Center top forecast priority list. The research on the use of machine learning methods for RI prediction is currently very limited. In this paper, we investigate the potential of popular machine learning methods to predict RI events. The evaluated models include support vector machines, logistic regression, naïve-Bayes classifiers, classification and regression trees and a wide range of ensemble methods including boosting and stacking. We also investigate dimensionality reduction and feature selection, and we address class imbalance using the Synthetic Minority Over-sampling Technique (SMOTE). The evaluation shows that some of the investigated models improve over the current operational Rapid Intensification Index model. Finally, we use RI predictions to make improved storm intensity predictions.
[3] Hadil Shaiba and Michael Hahsler. An experimental comparison of different classifiers for predicting tropical cyclone rapid intensification events. In Proceedings of the International Conference on Machine Learning, Electrical and Mechanical Engineering (ICMLEME'2014), Dubai, UAE, January 2014. [ preprint (PDF) ]
Accurately predicting the intensity and track of a Tropical Cyclone (TC) can save the lives of people and help to significantly reduce damage to property and infrastructure. Current track prediction models outperform intensity models which is partially due to the existence of rapid intensification (RI) events. RI appears in the lifecycle of most major hurricanes and can be defined as a change in intensity within 24 hours which exceeds 30 knots. Improving the predicting of RI events has been identified as one of the top priority problems by the National Hurricane Center (NHC). In this paper we compare the RI event prediction performance of several popular classification methods: Logistic regression, naive Bayes classifier, classification and regression tree (CART), and support vector machine. The dataset used is derived from the data used by the Statistical Hurricane Intensity Prediction Scheme (SHIPS) model for intensity prediction which contains large-scale weather, ocean, and earth condition predictors from 1982 to 2011. 10-fold cross validation is applied to compare the models. The probability of detection (POD) and false alarm ratio (FAR) are used to measure performance. Predicting RI events is a difficult problem but initial experiments show potential for improving forecast using data mining and machine learning techniques.
[4] Hadil Shaiba and Michael Hahsler. Intensity prediction model for tropical cyclone rapid intensification events. In Proceedings of the IADIS Applied Computing 2013 (AC 2013) Conference, Fort Worth, TX, October 2013.
Tropical Cyclones (TC) create strong wind and rain and can cause significant human and financial losses. Many major hurricanes in the Atlantic Ocean undergo rapid intensification (RI). RI events happen when the strength of the storm increases rapidly within 24 hours. Improving the hurricane’s intensity prediction model by accurately detecting the occurrence and predicting the intensity of an RI event can help avoid human and financial losses. In this paper we analyzed RI events in the Atlantic Basin and investigated the use of a combination of three different models to predict RI events. The first and second models are simple location and time-based models which use the conditional probability of intensification given the day of the year and the location of the storm. The third model uses a data mining technique which is based on the Extensible Markov Chain Model (EMM) which clusters the hurricane’s lifecycle into states and then uses the transition probabilities between these states for prediction. One of the main characteristics of a Markov Chain is that the next state depends only on the current state and by knowing the current state we can predict future states which provide us with estimates for future intensities. In future research we plan to test each model independently and combinations of the models by comparing them to the current best prediction models.
[5] Vladimir Jovanovic, Margaret H. Dunham, Michael Hahsler, and Yu Su. Evaluating hurricane intensity prediction techniques in real time. In Third IEEE ICDM Workshop on Knowledge Discovery from Climate Data, Proceedings of the of the 2011 IEEE International Conference on Data Mining Workshops (ICDMW 2011), pages 23--29. IEEE, December 2011. [ DOI | preprint (PDF) ]
While the accuracy of hurricane track prediction has been improving, predicting intensity, the maximum sustained wind speed, is still a very difficult challenge. This is problematic because the destructive power of a hurricane is directly related to its intensity. In this paper, we present Prediction Intensity Interval model for Hurricanes (PIIH) which combines sophisticated data mining techniques to create an online real time model for accurate intensity predictions and we present a web-based framework to dynamically compare PIIH to operational models used by the National Hurricane Center (NHC). The created dynamic website tracks, compares, and provides visualization to facilitate immediate comparisons of prediction techniques. This paper is a work in progress paper reporting on both, new features of the PIIH model and online visualization of the accuracy of that model as compared to other techniques.
[6] Yu Su, Sudheer Chelluboina, Michael Hahsler, and Margaret H. Dunham. A new data mining model for hurricane intensity prediction. In Second IEEE ICDM Workshop on Knowledge Discovery from Climate Data: Prediction, Extremes and Impacts, Proceedings of the of the 2010 IEEE International Conference on Data Mining Workshops (ICDMW 2010), pages 98--105. IEEE, December 2010. [ DOI ]
This paper proposes a new hurricane intensity prediction model, WFL-EMM, which is based on the data mining techniques of feature weight learning (WFL) and Extensible Markov Model (EMM). The data features used are those employed by one of the most popular intensity prediction models, SHIPS. In our algorithm, the weights of the features are learned by a genetic algorithm (GA) using historical hurricane data. As the GA's fitness function we use the error of the intensity prediction by an EMM learned using given feature weights. For fitness calculation we use a technique similar to k-fold cross validation on the training data. The best weights obtained by the genetic algorithm are used to build an EMM with all training data. This EMM is then applied to predict the hurricane intensities and compute prediction errors for the test data. Using historical data for the named Atlantic tropical cyclones from 1982 to 2003, experiments demonstrate that WFL-EMM provides significantly more accurate intensity predictions than SHIPS within 72 hours. Since we report here first results, we indicate how to improve WFL-EMM in the future.