Research Communication: Cover page, abstract, purpose statement, qualification statement, progress report memo, annotated bibliography
ALL ABSTRACTS
On an Unbiased Ridge Regression Estimator for the Stochastic Restricted Linear Regression Model
B M Golam Kibria, Ph.D., Professor, Dept. of Mathematics and Statistics, FIU A.K. (Presenter) and M.I. Alheety, Department of Mathematics, University of Anbar, Iraq
Abstract: Crouse et al. (1995) proposed an unbiased ridge regression estimator for the multicollinear linear regression model. Ozkale and Kacıranlar (2007) proposed two‐parameter estimator and Jibo Wu (2014) introduced an unbiased two‐parameter estimator based on prior information. Yalian and Hu (2011) proposed a new ridge‐type estimator by unifying the sample and prior information in linear model with additional stochastic linear restrictions. This paper considers an unbiased two‐parameter estimator for the stochastic restricted linear regression model. The properties and the performance of the proposed estimator compared to other common estimators using the mean squares error criterion are presented. A numerical example is given to illustrate the findings of the paper. Decomposition of Extremal Dependence and Application to US Extreme Precipitation
Data Keynote Speaker: Daniel S. Cooley, PhD., Professor, Department of Statistics, CSU
Abstract: Principal components (alternatively empirical orthogonal functions) are widely used to study modes of dependence in meteorological phenomena. However, PCA/EOFs arise by decomposing the covariance matrix, and therefore are not well‐suited for describing extremal dependence. Characterizing extreme dependence in high dimensions is difficult for most existing multivariate extreme modeling frameworks. Using recently developed methods, we summarize extremal dependence via the tail pairwise dependence matrix (TPDM), which can be seen as an extreme analog to the covariance matrix. An eigendecomposition of the TPDM results in an ordered orthonormal basis through which the modes of extremal dependence can be studied. Applying these methods to US precipitation data during the hurricane season, we investigate relationships between the extremes of the time series of basis coefficients and climatological drivers such as the El Nino/La Nina oscillation.
On Introduction to Flexible Hyperbolastic Growth and Survival Models Zoran Bursac, Ph.D., Professor, Chair and Director of the Consulting Center,
Department of Biostatistics, Robert Stempel College of Public Health & Social Work, FIU Abstract: The S‐shaped sigmoidal growth models have been extensively studied and applied in a wide range of medical and biological studies. We introduce a family of three and four parameter models called Hyperbolastic models for analyzing growth behavior of self‐limited growth. As a continuation of this work, we develop and introduce two related parametric Hyperbolastic survival models. To illustrate the application and utility of these models and to gain a more complete understanding of them, we apply these models to several sets of motivating data and compare their performance relative to other established and widely used growth and survival models.
An Empirical Comparison of Several Estimators for the Shape and Scale Parameters of the Two‐Parameter Weibull Distribution
Sergio Perez‐Melo, MS, (Presenter) Instructor, Division of Statistics of the Department of Mathematics and Statistics
Sneh Gulati, Ph.D. Professor, Division of Statistics, Department of Mathematics and Statistics
Abstract: The Weibull distribution is often the model of choice for data arising in different fields such as survival analysis, reliability engineering, extreme value theory, weather forecasting, hydrology, actuarial sciences, among many others. Several methods have been proposed to estimate the parameters of the Weibull distribution over the years leading to a plethora of papers. In this talk, several estimation methods for the shape and scale parameters of the two‐parameter Weibull distribution are reviewed and compared based on the mean square error. Because a theoretical comparison would be formidable given the number of estimators being considered, an extensive simulation study was used to compare them. It was observed that the Anderson‐Darling Maximum Goodness of Fit, Median Rank Regression, Quantile Least Squares and the Maximum Likelihood Estimators had better performance than other methods. The Anderson‐Darling Maximum Goodness of Fit, Median Rank Regression and Quantile Least Squares were superior to the Maximum Likelihood in small sample size situations for the shape parameter estimation. For sample sizes of more than 30 most of the analyzed methods performed equally well, with the exception of percentile based methods and the least absolute deviation method. Two real life data sets from software reliability and atmospheric science are used to illustrate the findings.
Predictive Analytics in Business: Case of Healthcare Management Tala Mirzaei, Ph.D.
Assistant Professor, Dept. of Information Systems and Business Analytics, FIU Abstract: Data analytics provide the ability to systematically identify patterns and insights from a variety of data as organizations pursue improvements in their processes, products, and services. When applied to the field of healthcare, analytics presents a new frontier for business intelligence. The principal question in health care management is about how advanced analytics can be used to improve medical outcomes, increase financial performance, and deepen the relationship between patients and care providers. We used state level data from the State Health Plan and implemented predictive mining techniques to identify the factors associated with hospital readmissions, which led to high cost and lower quality for the hospitals. We focused on heart failure, the number one cause of death and the biggest contributor to healthcare costs in the United States. Reduction of High Dimensional Data: A New Method for Reducing Data Dimensionality
in Linear Regression Eliser Nodarse, MS Candidate
Instructor, Division of Statistics, Department of Mathematics and Statistics, FIU Abstract: Regression is a statistical technique for modeling the relationship between a dependent variable Y and two or more predictor variables, also known as regressors. In the broad field of regression, there exists a special case in which the relationship between the dependent variable and the regressor(s) is linear. This is known as linear regression. The purpose of this thesis is to present a useful method that effectively selects a subset of regressors when dealing with high dimensional data and/or collinearity in linear regression. As the name depicts it, high dimensional data occurs when the number of predictor variables is far too large to use commonly known methods. Collinearity, on the other hand, occurs when there exists a linear relationship amongst one or more pairs of independent variables. The method we created, named Best Probable Subset, selects a subset of regressors using a probability scheme that depends on the linear relationship between the response variable and each predictor variable. We then compare the resulting “best probable subset” with other methods, including forward, backward and best subset selection.