###
Do Resampling Estimates Have Low Correlation to the Truth?

The Answer May Shock You. One criticism that is often leveled against using resampling methods (such as cross-validation) to measure model performance is that there is no correlation between the CV results and the true error rate. Let’s look at this with some simulated data. While this assertion is often correct, there are a few […]

###
Intro to Caret: Pre-Processing

Editor’s note: This is the third of a series of posts on the caret package. Creating Dummy Variables Zero- and Near Zero-Variance Predictors Identifying Correlated Predictors Linear Dependencies The preProcess Function Centering and Scaling Imputation Transforming Predictors Putting It All Together Class Distance Calculations caret includes several functions to pre-process the predictor data. It assumes that […]

###
Intro to caret: Visualizations

Editor’s note: This is the second of a series of posts on the caret package. The featurePlot function is a wrapper for different lattice plots to visualize the data. For example, the following figures show the default plot for continuous outcomes generated using the featurePlotfunction. For classification data sets, the iris data are used for illustration. […]

###
The caret Package

Editor’s note: This is the first of a long series of posts on the caret package. Introduction The caret package (short for _C_lassification _A_nd _RE_gression _T_raining) is a set of functions that attempt to streamline the process for creating predictive models. The package contains tools for: data splitting pre-processing feature selection model tuning using resampling […]

###
Optimizing with Nonlinear Programming

Rafael Ladeira asked on github: I was wondering why it doesn’t implement some others algorithms for search for optimal tuning parameters. What would be the caveats of using a genetic algorithm , for instance, instead of grid or random search? Do you think using some of those powerful optimization algorithms for tuning parameters is a […]

###
Three Aspects of Predictive Modeling

These slides were originally posted on appliedpredictivemodeling.com, and were kindly contributed to Open Data Science. Link to presentation: Three Aspects of Predictive Modeling By: Max Kuhn, Ph.D Presentation Overview: “Predictive modeling” definition Some example applications A short overview and example How is this dierent from what statisticians already do? Unmet challenges in applied modeling Predictive […]