Skip to main content

Data and Features

Machine Learning runs entirely on data. Before a model can "learn," it must understand the language of data.

The Anatomy of a Dataset​

In Classical Machine Learning, we usually deal with structured, tabular data (like an Excel spreadsheet). Let's take a look at what makes up a dataset:

Temperature (X1)Weekend (X2)Discount % (X3)Ice Cream Sales (y)
85110$450
6000$150
95120$800

1. Features (XX)​

Features (also known as independent variables, inputs, or predictors) are the characteristics used to make predictions. In our example above, Temperature, Weekend, and Discount % are our features. They are denoted by a capital XX because a dataset usually consists of a matrix of features.

2. Labels (yy)​

The label (also known as the dependent variable, output, or target) is what you are trying to predict. In our example, Ice Cream Sales is the label. It is denoted by a lowercase yy because it's usually a single vector.

The Model's Job​

A Machine Learning model takes in the Features (XX) and tries to approximate a mathematical function ff that successfully predicts the Label (yy).

y=f(X)y = f(X)

Once the model has successfully "learned" this function ff using historical data, you can pass in brand new features (tomorrow's weather forecast for Bigkart), and it will spit out predicted sales!

Feature Engineering​

Sometimes raw data isn't good enough. Feature Engineering is the process of using domain knowledge to extract new, more useful features from raw data. For instance, instead of just using "Time of Sale" and "Sunset Time", a better feature might be "Hours after Sunset".