Pandas
Pandas is an essential data analysis and manipulation library for Python. It provides fast, flexible, and expressive data structures designed to make working with "relational" or "labeled" data both easy and intuitive.
Key Data Structures
- Series: 1D labeled homogeneously-typed array.
- DataFrame: General 2D labeled, size-mutable tabular structure with potentially heterogeneously-typed columns.
Basic Usage
import pandas as pd
# Create a DataFrame from a dictionary
data = {
"Name": ["Alice", "Bob", "Charlie"],
"Age": [25, 30, 35],
"City": ["New York", "Paris", "London"]
}
df = pd.DataFrame(data)
# Reading data from CSV
# df = pd.read_csv("data.csv")
# Data inspection
print(df.head())
print(df.describe())
# Filtering data
adults = df[df["Age"] >= 30]
# Grouping and aggregation
# df.groupby("City")["Age"].mean()
Why it is essential for AI
Before training any ML model, you must clean, explore, and preprocess your dataset. Pandas is the industry standard for this EDA (Exploratory Data Analysis) and feature engineering phase.