Advanced24 min

Pipelines, scaling, and clustering

Encode categories, scale features, chain it all in a Pipeline, then tune with cross-validation and grid search — plus K-means for unlabelled data.

Why this matters in AI / ML / GenAI

Pipelines are how professionals prevent data leakage and ship a single deployable object. Cross-validation gives an honest score, grid search tunes it, and K-means handles the very common case of having no labels at all.

Preprocessing

Scaling puts features on a comparable range. StandardScaler gives mean 0 and standard deviation 1; MinMaxScaler squeezes into [0, 1]. Required for KNN, SVM, K-means, and regularised linear models. Irrelevant for trees.

Categorical encoding: OneHotEncoder makes one binary column per category and is correct for unordered categories like country or model name. OrdinalEncoder assigns integers and is only appropriate when order genuinely exists (small < medium < large). Using ordinal encoding on unordered data tells the model that "Germany" is greater than "Brazil", which is meaningless.

Missing values: SimpleImputer fills with mean, median, or a constant. Median is safer for skewed data. Consider adding a boolean "was missing" column — the fact that a value was absent is often predictive.

Pipelines and ColumnTransformer

A Pipeline chains preprocessing and a model into one object that itself has fit and predict.

Two reasons this matters:

  1. No leakage. Inside cross-validation, the scaler is refit on each training fold rather than on the whole dataset. Scaling before splitting is the most common leak in beginner code, and it silently inflates scores.
  2. One deployable artifact. Pickle the pipeline and production applies exactly the same transforms as training. Mismatched preprocessing between training and serving is a top cause of models that work in a notebook and fail in production.

ColumnTransformer applies different steps to different columns — scale the numbers, one-hot the categories — in a single object.

Cross-validation, grid search, and K-means

A single train/test split is one sample of performance and can be lucky. K-fold cross-validation splits into k parts, trains k times, and reports mean and standard deviation. That standard deviation tells you how much to trust the mean.

GridSearchCV tries every hyperparameter combination with cross-validation and refits the best one. Use RandomizedSearchCV when the grid is large.

K-means groups unlabelled data into k clusters. You must choose k: the elbow method plots inertia against k and you take the bend. Always scale before clustering, since K-means measures raw distance.

Clusters are not labels. They are groupings you still have to interpret and name.

Copy-paste examples

Copy into your own editor, or load one into the compiler below and press Run.

A pipeline with mixed column types

ColumnTransformer scales numbers and one-hot encodes categories in one object.

import numpy as np
from sklearn.compose import ColumnTransformer
from sklearn.ensemble import RandomForestClassifier
from sklearn.impute import SimpleImputer
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
from sklearn.metrics import accuracy_score
import pandas as pd

rng = np.random.default_rng(42)
n = 500
df = pd.DataFrame({
    "tokens": rng.integers(50, 4000, n).astype(float),
    "latency_ms": rng.normal(800, 300, n),
    "model": rng.choice(["mini", "large", "local"], n),
    "region": rng.choice(["in", "us", "eu"], n),
})
df.loc[rng.choice(n, 30, replace=False), "latency_ms"] = np.nan
df["escalated"] = ((df["tokens"] > 2000) | (df["model"] == "local")).astype(int)

numeric = ["tokens", "latency_ms"]
categorical = ["model", "region"]

preprocess = ColumnTransformer([
    ("num", Pipeline([("impute", SimpleImputer(strategy="median")), ("scale", StandardScaler())]), numeric),
    ("cat", OneHotEncoder(handle_unknown="ignore"), categorical),
])

pipeline = Pipeline([
    ("prep", preprocess),
    ("model", RandomForestClassifier(n_estimators=120, random_state=0)),
])

X_train, X_test, y_train, y_test = train_test_split(
    df[numeric + categorical], df["escalated"], test_size=0.25, random_state=0, stratify=df["escalated"]
)
pipeline.fit(X_train, y_train)
print("missing values handled:", int(df["latency_ms"].isna().sum()))
print("test accuracy:", round(accuracy_score(y_test, pipeline.predict(X_test)), 4))
print("\none pipeline object holds imputation, scaling, encoding, and the model")

Why scaling before splitting leaks

The leaky version scores higher than it deserves. Cross-validation inside a pipeline is honest.

import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(0)
X = rng.normal(0, 1, (300, 20))
y = rng.integers(0, 2, 300)                  # labels are pure noise

leaky_X = StandardScaler().fit_transform(X)  # fit on ALL data, including test folds
leaky = cross_val_score(LogisticRegression(max_iter=500), leaky_X, y, cv=5).mean()

honest = cross_val_score(
    make_pipeline(StandardScaler(), LogisticRegression(max_iter=500)), X, y, cv=5
).mean()

print("labels are random, so the true score should be about 0.50")
print(f"scaled before splitting : {leaky:.4f}")
print(f"scaled inside pipeline  : {honest:.4f}")

Cross-validation and grid search

The standard deviation tells you how much to trust the mean.

import numpy as np
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import GridSearchCV, cross_val_score

rng = np.random.default_rng(7)
X = rng.normal(0, 1, (400, 5))
y = ((X[:, 0] + X[:, 1] ** 2) > 1).astype(int)

scores = cross_val_score(RandomForestClassifier(n_estimators=80, random_state=0), X, y, cv=5)
print("fold scores:", [round(float(s), 4) for s in scores])
print(f"mean {scores.mean():.4f} +/- {scores.std():.4f}")

grid = GridSearchCV(
    RandomForestClassifier(random_state=0),
    {"n_estimators": [50, 150], "max_depth": [3, 6, None]},
    cv=4,
    scoring="f1",
)
grid.fit(X, y)
print("\nbest params:", grid.best_params_)
print("best CV F1 :", round(grid.best_score_, 4))
print("combinations tried:", len(grid.cv_results_["params"]))

K-means with the elbow method

The bend in the curve suggests how many clusters the data really has.

import matplotlib.pyplot as plt
import numpy as np
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

rng = np.random.default_rng(42)
blobs = np.vstack([
    rng.normal([0, 0], 0.6, (120, 2)),
    rng.normal([4, 4], 0.6, (120, 2)),
    rng.normal([0, 5], 0.6, (120, 2)),
])
X = StandardScaler().fit_transform(blobs)

inertias = []
ks = range(1, 8)
for k in ks:
    inertias.append(KMeans(n_clusters=k, n_init=10, random_state=0).fit(X).inertia_)

model = KMeans(n_clusters=3, n_init=10, random_state=0).fit(X)

fig, axes = plt.subplots(1, 2, figsize=(9.5, 3.6))
axes[0].plot(list(ks), inertias, marker="o", color="#1d4ed8")
axes[0].axvline(3, color="#ea580c", linestyle="--", label="elbow at k=3")
axes[0].set_title("Elbow method")
axes[0].set_xlabel("Number of clusters (k)")
axes[0].set_ylabel("Inertia (within-cluster sum of squares)")
axes[0].legend()

axes[1].scatter(X[:, 0], X[:, 1], c=model.labels_, cmap="viridis", alpha=0.7)
axes[1].scatter(model.cluster_centers_[:, 0], model.cluster_centers_[:, 1],
                marker="X", s=200, color="#dc2626", label="centroids")
axes[1].set_title("K-means clusters (k=3)")
axes[1].set_xlabel("Feature 1 (standardised)")
axes[1].set_ylabel("Feature 2 (standardised)")
axes[1].legend()
fig.tight_layout()

unique, counts = np.unique(model.labels_, return_counts=True)
print("cluster sizes:", dict(zip(unique.tolist(), counts.tolist())))
print("inertia at k=3:", round(float(model.inertia_), 2))

Tune a pipeline with grid search

Try it — in-browser Python

Packages: numpy, scikit-learn, matplotlib

Add 0.01 to the C values list and see whether a stronger penalty wins.

Output

Python runs in your browser. First run downloads the runtime.

Press Run (or Ctrl+Enter) to execute.

CPython in WebAssembly. Stdlib works. NumPy, pandas, scikit-learn and Matplotlib load on demand, and charts render below. No input(), no GPU, no network installs.

Takeaways

  • Put every transform inside a Pipeline — it prevents leakage and ships as one artifact.
  • One-hot unordered categories; ordinal encoding implies an order that may not exist.
  • Cross-validate for an honest mean and spread, then GridSearchCV to tune.