scikit-learn: regression
Train/test split, linear and polynomial regression, and the metrics that tell you whether a model is any good.
Why this matters in AI / ML / GenAI
scikit-learn is the fastest path from a table of numbers to a working model, and its fit/predict interface is the mental model behind every other framework. Regression also teaches overfitting and evaluation in a form you can see in one chart.
The estimator interface
Every scikit-learn model follows the same three steps:
model = SomeEstimator(**hyperparameters)model.fit(X_train, y_train)predictions = model.predict(X_test)
X is 2-dimensional — rows are samples, columns are features — even with a single feature, where you need .reshape(-1, 1). y is 1-dimensional. Getting this wrong produces the most common scikit-learn error message you will ever see.
Because the interface is uniform, swapping LinearRegression for RandomForestRegressor is a one-line change.
Train/test split
Never evaluate on data the model trained on — it has memorised it, and the score is meaningless.
train_test_split(X, y, test_size=0.2, random_state=42) holds back 20%. Always pass random_state so the split is reproducible.
The rule that gets broken most often: fit scalers and encoders on the training set only, then apply them to test data. Fitting on everything leaks information about the test set into training and inflates your score. A Pipeline enforces this automatically.
For time series, never split randomly. Future data must not leak into past training data — split by date.
Regression metrics
- MAE (mean absolute error) — average error in original units. Easy to explain to a stakeholder.
- MSE — squares the errors, so large mistakes dominate. It is what most models optimise.
- RMSE — square root of MSE, back in original units.
- R² — the share of variance explained. 1.0 is perfect, 0 is no better than predicting the mean, and negative is worse than the mean.
Report MAE alongside R². R² of 0.85 sounds strong until you learn the MAE is ₹40,000 on a ₹50,000 prediction.
Polynomial regression fits curves by adding x², x³ as features. It is also the clearest demonstration of overfitting: raise the degree far enough and the curve passes through every training point while predicting nonsense between them.
Copy-paste examples
Copy into your own editor, or load one into the compiler below and press Run.
Linear regression end to end
First run downloads scikit-learn, which takes a few seconds.
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
rng = np.random.default_rng(42)
X = rng.uniform(0, 10, 200).reshape(-1, 1) # 2D: (200 rows, 1 feature)
y = 3.5 * X.ravel() + 12 + rng.normal(0, 2, 200)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
print("train:", X_train.shape, "test:", X_test.shape)
model = LinearRegression().fit(X_train, y_train)
pred = model.predict(X_test)
print(f"\nlearned: y = {model.coef_[0]:.3f}x + {model.intercept_:.3f}")
print("true : y = 3.500x + 12.000")
print(f"\nMAE : {mean_absolute_error(y_test, pred):.3f}")
print(f"RMSE: {mean_squared_error(y_test, pred) ** 0.5:.3f}")
print(f"R2 : {r2_score(y_test, pred):.4f}")Multiple regression and feature importance
Coefficients are only comparable when features are on the same scale.
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.metrics import r2_score
rng = np.random.default_rng(7)
n = 400
experience = rng.uniform(0, 15, n)
projects = rng.integers(0, 30, n).astype(float)
noise_feature = rng.normal(0, 1, n)
salary = 300000 + 85000 * experience + 12000 * projects + rng.normal(0, 40000, n)
X = np.column_stack([experience, projects, noise_feature])
y = salary
names = ["experience", "projects", "irrelevant"]
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.25, random_state=0)
pipeline = make_pipeline(StandardScaler(), LinearRegression()).fit(X_train, y_train)
coefs = pipeline.named_steps["linearregression"].coef_
print("R2 on test:", round(r2_score(y_test, pipeline.predict(X_test)), 4))
print("\nstandardised coefficients (impact per 1 sd):")
for name, coef in sorted(zip(names, coefs), key=lambda p: -abs(p[1])):
print(f" {name:12} {coef:>12,.0f}")Polynomial regression and overfitting
Watch the degree-15 model score perfectly on train and badly on test.
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import r2_score
rng = np.random.default_rng(3)
X = np.sort(rng.uniform(-3, 3, 60)).reshape(-1, 1)
y = 0.5 * X.ravel() ** 3 - 2 * X.ravel() + rng.normal(0, 2.5, 60)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=1)
print(f"{'degree':>7}{'train R2':>11}{'test R2':>11} verdict")
for degree in [1, 3, 8, 15]:
model = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(X_train, y_train)
train_r2 = r2_score(y_train, model.predict(X_train))
test_r2 = r2_score(y_test, model.predict(X_test))
if test_r2 < 0.4 and train_r2 > 0.9:
verdict = "overfitting"
elif train_r2 < 0.6:
verdict = "underfitting"
else:
verdict = "good fit"
print(f"{degree:>7}{train_r2:>11.4f}{test_r2:>11.4f} {verdict}")Plot the fitted curves
The chart makes overfitting unmistakable — the wiggly line is memorising noise.
import matplotlib.pyplot as plt
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline
rng = np.random.default_rng(3)
X = np.sort(rng.uniform(-3, 3, 40)).reshape(-1, 1)
y = 0.5 * X.ravel() ** 3 - 2 * X.ravel() + rng.normal(0, 2.5, 40)
grid = np.linspace(-3, 3, 300).reshape(-1, 1)
fig, ax = plt.subplots(figsize=(7, 4))
ax.scatter(X, y, color="#0f172a", alpha=0.7, label="training data", zorder=3)
for degree, colour in [(1, "#64748b"), (3, "#1d4ed8"), (15, "#dc2626")]:
model = make_pipeline(PolynomialFeatures(degree), LinearRegression()).fit(X, y)
ax.plot(grid, model.predict(grid), color=colour, linewidth=2, label=f"degree {degree}")
ax.set_ylim(y.min() - 5, y.max() + 5)
ax.set_title("Polynomial fits: underfit, good fit, overfit")
ax.set_xlabel("Feature x")
ax.set_ylabel("Target y")
ax.legend()
print("degree 1 is too rigid, degree 3 matches the truth, degree 15 chases noise")Fit and evaluate your own regression
Try it — in-browser Python
Packages: numpy, scikit-learn, matplotlib
Change noise_level to 30 and watch R² fall while the coefficients stay close.
Output
Python runs in your browser. First run downloads the runtime.
Press Run (or Ctrl+Enter) to execute.
CPython in WebAssembly. Stdlib works. NumPy, pandas, scikit-learn and Matplotlib load on demand, and charts render below. No input(), no GPU, no network installs.
Takeaways
- Every estimator is fit(X, y) then predict(X); X must be 2D and y 1D.
- Always split before fitting, pass random_state, and fit scalers on training data only.
- Report MAE alongside R²; rising train score with falling test score means overfitting.