Advanced22 min

The standard library toolkit

collections, itertools, functools, hashlib, uuid, os and sys — the batteries that ship with Python.

Why this matters in AI / ML / GenAI

Counting label frequencies, grouping records, caching embeddings, fingerprinting datasets, and generating request ids are daily tasks. Each has a one-line standard-library answer that most people reimplement badly.

collections

  • Counter — frequency counting with .most_common(n). Perfect for label distributions, token frequencies, and spotting class imbalance.
  • defaultdict — a dict that creates missing values automatically, so grouping needs no if key not in d dance.
  • namedtuple — a lightweight immutable record with named fields.
  • deque — a double-ended queue with O(1) appends and pops at both ends, and a maxlen that discards old items. Ideal for rolling windows and sliding-window metrics.
  • OrderedDict — rarely needed now that regular dicts preserve insertion order.

functools and itertools

functools: lru_cache / cache for memoizing pure functions, partial to pre-fill arguments, wraps for decorators, reduce for custom folds, and cached_property for expensive attributes computed once per instance.

itertools: chain, islice, groupby (requires sorted input — the number one gotcha), product for hyperparameter grids, combinations and permutations, cycle, count, and accumulate for running totals.

Both modules are implemented in C, so they are faster than hand-written equivalents as well as shorter.

Identity, hashing, and the environment

uuiduuid.uuid4() for request ids and run ids. Unique without coordination between machines.

hashlibsha256 to fingerprint a dataset, cache key, or prompt. A stable fingerprint lets you prove two runs used identical data. Never use md5 for security, and never hash passwords with a plain digest.

os / sysos.getenv for configuration, os.cpu_count() for worker sizing, sys.argv for CLI arguments (though argparse is better), sys.version_info for version checks, and sys.exit(1) to signal failure to CI.

pathlib, json, csv, tempfile, shutil cover the rest of everyday file work.

Copy-paste examples

Copy into your own editor, or load one into the compiler below and press Run.

Counter for label distribution

Class imbalance shows up immediately.

from collections import Counter

labels = ["positive"] * 70 + ["negative"] * 25 + ["neutral"] * 5
counts = Counter(labels)

print(counts)
print("most common:", counts.most_common(2))
total = sum(counts.values())
for label, n in counts.most_common():
    bar = "#" * (n // 2)
    print(f"{label:9} {n:3} ({n / total:5.1%}) {bar}")

print("\nimbalance ratio:", round(max(counts.values()) / min(counts.values()), 1), ": 1")
print("token frequency:", Counter("the cat sat on the mat".split()).most_common(2))

defaultdict and deque

Grouping without key checks, and a fixed-size rolling window.

from collections import defaultdict, deque

records = [
    ("mlops", "doc-1"), ("rag", "doc-2"),
    ("mlops", "doc-3"), ("agents", "doc-4"), ("rag", "doc-5"),
]

grouped = defaultdict(list)
for topic, doc in records:
    grouped[topic].append(doc)
for topic, docs in grouped.items():
    print(f"{topic:8} -> {docs}")

window = deque(maxlen=5)
print("\nrolling mean of last 5 latencies:")
for latency in [100, 120, 900, 130, 110, 105, 98]:
    window.append(latency)
    print(f"  add {latency:3} -> window={list(window)} mean={sum(window) / len(window):.1f}")

namedtuple and cached results

lru_cache turns a repeated computation into a lookup.

import functools
from collections import namedtuple

Prediction = namedtuple("Prediction", ["doc_id", "label", "score"])

preds = [
    Prediction("d1", "positive", 0.91),
    Prediction("d2", "negative", 0.44),
]
for p in preds:
    print(f"{p.doc_id}: {p.label} ({p.score:.2f})")
print("as dict:", preds[0]._asdict())

@functools.lru_cache(maxsize=256)
def expensive_embed(text):
    print("  computing:", text)
    return sum(ord(c) for c in text) % 9973

print("\nfirst call :", expensive_embed("python"))
print("second call:", expensive_embed("python"), "(cached)")
print("cache stats:", expensive_embed.cache_info())

itertools groupby and accumulate

groupby needs sorted input — that is the classic mistake.

import itertools

rows = [
    {"model": "mini", "tokens": 500},
    {"model": "large", "tokens": 1200},
    {"model": "mini", "tokens": 300},
    {"model": "large", "tokens": 900},
]

rows.sort(key=lambda r: r["model"])          # required before groupby
for model, group in itertools.groupby(rows, key=lambda r: r["model"]):
    items = list(group)
    print(f"{model:6} calls={len(items)} tokens={sum(i['tokens'] for i in items)}")

daily = [120, 340, 88, 502]
print("\nrunning total:", list(itertools.accumulate(daily)))
print("grid combos  :", len(list(itertools.product([1e-5, 2e-5], [16, 32], [3, 5]))))

uuid, hashlib, and environment

Request ids and dataset fingerprints in three lines.

import hashlib
import json
import os
import sys
import uuid

request_id = str(uuid.uuid4())
print("request id:", request_id, "| short:", request_id[:8])

dataset = [{"text": "sample a", "label": 1}, {"text": "sample b", "label": 0}]
fingerprint = hashlib.sha256(json.dumps(dataset, sort_keys=True).encode()).hexdigest()
print("dataset fingerprint:", fingerprint[:16])

prompt = "Summarise the quarterly report"
cache_key = hashlib.sha256(f"gpt-4.1-mini|0.0|{prompt}".encode()).hexdigest()[:20]
print("cache key:", cache_key)

os.environ.setdefault("LOG_LEVEL", "INFO")
print("\nlog level :", os.getenv("LOG_LEVEL"))
print("python    :", ".".join(map(str, sys.version_info[:3])))
print("cpu count :", os.cpu_count())

Analyse a batch of model calls

Try it — in-browser Python

Add more calls and watch the Counter and rolling window update.

Output

Python runs in your browser. First run downloads the runtime.

Press Run (or Ctrl+Enter) to execute.

CPython in WebAssembly. Stdlib works. NumPy, pandas, scikit-learn and Matplotlib load on demand, and charts render below. No input(), no GPU, no network installs.

Takeaways

  • Counter for distributions, defaultdict for grouping, deque for rolling windows.
  • lru_cache memoizes pure functions; itertools.groupby needs sorted input.
  • uuid4 for request ids, sha256 for dataset and cache fingerprints.