{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# Lesson 41: Object Detection as Regression\n", "\n", "Lesson 40's sliding-window detector classified thousands of fixed-size windows and merged the results with non-maximum suppression (NMS). That works, but it is fundamentally a *classification* approach bolted onto a search — the network never predicts a box directly, only \"face or not, at this exact window.\" Modern detectors instead treat localization as a **regression** problem: given an image, directly predict bounding box coordinates. This lesson builds the simplest possible version of that idea, then surveys how real detectors (R-CNN, YOLO, SSD) scale it up." ] }, { "cell_type": "code", "execution_count": null, "id": "a6d0f365", "metadata": {}, "outputs": [], "source": [ "import numpy as np\n", "import torch\n", "import torch.nn as nn\n", "import torch.nn.functional as F\n", "import matplotlib.pyplot as plt\n", "import matplotlib.patches as patches" ] }, { "cell_type": "markdown", "id": "dbe19032", "metadata": {}, "source": [ "## One object per image: predict a box directly\n", "\n", "The simplest possible regression problem: one circular blob per image, at an unknown location and size. Instead of a class label, the network's target is now four numbers — `(cx, cy, width, height)` indicating the bounding box, normalized to `[0, 1]` by image size." ] }, { "cell_type": "code", "execution_count": null, "id": "f2b18fc7", "metadata": {}, "outputs": [], "source": [ "SIZE = 32\n", "\n", "def make_scene(rng, size=SIZE, obj_size=8):\n", " scene = np.zeros((size, size), dtype=np.float32)\n", " cx = rng.integers(obj_size, size - obj_size)\n", " cy = rng.integers(obj_size, size - obj_size)\n", " yy, xx = np.mgrid[0:size, 0:size]\n", " scene[((xx - cx) ** 2 + (yy - cy) ** 2) <= (obj_size * 0.5) ** 2] = 1.0\n", " scene = np.clip(scene + rng.normal(0, 0.05, scene.shape), 0, 1).astype(np.float32)\n", " box = (cx, cy, obj_size, obj_size) # cx, cy, w, h\n", " return scene, box\n", "\n", "rng = np.random.default_rng(9)\n", "N = 400\n", "scenes, boxes = [], []\n", "for _ in range(N):\n", " s, b = make_scene(rng)\n", " scenes.append(s); boxes.append(b)\n", "scenes = np.array(scenes, dtype=np.float32)\n", "boxes = np.array(boxes, dtype=np.float32)\n", "\n", "split = int(0.85 * N)\n", "Xtr, Btr = scenes[:split], boxes[:split] / SIZE\n", "Xte, Bte = scenes[split:], boxes[split:] / SIZE\n", "\n", "fig, axes = plt.subplots(1, 4, figsize=(9, 2.5))\n", "for ax, im, b in zip(axes, Xtr[:4], boxes[:4]):\n", " ax.imshow(im, cmap='gray')\n", " ax.add_patch(patches.Rectangle((b[0] - b[2] / 2, b[1] - b[3] / 2), b[2], b[3], edgecolor='lime', facecolor='none', linewidth=2))\n", " ax.axis('off')\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "f532b10a", "metadata": {}, "source": [ "## The model and the IoU metric\n", "\n", "The network is a CNN backbone (Lesson 34's pattern) followed by a 4-output regression head with a sigmoid, so every prediction lands in `[0, 1]` — a valid normalized box coordinate. It's trained with plain MSE loss against the true box, and evaluated with **IoU** (Lesson 40's intersection-over-union), the metric that actually matters for detection: how much the predicted and true boxes overlap, not how close the four numbers are in isolation." ] }, { "cell_type": "code", "execution_count": null, "id": "65976fc1", "metadata": {}, "outputs": [], "source": [ "class Detector(nn.Module):\n", " def __init__(self):\n", " super().__init__()\n", " self.conv = nn.Sequential(\n", " nn.Conv2d(1, 16, 5, padding=2), nn.ReLU(), nn.MaxPool2d(2),\n", " nn.Conv2d(16, 32, 5, padding=2), nn.ReLU(), nn.AdaptiveMaxPool2d(1),\n", " )\n", " self.fc = nn.Linear(32, 4) # cx, cy, w, h, normalized\n", "\n", " def forward(self, x):\n", " return torch.sigmoid(self.fc(self.conv(x).flatten(1)))\n", "\n", "def iou_batch(pred, target):\n", " pcx, pcy, pw, ph = pred[:, 0], pred[:, 1], pred[:, 2], pred[:, 3]\n", " tcx, tcy, tw, th = target[:, 0], target[:, 1], target[:, 2], target[:, 3]\n", " px0, py0, px1, py1 = pcx - pw / 2, pcy - ph / 2, pcx + pw / 2, pcy + ph / 2\n", " tx0, ty0, tx1, ty1 = tcx - tw / 2, tcy - th / 2, tcx + tw / 2, tcy + th / 2\n", " ix0, iy0 = torch.maximum(px0, tx0), torch.maximum(py0, ty0)\n", " ix1, iy1 = torch.minimum(px1, tx1), torch.minimum(py1, ty1)\n", " inter = (ix1 - ix0).clamp(min=0) * (iy1 - iy0).clamp(min=0)\n", " union = pw * ph + tw * th - inter\n", " return inter / union.clamp(min=1e-8)\n", "\n", "torch.manual_seed(0)\n", "model = Detector()\n", "opt = torch.optim.Adam(model.parameters(), lr=0.005)\n", "Xt = torch.tensor(Xtr).unsqueeze(1); Bt = torch.tensor(Btr)\n", "for _ in range(400):\n", " opt.zero_grad()\n", " loss = F.mse_loss(model(Xt), Bt)\n", " loss.backward()\n", " opt.step()\n", "\n", "with torch.no_grad():\n", " pred_te = model(torch.tensor(Xte).unsqueeze(1))\n", " ious = iou_batch(pred_te, torch.tensor(Bte))\n", "\n", "print(f'mean IoU on test set: {ious.mean().item():.3f}')\n", "print(f'fraction of test boxes with IoU > 0.5: {(ious > 0.5).float().mean().item():.1%}')" ] }, { "cell_type": "code", "execution_count": null, "id": "b84319ee", "metadata": {}, "outputs": [], "source": [ "fig, axes = plt.subplots(1, 4, figsize=(9, 2.5))\n", "for i, ax in enumerate(axes):\n", " ax.imshow(Xte[i], cmap='gray')\n", " tb = Bte[i] * SIZE\n", " pb = pred_te[i].numpy() * SIZE\n", " ax.add_patch(patches.Rectangle((tb[0] - tb[2] / 2, tb[1] - tb[3] / 2), tb[2], tb[3], edgecolor='lime', facecolor='none', linewidth=2, label='true'))\n", " ax.add_patch(patches.Rectangle((pb[0] - pb[2] / 2, pb[1] - pb[3] / 2), pb[2], pb[3], edgecolor='red', facecolor='none', linewidth=1.5, linestyle='--', label='pred'))\n", " ax.set_title(f'IoU={ious[i]:.2f}', fontsize=9)\n", " ax.axis('off')\n", "axes[0].legend(fontsize=6, loc='upper left')\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "fc21329f", "metadata": {}, "source": [ "## Scaling this up: two families of real detectors\n", "\n", "This detector above only handles exactly one object per image, because a fixed-size output vector (4 numbers) can only describe one box. But real scenes have a variable, unknown number of objects. Two different families of detectors dominate the landscape:\n", "\n", "**Two-stage (R-CNN family):** first generate a modest number of *region proposals* — candidate boxes likely to contain something, via a cheap, class-agnostic method (the original R-CNN used classical segmentation; **Faster R-CNN** (Ren et al., 2015★) learns a small \"region proposal network\" instead) — then run a classifier-plus-box-regressor (this lesson's whole architecture) on each proposal independently, exactly like running the sliding-window classifier from Lesson 40 but only at a handful of promising locations instead of every window. Accurate, but only as fast as (proposals) x (one forward pass) allows.\n", "\n", "**Single-stage (YOLO, SSD):** skip proposals entirely. **YOLO** (You Only Look Once, Redmon et al., 2016★) divides the image into a coarse grid of cells, and has each grid cell directly predict (as this lesson's network does) a fixed number of boxes plus a class label plus a confidence score, all in one forward pass. **SSD** (Single Shot MultiBox Detector, Liu et al., 2016★) follows the same one-pass recipe but predicts boxes from *several* feature-map resolutions at once (not just one final grid), so coarser layers naturally catch larger objects and finer layers catch smaller ones. To let a single cell describe objects of different aspect ratios, single-stage detectors use **anchor boxes**: several predefined box shapes (tall, wide, square) per cell, with the network predicting an *offset* from each anchor rather than a box from scratch. Faster to run, historically somewhat less accurate than two-stage methods, though the gap has narrowed considerably.\n", "\n", "Both families end with the same postprocessing step: non-maximum suppression (NMS, Lesson 40) to merge the overlapping candidate boxes any real multi-object scene produces. (For simplicity, this lesson omitted NMS by only predicting a single bounding box.)" ] }, { "cell_type": "markdown", "id": "fb52557c", "metadata": {}, "source": [ "## In practice: real detectors on real images\n", "\n", "OpenCV ships a single-stage detector, `cv2.FaceDetectorYN` (**YuNet**, Wu et al., 2023), that follows the single-stage recipe above for faces specifically. Unlike Lesson 40's Viola-Jones cascade, it's a small single-shot CNN — the same family as SSD above, just specialized to one class — and it runs its own NMS internally before returning boxes." ] }, { "cell_type": "code", "execution_count": null, "id": "2f425034", "metadata": {}, "outputs": [], "source": [ "import urllib.request\n", "from pathlib import Path\n", "import cv2\n", "\n", "CACHE_DIR = Path.home() / '.cache' / 'cvintro'\n", "YUNET_URL = 'https://media.githubusercontent.com/media/opencv/opencv_zoo/main/models/face_detection_yunet/face_detection_yunet_2023mar.onnx'\n", "YUNET_PATH = CACHE_DIR / 'face_detection_yunet_2023mar.onnx'\n", "\n", "def ensure_yunet():\n", " if YUNET_PATH.exists():\n", " return\n", " CACHE_DIR.mkdir(parents=True, exist_ok=True)\n", " print('Downloading the YuNet face detector (one-time, cached under ~/.cache/cvintro)...')\n", " urllib.request.urlretrieve(YUNET_URL, YUNET_PATH)\n", "\n", "ensure_yunet()\n", "\n", "photo = cv2.imread('../img/apollo11_crew.jpg') # same photo Lesson 40 ran Viola-Jones on\n", "h, w = photo.shape[:2]\n", "yunet = cv2.FaceDetectorYN_create(str(YUNET_PATH), '', (w, h))\n", "_, yunet_faces = yunet.detect(photo)\n", "\n", "photo_rgb = cv2.cvtColor(photo, cv2.COLOR_BGR2RGB)\n", "fig, ax = plt.subplots(figsize=(8, 6))\n", "ax.imshow(photo_rgb)\n", "for x, y, fw, fh, *_, score in yunet_faces:\n", " ax.add_patch(patches.Rectangle((x, y), fw, fh, edgecolor='lime', facecolor='none', linewidth=2))\n", " ax.text(x, y - 8, f'{score:.2f}', color='lime', fontsize=9, weight='bold')\n", "ax.set_title(f'cv2.FaceDetectorYN: {len(yunet_faces)} detections')\n", "ax.axis('off')\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "98b0dd9f", "metadata": {}, "source": [ "All three astronauts are found correctly, with no false positive — which is an improvement over Lesson 40's Viola-Jones cascade. This is the payoff of a learned single-shot CNN over a cascade of hand-designed Haar-like features: richer features, trained end-to-end on real face/non-face data rather than assembled stage by stage." ] }, { "cell_type": "markdown", "id": "3a9a4ad3", "metadata": {}, "source": [ "But YuNet only answers \"face or not\" — one class. \n", "\n", "For general, multi-class detection, Faster R-CNN pretrained on **COCO** (Lin et al., 2014★ — about 330k real photos labeled across 80 object categories) is a few lines away via `torchvision`. This is the same \"pretrained model in one line\" pattern as Lesson 38's ResNet-18. It natively handles a variable, unknown number of objects per image — the actual payoff of the region-proposal architecture described above." ] }, { "cell_type": "code", "execution_count": null, "id": "37f473ca", "metadata": {}, "outputs": [], "source": [ "import urllib.request\n", "from pathlib import Path\n", "import cv2\n", "import torchvision\n", "\n", "CACHE_DIR = Path.home() / '.cache' / 'cvintro'\n", "COCO_IMG_URL = 'http://images.cocodataset.org/val2017/000000039769.jpg'\n", "COCO_IMG_PATH = CACHE_DIR / 'coco_sample.jpg'\n", "\n", "def ensure_coco_sample():\n", " if COCO_IMG_PATH.exists():\n", " return\n", " CACHE_DIR.mkdir(parents=True, exist_ok=True)\n", " print('Downloading a COCO val2017 sample image (one-time, cached under ~/.cache/cvintro)...')\n", " urllib.request.urlretrieve(COCO_IMG_URL, COCO_IMG_PATH)\n", "\n", "ensure_coco_sample()\n", "\n", "weights = torchvision.models.detection.FasterRCNN_ResNet50_FPN_Weights.COCO_V1\n", "real_detector = torchvision.models.detection.fasterrcnn_resnet50_fpn(weights=weights)\n", "real_detector.eval() # frozen, no training at all -- Faster R-CNN exactly as released\n", "coco_classes = weights.meta['categories']\n", "\n", "img_bgr = cv2.imread(str(COCO_IMG_PATH))\n", "img_rgb = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2RGB)\n", "img_tensor = torch.tensor(img_rgb / 255.0, dtype=torch.float32).permute(2, 0, 1)\n", "\n", "with torch.no_grad():\n", " result = real_detector([img_tensor])[0]" ] }, { "cell_type": "code", "execution_count": null, "id": "2249bd53", "metadata": {}, "outputs": [], "source": [ "score_thresh = 0.7\n", "keep = result['scores'] > score_thresh\n", "\n", "fig, ax = plt.subplots(figsize=(7, 5.5))\n", "ax.imshow(img_rgb)\n", "for box, label, score in zip(result['boxes'][keep], result['labels'][keep], result['scores'][keep]):\n", " x0, y0, x1, y1 = box.numpy()\n", " ax.add_patch(patches.Rectangle((x0, y0), x1 - x0, y1 - y0, edgecolor='lime', facecolor='none', linewidth=2))\n", " ax.text(x0, y0 - 4, f'{coco_classes[label]} {score:.2f}', color='lime', fontsize=9, weight='bold')\n", "ax.set_title(f'Faster R-CNN pretrained on COCO ({int(keep.sum())} detections above score {score_thresh})')\n", "ax.axis('off')\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "c0ffcbea", "metadata": {}, "source": [ "
Image source: COCO dataset (val2017, image 000000039769)
" ] }, { "cell_type": "markdown", "id": "e2c60d2b", "metadata": {}, "source": [ "Four confident, correct detections survive the `score_thresh=0.7` cutoff — two cats and two remote controls, each above 0.78 — despite this network using only pretrained weights. Lower `score_thresh` and rerun to see what the cutoff was hiding: several overlapping, lower-confidence boxes (a \"couch\" and a \"bed\" guess both covering most of the image, around score 0.54)." ] }, { "cell_type": "markdown", "id": "950ad0f8", "metadata": {}, "source": [ "### Exercises\n", "\n", "1. Change `obj_size` in `make_scene` from a fixed `8` to a random value (e.g. `rng.integers(4, 12)`) so objects vary in size, and retrain. Does mean IoU hold up, get worse, or barely change — and why would variable object scale be harder for a single fixed-size regression head than variable position?\n", "2. The loss function here is plain MSE on `(cx, cy, w, h)`, but the metric that matters is IoU. Replace the loss with `1 - iou_batch(pred, target).mean()` (directly optimizing IoU) and compare final mean test IoU to the MSE-trained version. Real detectors (e.g. Faster R-CNN, YOLO variants) do exactly this with generalized IoU losses — can you see why MSE loss and IoU metric might disagree on which of two similar predictions is \"better\"?\n", "3. This lesson's detector has no anchor boxes and only ever handles one object. Sketch (in words) how you would modify the architecture's *output* to handle up to 3 objects per image using a 3x3 grid where each grid cell predicts one box and a \"there's an object here\" confidence — the core idea behind YOLO's grid.\n", "4. Lower `score_thresh` on the real Faster R-CNN to `0.3` and rerun. Count how many boxes appear versus at `0.7`; which of the new ones are genuinely additional objects and which are duplicates/near-misses on the cats or remotes that NMS would normally merge away? Then try a different COCO val2017 image URL (swap the filename in `COCO_IMG_URL`, e.g. any `http://images.cocodataset.org/val2017/<12-digit-id>.jpg`) and see what classes the model finds." ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python", "version": "3.10.0" } }, "nbformat": 4, "nbformat_minor": 5 }