{ "cells": [ { "cell_type": "markdown", "metadata": {}, "source": [ "# Lesson 43: Instance Segmentation\n", "\n", "Semantic segmentation (Lesson 42) labels each pixel with a class, but it has no notion of *how many* instances of each class there are — two touching instances are just labeled as one connected blob of pixels. **Instance segmentation** answers the harder question: which pixels belong to *this specific* object, as opposed to some other object of the same class. This lesson augments Lesson 42's segmentation network directly to solve it, with an idea that traces straight back to Lesson 12's Hough transform: instead of only classifying pixels, have the network also predict, for every foreground pixel, a vote for where its object's center is — then cluster the votes to recover individual instances." ] }, { "cell_type": "code", "execution_count": null, "id": "004a784d", "metadata": {}, "outputs": [], "source": [ "import numpy as np\n", "import cv2\n", "import torch\n", "import torch.nn as nn\n", "import torch.nn.functional as F\n", "import matplotlib.pyplot as plt" ] }, { "cell_type": "markdown", "id": "557bc204", "metadata": {}, "source": [ "## Overlapping instances\n", "\n", "Two circles per scene, deliberately placed close enough to touch or overlap. Ground truth includes both a semantic mask (0=background, 1=circle) and an *instance* mask (0=background, 1=first circle, 2=second circle)." ] }, { "cell_type": "code", "execution_count": null, "id": "68daaf46", "metadata": {}, "outputs": [], "source": [ "SIZE = 32\n", "\n", "def make_scene(rng, size=SIZE, n_circles=2, r=5):\n", " scene = np.zeros((size, size), dtype=np.float32)\n", " sem_mask = np.zeros((size, size), dtype=np.int64)\n", " inst_mask = np.zeros((size, size), dtype=np.int64)\n", " centers = []\n", " for k in range(n_circles):\n", " if k == 0:\n", " cx, cy = rng.integers(r + 8, size - r - 8), rng.integers(r + 8, size - r - 8)\n", " else:\n", " px, py = centers[-1]\n", " angle = rng.uniform(0, 2 * np.pi)\n", " dist = rng.uniform(6, 9) # overlapping (< 2r) but centers still separable\n", " cx = int(np.clip(px + dist * np.cos(angle), r, size - r - 1))\n", " cy = int(np.clip(py + dist * np.sin(angle), r, size - r - 1))\n", " centers.append((cx, cy))\n", " yy, xx = np.mgrid[0:size, 0:size]\n", " m = ((xx - cx) ** 2 + (yy - cy) ** 2) <= r ** 2\n", " scene[m] = 1.0\n", " sem_mask[m] = 1\n", " inst_mask[m] = k + 1 # later circles paint over earlier ones at overlaps\n", " scene = np.clip(scene + rng.normal(0, 0.05, scene.shape), 0, 1).astype(np.float32)\n", " return scene, sem_mask, inst_mask, centers\n", "\n", "rng = np.random.default_rng(17)\n", "N = 300\n", "scenes, sem_masks, inst_masks, all_centers = [], [], [], []\n", "for _ in range(N):\n", " s, sm, im, c = make_scene(rng)\n", " scenes.append(s); sem_masks.append(sm); inst_masks.append(im); all_centers.append(c)\n", "scenes = np.array(scenes, dtype=np.float32)\n", "sem_masks = np.array(sem_masks, dtype=np.int64)\n", "\n", "fig, axes = plt.subplots(2, 4, figsize=(9, 4.5))\n", "for i in range(4):\n", " axes[0, i].imshow(scenes[i], cmap='gray'); axes[0, i].axis('off')\n", " axes[1, i].imshow(inst_masks[i], cmap='viridis'); axes[1, i].axis('off')\n", "axes[0, 0].set_title('input', fontsize=9, loc='left')\n", "axes[1, 0].set_title('ground truth instance mask', fontsize=9, loc='left')\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "7c31f82a", "metadata": {}, "source": [ "## Augmenting the network with an instance head\n", "\n", "Take Lesson 42's exact U-Net architecture and give it a second output head alongside the semantic head: for every foreground pixel, predict a `(dx, dy)` offset pointing toward its own instance's center. The semantic head still does the same job as Lesson 42 (background vs. circle); the new offset head is what turns plain segmentation into *instance* segmentation." ] }, { "cell_type": "code", "execution_count": null, "id": "d4c3a3a6", "metadata": {}, "outputs": [], "source": [ "# build offset targets: for every foreground pixel, the (dx, dy) to ITS instance's center\n", "yy, xx = np.mgrid[0:SIZE, 0:SIZE]\n", "offset_targets = np.zeros((N, 2, SIZE, SIZE), dtype=np.float32)\n", "for i in range(N):\n", " im = inst_masks[i]\n", " for k, (cx, cy) in enumerate(all_centers[i]):\n", " m = im == (k + 1)\n", " offset_targets[i, 0][m] = (cx - xx[m]) / SIZE\n", " offset_targets[i, 1][m] = (cy - yy[m]) / SIZE\n", "\n", "split = int(0.85 * N)\n", "Xtr, Str, Otr = scenes[:split], sem_masks[:split], offset_targets[:split]\n", "Xte, Ste, Ote = scenes[split:], sem_masks[split:], offset_targets[split:]\n", "inst_te, centers_te = inst_masks[split:], all_centers[split:]\n", "\n", "class InstanceNet(nn.Module):\n", " \"\"\"Lesson 42's U-Net, with a second head: per-pixel offset-to-center regression\"\"\"\n", " def __init__(self):\n", " super().__init__()\n", " self.enc1 = nn.Sequential(nn.Conv2d(1, 16, 3, padding=1), nn.ReLU())\n", " self.enc2 = nn.Sequential(nn.Conv2d(16, 32, 3, padding=1), nn.ReLU())\n", " self.pool = nn.MaxPool2d(2)\n", " self.up = nn.Upsample(scale_factor=2, mode='nearest')\n", " self.dec1 = nn.Sequential(nn.Conv2d(32 + 16, 16, 3, padding=1), nn.ReLU())\n", " self.sem_head = nn.Conv2d(16, 2, 1) # background vs. circle\n", " self.offset_head = nn.Conv2d(16, 2, 1) # (dx, dy) to this pixel's instance center\n", "\n", " def forward(self, x):\n", " f1 = self.enc1(x)\n", " f2 = self.enc2(self.pool(f1))\n", " d1 = self.dec1(torch.cat([self.up(f2), f1], dim=1))\n", " return self.sem_head(d1), self.offset_head(d1)\n", "\n", "torch.manual_seed(0)\n", "model = InstanceNet()\n", "opt = torch.optim.Adam(model.parameters(), lr=0.01)\n", "Xt = torch.tensor(Xtr).unsqueeze(1); St = torch.tensor(Str); Ot = torch.tensor(Otr)\n", "fg_mask_t = (St == 1).unsqueeze(1).float()\n", "\n", "for _ in range(300):\n", " opt.zero_grad()\n", " sem_logits, offset_pred = model(Xt)\n", " sem_loss = F.cross_entropy(sem_logits, St)\n", " offset_loss = (F.mse_loss(offset_pred, Ot, reduction='none') * fg_mask_t).sum() / fg_mask_t.sum().clamp(min=1)\n", " (sem_loss + 2.0 * offset_loss).backward()\n", " opt.step()\n", "\n", "with torch.no_grad():\n", " sem_logits_te, offset_pred_te = model(torch.tensor(Xte).unsqueeze(1))\n", " sem_preds = sem_logits_te.argmax(1).numpy()\n", " offset_preds = offset_pred_te.numpy()\n", "\n", "print(f'semantic pixel accuracy: {(sem_preds == Ste).mean():.1%}')\n" ] }, { "cell_type": "markdown", "id": "d11fb241", "metadata": {}, "source": [ "## Recovering instances: vote for the center, then cluster (Hough, again)\n", "\n", "The offset head above was trained to predict, for every foreground pixel, an offset pointing toward *its own instance's* center — supervised directly from the ground-truth instance masks. Each foreground pixel casts one vote (its predicted center location) into an accumulator, exactly like Lesson 12's Hough line/circle voting. Because every pixel belonging to the same circle votes for approximately the same point, the accumulator forms one tight cluster of votes per instance, however tangled the pixels themselves are. Finding instances becomes: find the vote clusters, then assign each pixel to its nearest cluster." ] }, { "cell_type": "code", "execution_count": null, "id": "bdf53382", "metadata": {}, "outputs": [], "source": [ "def cluster_instances(binary_mask, offset_pred, size=SIZE, peak_dist=4, vote_thresh=3):\n", " ys, xs = np.where(binary_mask)\n", " if len(xs) == 0:\n", " return np.zeros_like(binary_mask, dtype=np.int64), []\n", " pred_cx = xs + offset_pred[0][ys, xs] * size\n", " pred_cy = ys + offset_pred[1][ys, xs] * size\n", "\n", " votes = np.zeros((size, size), dtype=np.float32)\n", " for cx, cy in zip(pred_cx, pred_cy):\n", " cxi = int(np.clip(round(cx), 0, size - 1)); cyi = int(np.clip(round(cy), 0, size - 1))\n", " votes[cyi, cxi] += 1\n", "\n", " # greedily take the highest-voted peak, suppress its neighborhood, repeat\n", " # (the same non-max-suppression idea as Lesson 40's detection boxes)\n", " peaks = []\n", " votes_work = votes.copy()\n", " for _ in range(6):\n", " idx = np.unravel_index(votes_work.argmax(), votes_work.shape)\n", " if votes_work[idx] < vote_thresh:\n", " break\n", " peaks.append((idx[1], idx[0]))\n", " y0, y1 = max(0, idx[0] - peak_dist), min(size, idx[0] + peak_dist + 1)\n", " x0, x1 = max(0, idx[1] - peak_dist), min(size, idx[1] + peak_dist + 1)\n", " votes_work[y0:y1, x0:x1] = 0\n", "\n", " if not peaks:\n", " return np.zeros_like(binary_mask, dtype=np.int64), []\n", " inst_pred = np.zeros_like(binary_mask, dtype=np.int64)\n", " peaks_arr = np.array(peaks)\n", " for y, x in zip(ys, xs):\n", " dists = (peaks_arr[:, 0] - x) ** 2 + (peaks_arr[:, 1] - y) ** 2\n", " inst_pred[y, x] = np.argmin(dists) + 1\n", " return inst_pred, peaks\n", "\n", "n_correct = 0\n", "inst_preds, all_peaks = [], []\n", "for i in range(len(Xte)):\n", " binary_mask = sem_preds[i] == 1\n", " inst_pred, peaks = cluster_instances(binary_mask, offset_preds[i])\n", " inst_preds.append(inst_pred); all_peaks.append(peaks)\n", " true_n = len(set(inst_te[i].ravel()) - {0})\n", " if len(peaks) == true_n:\n", " n_correct += 1\n", "\n", "print(f'offset-voting recovers the correct instance count: {n_correct}/{len(Xte)} scenes '\n", " f'({n_correct/len(Xte):.1%})')" ] }, { "cell_type": "code", "execution_count": null, "id": "afe7a96d", "metadata": {}, "outputs": [], "source": [ "fig, axes = plt.subplots(3, 4, figsize=(9, 6.5))\n", "for i in range(4):\n", " axes[0, i].imshow(scenes[split + i], cmap='gray'); axes[0, i].axis('off')\n", " axes[1, i].imshow(inst_te[i], cmap='viridis'); axes[1, i].axis('off')\n", " axes[2, i].imshow(inst_preds[i], cmap='viridis'); axes[2, i].axis('off')\n", "for r, name in enumerate(['input', 'ground truth instances', 'recovered via offset voting']):\n", " axes[r, 0].set_title(name, fontsize=9, loc='left')\n", "plt.tight_layout()\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "0a46bfd2", "metadata": {}, "source": [ "## Mask R-CNN: Faster R-CNN plus a mask branch\n", "\n", "The offset-voting approach above is one family of instance segmentation methods (related to techniques sometimes called \"instance embedding\" or center-voting). The other dominant approach, **Mask R-CNN** (He et al., 2017★), does not start from scratch — it is **Faster R-CNN** (Lesson 41) with one new piece bolted on. The region proposal network, the per-box classifier, and the per-box coordinate refinement are all unchanged. Mask R-CNN just adds a third, parallel branch on every detected box: alongside the existing class label and refined box coordinates, a small fully-convolutional head predicts a binary mask *within that box*. Detection already solves \"how many objects, and roughly where\" via region proposals and NMS; the mask branch adds \"and here’s this one’s exact silhouette\" — exactly the piece built and tested in isolation below.\n", "\n", "One more change was needed to make that mask branch actually work well. Faster R-CNN extracts a fixed-size feature grid for each proposed box by snapping (quantizing) the box’s coordinates to the nearest feature-map cell — a step called RoIPool. That rounding barely affects classifying a box or nudging its coordinates by a few pixels, but it visibly misaligns a per-pixel mask against the object it is supposed to outline. Mask R-CNN replaces RoIPool with **RoIAlign**, which reads features via bilinear interpolation instead of rounding to the nearest cell, so the extracted features — and therefore the predicted mask — stay aligned with the actual box coordinates. This is a small change to the detector, but it is the specific fix that makes pixel-accurate masks possible.\n", "\n", "Putting Lessons 36, 40, 41, and this lesson together: **classification** (Lesson 36) gives one label for the whole image, **semantic segmentation** (Lesson 42) gives one label per pixel with no notion of separate objects of the same class, and **instance segmentation** (this lesson) adds a distinct identity per object — \"background,\" \"circle #1,\" \"circle #2,\" not just \"background,\" \"circle.\" A fourth term, **panoptic segmentation**, unifies the last two: every pixel gets both a semantic class and, for countable \"thing\" classes like circles or cars, an instance ID — while uncountable \"stuff\" classes like sky or road are labeled semantically only, since instance identity doesn’t make sense for them." ] }, { "cell_type": "markdown", "id": "cc1efa38", "metadata": {}, "source": [ "## A mask head\n", "\n", "Detection (Lesson 41) and per-instance masking are two separate, already-solved pieces at this point — the only genuinely new ingredient Mask R-CNN adds is the **mask head**: a small network that looks *only* inside one proposed box and predicts a binary mask for the single object that box was proposed for, ignoring anything else visible in the crop.\n", "\n", "Build that piece in isolation. Take the ground-truth box around each instance's true center (assume localization is already solved, so only the new mask-head idea is under test), crop a fixed 14x14 window there, and train a tiny fully-convolutional head to predict which pixels in that crop belong to *this* instance — not to any neighboring circle that happens to also be visible in the same crop." ] }, { "cell_type": "code", "execution_count": null, "id": "d2ea2c43", "metadata": {}, "outputs": [], "source": [ "CROP = 14\n", "HALF = CROP // 2\n", "\n", "def make_crops(indices):\n", " crops, masks = [], []\n", " for i in indices:\n", " scene, inst = scenes[i], inst_masks[i]\n", " for k, (cx, cy) in enumerate(all_centers[i]):\n", " x0 = int(np.clip(cx - HALF, 0, SIZE - CROP))\n", " y0 = int(np.clip(cy - HALF, 0, SIZE - CROP))\n", " crop = scene[y0:y0 + CROP, x0:x0 + CROP]\n", " inst_crop = inst[y0:y0 + CROP, x0:x0 + CROP]\n", " mask = (inst_crop == (k + 1)).astype(np.float32) # only THIS instance, never the neighbor\n", " crops.append(crop); masks.append(mask)\n", " return np.array(crops, dtype=np.float32), np.array(masks, dtype=np.float32)\n", "\n", "X_crop_train, M_crop_train = make_crops(range(split))\n", "X_crop_test, M_crop_test = make_crops(range(split, N))\n", "print(f'{len(X_crop_train)} training crops, {len(X_crop_test)} test crops, each {CROP}x{CROP}')\n", "\n", "class MaskHead(nn.Module):\n", " def __init__(self):\n", " super().__init__()\n", " self.net = nn.Sequential(\n", " nn.Conv2d(1, 16, 3, padding=1), nn.ReLU(),\n", " nn.Conv2d(16, 16, 3, padding=1), nn.ReLU(),\n", " nn.Conv2d(16, 1, 1), # per-pixel mask logit, full crop resolution throughout\n", " )\n", "\n", " def forward(self, x):\n", " return self.net(x)\n", "\n", "torch.manual_seed(0)\n", "mask_head = MaskHead()\n", "opt = torch.optim.Adam(mask_head.parameters(), lr=0.01)\n", "Xc = torch.tensor(X_crop_train).unsqueeze(1)\n", "Mc = torch.tensor(M_crop_train).unsqueeze(1)\n", "for _ in range(300):\n", " opt.zero_grad()\n", " loss = F.binary_cross_entropy_with_logits(mask_head(Xc), Mc)\n", " loss.backward()\n", " opt.step()\n", "\n", "with torch.no_grad():\n", " pred_masks = (mask_head(torch.tensor(X_crop_test).unsqueeze(1)) > 0).float().squeeze(1).numpy()\n", "\n", "def mean_iou(pred, true):\n", " inter = (pred * true).sum(axis=(1, 2))\n", " union = ((pred + true) > 0).sum(axis=(1, 2))\n", " return (inter / np.clip(union, 1, None)).mean()\n", "\n", "naive_masks = (X_crop_test > 0.5).astype(np.float32) # \"everything foreground in the crop\", no per-instance awareness\n", "print(f'mask-head IoU (per-instance): {mean_iou(pred_masks, M_crop_test):.3f}')\n", "print(f'naive foreground-in-crop baseline IoU: {mean_iou(naive_masks, M_crop_test):.3f}')" ] }, { "cell_type": "code", "execution_count": null, "id": "f3dcf333", "metadata": {}, "outputs": [], "source": [ "fig, axes = plt.subplots(3, 4, figsize=(9, 5))\n", "for i in range(4):\n", " axes[0, i].imshow(X_crop_test[i], cmap='gray'); axes[0, i].axis('off')\n", " axes[1, i].imshow(M_crop_test[i], cmap='viridis', vmin=0, vmax=1); axes[1, i].axis('off')\n", " axes[2, i].imshow(pred_masks[i], cmap='viridis', vmin=0, vmax=1); axes[2, i].axis('off')\n", "for r, name in enumerate(['crop (neighbor circle visible)', 'ground truth mask (this instance only)', 'predicted mask']):\n", " axes[r, 0].set_title(name, fontsize=9, loc='left')\n", "plt.tight_layout()\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "0651fd8a", "metadata": {}, "source": [ "Because this dataset's circles are generated close enough to always touch or overlap, *every single test crop* contains visible pixels from the neighboring circle — there is no \"easy\" case here to inflate the numbers. The per-instance mask head still reaches a respectably high IoU against the true single-instance mask, clearly ahead of the naive baseline that just marks every foreground pixel in the crop as belonging to \"the\" object (which is systematically wrong whenever a neighbor intrudes — the same problem that motivated giving the network an instance-center head in the first place, just replayed at crop scale). The mask head doesn't need to solve \"where are the circles\" — the box already answered that — it only has to learn \"given that this box was proposed for one specific circle, which pixels are that circle's, and not its neighbor's.\" That's a strictly easier, more local question, which is exactly why Mask R-CNN's mask branch can be small and fast: a good box already does most of the work.\n", "\n", "This also clarifies why Mask R-CNN and this lesson's offset-voting approach rarely fail the same way. Offset-voting can mis-cluster votes when two instances' centers are close enough that their accumulator peaks blur together (Exercise 1 below explores this). Mask R-CNN instead depends on the upstream detector proposing a correct box for every instance in the first place — miss a box, and that instance's mask is never even attempted, no matter how good the mask head is." ] }, { "cell_type": "markdown", "id": "f6b55760", "metadata": {}, "source": [ "## In practice: Mask R-CNN on a real photo\n", "\n", "The mask head above proves the idea works once localization is solved for free (ground-truth box centers). Here is the same idea, fully assembled and pretrained: a real **Mask R-CNN** (`torchvision.models.detection.maskrcnn_resnet50_fpn`, trained on COCO) run end to end — detection and mask head both included, nothing given away — on a photo with exactly the kind of touching, same-class instances this lesson has been building toward." ] }, { "cell_type": "code", "execution_count": null, "id": "1f7ac4f3", "metadata": {}, "outputs": [], "source": [ "import torchvision\n", "import cv2\n", "import matplotlib.patches as mpatches\n", "\n", "weights = torchvision.models.detection.MaskRCNN_ResNet50_FPN_Weights.COCO_V1\n", "maskrcnn = torchvision.models.detection.maskrcnn_resnet50_fpn(weights=weights)\n", "maskrcnn.eval() # frozen, no training at all\n", "coco_classes = weights.meta['categories']\n", "\n", "photo = cv2.imread('../img/apollo11_crew.jpg')\n", "photo_rgb = cv2.cvtColor(photo, cv2.COLOR_BGR2RGB)\n", "\n", "preprocess = weights.transforms()\n", "photo_tensor = torch.tensor(photo_rgb / 255.0, dtype=torch.float32).permute(2, 0, 1)\n", "batch = preprocess(photo_tensor).unsqueeze(0)\n", "\n", "with torch.no_grad():\n", " result = maskrcnn(batch)[0]\n", "\n", "score_thresh = 0.8\n", "keep = result['scores'] > score_thresh\n", "labels = result['labels'][keep].numpy()\n", "scores = result['scores'][keep].numpy()\n", "masks = result['masks'][keep, 0].numpy() # soft per-instance masks\n", "\n", "print(f'{keep.sum().item()} instances above {score_thresh:.0%} confidence:')\n", "for lbl, sc in zip(labels, scores):\n", " print(f' {coco_classes[lbl]:>10} {sc:.1%}')\n", "\n", "colors = plt.cm.tab10(np.linspace(0, 1, 10))\n", "inst_label_rgb = np.full_like(photo_rgb, 255)\n", "for i in range(len(labels)):\n", " color = (np.array(colors[i][:3]) * 255).astype(np.uint8)\n", " inst_label_rgb[masks[i] > 0.5] = color\n", "\n", "fig, axes = plt.subplots(1, 2, figsize=(12, 5.5))\n", "axes[0].imshow(photo_rgb); axes[0].set_title('input'); axes[0].axis('off')\n", "axes[1].imshow(inst_label_rgb); axes[1].set_title('Mask R-CNN per-instance masks'); axes[1].axis('off')\n", "legend_handles = [mpatches.Patch(color=colors[i], label=f'{coco_classes[labels[i]]} #{i+1} ({scores[i]:.0%})')\n", " for i in range(len(labels))]\n", "axes[1].legend(handles=legend_handles, loc='upper right', fontsize=8)\n", "plt.tight_layout()\n", "plt.show()" ] }, { "cell_type": "markdown", "id": "f29194a7", "metadata": {}, "source": [ "

Image source: NASA (public domain), via Wikimedia Commons

" ] }, { "cell_type": "markdown", "id": "a7896a04", "metadata": {}, "source": [ "The three astronauts are shoulder-to-shoulder, touching along multiple arm and torso boundaries, and belong to the exact same COCO class (`person`) — precisely the case where a semantic mask would merge them into one undifferentiated blob. Mask R-CNN separates them cleanly into three distinct instances anyway, because its detector proposes three separate boxes first (one per astronaut) and the mask head only looks inside a single box at a time, exactly like the mask head above. A fourth, spurious detection also shows up: a small `clock` mask at 84% confidence, on the left astronaut's wrist — the round, mirrored cuff-mounted checklist on Armstrong's arm, misread as a clock face." ] }, { "cell_type": "markdown", "id": "a791cfad", "metadata": {}, "source": [ "### Exercises\n", "\n", "1. Reduce `dist` in `make_scene` from the range `(6, 9)` to `(2, 5)`, making the circles overlap much more heavily (centers closer together). Does the offset-voting method's instance-count accuracy hold up, or does it start failing too — and if so, what does the failure look like (merged instances, split instances, or something else)?\n", "2. The clustering here uses `peak_dist=4` for non-max suppression on the vote accumulator. Try `peak_dist=8`. Does accuracy improve, get worse, or become sensitive to exactly how close two true circle centers happen to be in a given scene?\n", "3. This lesson only ever has 2 circles per scene. Extend `make_scene` to place a random number of circles (1 to 4) and adjust `cluster_instances`'s loop bound accordingly. Does the offset-voting approach's instance count accuracy hold up as the true number of instances grows?\n", "4. The mask head above was trained and evaluated using the *ground-truth* box center, sidestepping detection entirely. Perturb each crop's center by a few random pixels before cropping (simulating an imperfect detector) and rerun. How much does mask IoU degrade as the box gets less accurately centered — and does that match Mask R-CNN's real-world dependence on detector quality?" ] } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "name": "python", "version": "3.10.0" } }, "nbformat": 4, "nbformat_minor": 5 }