1. Keypoint Evaluation

This page describes the keypoint evaluation metrics used by COCO. The evaluation code provided here can be used to obtain results on the publicly available COCO validation set. It computes multiple metrics described below. To obtain results on the COCO test set, for which ground-truth annotations are hidden, generated results must be uploaded to the evaluation server. The exact same evaluation code, described below, is used to evaluate results on the test set.

1.1. Evaluation Overview

The COCO keypoint task requires simultaneously detecting objects and localizing their keypoints (object locations are not given at test time). As the task of simultaneous detection and keypoint estimation is relatively new, we chose to adopt a novel metric inspired by object detection metrics. For simplicity, we refer to this task as keypoint detection and the prediction algorithm as the keypoint detector. We suggest reviewing the evaluation metrics for object detection before proceeding.

The core idea behind evaluating keypoint detection is to mimic the evaluation metrics used for object detection, namely average precision (AP) and average recall (AR) and their variants. At the heart of these metrics is a similarity measure between ground truth objects and predicted objects. In the case of object detection, the IoU serves as this similarity measure (for both boxes and segments). Thesholding the IoU defines matches between the ground truth and predicted objects and allows computing precision-recall curves. To adopt AP/AR for keypoints detection, we only need to define an analogous similarity measure. We do so by defining an object keypoint similarity (OKS) which plays the same role as the IoU.

1.2. Object Keypoint Similarity

For each object, ground truth keypoints have the form [x1,y1,v1,...,xk,yk,vk], where x,y are the keypoint locations and v is a visibility flag defined as v=0: not labeled, v=1: labeled but not visible, and v=2: labeled and visible. Each ground truth object also has a scale s which we define as the square root of the object segment area. For details on the ground truth format please see the download page.

For each object, the keypoint detector must output keypoint locations and an object-level confidence. Predicted keypoints for an object should have the same form as the ground truth: [x1,y1,v1,...,xk,yk,vk]. However, the detector's predicted vi are not currently used during evaluation, that is the keypoint detector is not required to predict per-keypoint visibilities or confidences.

We define the object keypoint similarity (OKS) as:

OKS = Σi[exp(-di2/2s2κi2)δ(vi>0)] / Σi[δ(vi>0)]

The di are the Euclidean distances between each corresponding ground truth and detected keypoint and the vi are the visibility flags of the ground truth (the detector's predicted vi are not used). To compute OKS, we pass the di through an unnormalized Guassian with standard deviation sκi, where s is the object scale and κi is a per-keypont constant that controls falloff. For each keypoint this yields a keypoint similarity that ranges between 0 and 1. These similarities are averaged over all labeled keypoints (keypoints for which vi>0). Predicted keypoints that are not labeled (vi=0) do not affect the OKS. Perfect predictions will have OKS=1 and predictions for which all keypoints are off by more than a few standard deviations sκi will have OKS~0. The OKS is analogous to the IoU. Given the OKS, we can compute AP and AR just as the IoU allows us to compute these metrics for box/segment detection.

1.3. Tuning OKS

We tune the κi such that the OKS is a perceptually meaningful and easy to interpret similarity measure. First, using 5000 redundantly annotated images in val, for each keypoint type i we measured the per-keypoint standard deviation σi with respect to object scale s. That is we compute σi2=E[di2/s2]. σi varies substantially for different keypoints: keypoints on a person's body (shoulders, knees, hips, etc.) tend to have a σ much larger than on a person's head (eyes, nose, ears).

To obtain a perceptually meaningful and interpretable similarity metric we set κi=2σi. With this setting of κi, at one, two, and three standard deviations of di/s the keypoint similarity exp(-di2/2s2κi2) takes on values of e-1/8=.88, e-4/8=.61 and e-9/8=.32. As expected, human annotated keypoints are normally distributed (ignoring occasional outliers). Thus, recalling the 68–95–99.7 rule, setting κi=2σi means that 68%, 95%, and 99.7% of human annotated keypoints should have a keypoint similarity of .88, .61, or .32 or higher, respectively (in practice the percentages are 75%, 95% and 98.7%).

The OKS is the average keypoint similarity across all (labeled) object keypoints. Below we plot the predicted OKS distribution with κi=2σi assuming 10 independent keypoints per object (blue curve) and the actual distribution of human OKS scores on the dually annotated data (green curve):

The curves don't match exactly for a few reasons: (1) object keypoints are not independent, (2) the number of labeled keypoints per objects varies, and (3) the real data contains 1-2% outliers (most of which are caused by annotators mistaking left for right or annotating the wrong person when two people are nearby). Nevertheless, the behavior is roughly as expected. We conclude with a few observations about human performance: (1) at OKS of .50, human performance is nearly perfect (95%), (2) median human OKS is ~.91, (3) human performance drops rapidly after an OKS of .95. Note that this OKS distribution can be used to predict human AR (as AR doesn't depend on false positives).

2. Metrics

The following 10 metrics are used for characterizing the performance of a keypoint detector on COCO:

Average Precision (AP):
AP
% AP at OKS=.50:.05:.95 (primary challenge metric)
APOKS=.50
% AP at OKS=.50 (loose metric)
APOKS=.75
% AP at OKS=.75 (strict metric)
AP Across Scales:
APmedium
% AP for medium objects: 322 < area < 962
APlarge
% AP for large objects: area > 962
Average Recall (AR):
AR
% AR at OKS=.50:.05:.95
AROKS=.50
% AR at OKS=.50
AROKS=.75
% AR at OKS=.75
AR Across Scales:
ARmedium
% AR for medium objects: 322 < area < 962
ARlarge
% AR for large objects: area > 962

  1. Unless otherwise specified, AP and AR are averaged over multiple OKS values (.50:.05:.95).
  2. As discussed, we set κi=2σi for each keypoint type i. For people, the σ's are .026, .025, .035, .079, .072, .062, .107, .087, & .089 for the nose, eyes, ears, shoulders, elbows, wrists, hips, knees, & ankles, respectively.
  3. AP (averaged across all 10 OKS thresholds) will determine the challenge winner. This should be considered the single most important metric when considering keypoint performance on COCO.
  4. All metrics are computed allowing for at most 20 top-scoring detections per image (we use 20 detections, not 100 as in the object detection challenge, as currently person is the only category with keypoints).
  5. Small objects (segment area < 322) do not contain keypoint annotations.
  6. For objects without labeled keypoints, including crowds, we use a lenient heuristic that allows matching of detections based on hallucinated keypoints (placed within the ground truth objects so as to maximize OKS). This is very similar to how ignore regions are handled for detection with boxes/segments. See the code for details.
  7. Each object is given equal importance, regardless of the number of labeled/visible keypoints. We do not filter objects with only a few keypoints, nor do we weight object examples by the number of keypoints present.

3. Evaluation Code

Evaluation code is available on the COCO github. Specifically, see either CocoEval.m or cocoeval.py in the Matlab or Python code, respectively. Also see evalDemo in either the Matlab or Python code (demo). Before running the evaluation code, please prepare your results in the format described on the results format page.

4. Analysis Code

In addition to the evaluation code, we also provide a function analyze() for performing a detailed breakdown of the errors in multi-instance keypoint estimation. This is described extensively in the paper Benchmarking and Error Diagnosis in Multi-Instance Pose Estimation by Ronchi et al. The code generates plots like this:

We show the results of the analysis of the Pose Affinity Fields detector from Zhe Cao et al., winner of the 2016 Keypoint Challenge at ECCV 2016.

The plot summarizes the impact of all types of error on the performance of a multi-instance pose estimation algorithm. It is composed of a series of Precision Recall (PR) curves where each curve is guaranteed to be strictly higher than the previous as the algorithm's detections are progressively corrected at an (arbitrary) OKS threshold of .9. The legend shows the Area Under the Curve (AUC). The curves are as follows (check project page for a full description):

  1. Original Dts.: PR obtained with the original detections at OKS=.9 (AP at strict OKS), area under curve corresponds to APOKS=.9 metric.
  2. Miss: PR at OKS=.9 (AP at strict OKS), after all miss errors have been corrected. A miss is a large localization error: the detected keypoint is not within the proximity of the correct body part.
  3. Swap: PR at OKS=.9 (AP at stric OKS), after all swap errors have been corrected. A swap is due to the confusion between the same body part of different people in an image (i.e. right elbow).
  4. Inversion: PR at OKS=.9 (AP at stric OKS), after all inversion errors have been corrected. An inversion is due to the confusion of body parts within the same person (i.e. left and right elbow).
  5. Jitter: PR at OKS=.9 (AP at strict OKS), after all jitter errors have been corrected. A jitter is a small localization error: the detected keypoint is within the proximity of the correct body part.
  6. Opt. Score: PR at OKS=.9 (AP at strict OKS), after all the algorithm's detections have been rescored using an oracle function computed at evaluation time. As a result of the rescoring the number of matches between detections and ground-truth instances is maximized.
  7. FP: PR after all background fps are removed. FP is a step function that is 1 until max recall is reached then drops to 0 (the curve is smoother after averaging across categories).
  8. FN: PR after all remaining errors are removed (trivially AP=1).

In the case of the above detector, overall AP at OKS=.9 is .327. Correcting all the miss errors results in a large improvement of the AP to .415. Smaller gains are obtained when correcting swaps, .448, and inversions, .545. Another large improvement is obtained when jiitter errors are removed, resulting in an AUC of .859. This shows what would the performance be if the CMU algorithm had perfect localization of keypoints. When localization is very good, the impact of confidence score errors is not as significant, but still results in an AUC improvement of about 2% (.879). Optimally scoring detections greatly diminishes the impact of Background False Positives, as detections rarely remain unmatched. Finally, removing Background False Negatives provides the remaining AUC to obtain perfect performance. In summary, CMU’s errors at OKS=.9 are dominated by imperfect localization, mostly jitter errors, and missed detections.

For a given detector, the code generates a total of 180 plots, analyzing all the types of errors at 3 area ranges (medium, large, all) and 10 evaluation thresholds (.5::.05::.95). The analysis code will automatically generate a pdf report containing a summary of the overall performance, the sensitivity of a method's behaviour to the different types of errors and their impact on performance, and several examples of the most significant failure cases.

Note: analyze() can take significant time to run, please be patient. As such, we typically do not run this code on the evaluation server; you must run the code locally using the validation set. You can find the analyze() function as part of this GitHub repository.