How To Detect Targets on Image with Deep Learning and OpenCV

Object detection projects sometimes need to find groups of people in a single frame for further processing. There are many ways to do this; this article describes one that combines deep learning with OpenCV.
First, define what counts as a group. We treated a group as three or more people standing close enough together that their bounding boxes overlap.
Tools
We used Python, because most machine learning libraries are available for it. The video is processed frame by frame: the frames come from the OpenCV library, and each frame is treated as a separate image, as in the OpenCV tutorial.
To detect people in each frame, we needed a detector that is light and fast, and chose YOLO. It runs reasonably fast even on a CPU and has many implementations and pre-trained models.
How YOLO Works
According to the YOLO authors, earlier detection systems repurpose classifiers or localisers: they run the model on an image at many locations and scales and treat high-scoring regions as detections.
YOLO instead applies a single neural network to the whole image. The network divides the image into regions and predicts bounding boxes and probabilities for each region, and the boxes are weighted by those probabilities.
Because YOLO sees the whole image at once, its predictions use the global context of the image. It also needs only one network evaluation per image, while R-CNN needs thousands; the authors report it as more than 1000x faster than R-CNN and 100x faster than Fast R-CNN. YOLO is implemented in Darknet, an open-source neural network framework. This project used YOLOv3 (2018); newer detectors have appeared since.
Deep Learning Backend
The model loads easily through GluonCV, a deep learning toolkit for computer vision that uses the MXNet library as its backend. Apache MXNet was retired in September 2023 and moved to the Apache Attic in February 2024, so choose a maintained framework for a new project.
To pass a video frame to MXNet, convert its NumPy array into an MXNet array.
right_frame = mxnet.nd.array(frame)Processing and finding groups
We chose the yolo3_darknet53_coco implementation of YOLO (YOLOv3 with a Darknet-53 backbone), pre-trained on the COCO dataset. COCO already includes a “person” class, so we didn't need to train the network.
net = model_zoo.get_model(
'yolo3_darknet53_coco',
pretrained=True
)
x, img = data.transforms.presets.yolo.transform_test(
right_frame,
short=512,
max_size=1024
)
class_IDs, scores, bounding_boxes = net(x)Processing returns three arrays: class IDs (class_IDs), class probabilities (scores) and bounding boxes (bounding_boxes). With the bounding boxes, you can find groups through intersecting boxes, using the coordinates of each rectangle's corners.
Then, if the bounding boxes of two people intersect, we add the pair to an array (target_intersection). Looping over the pairs, we check whether the same person appears in two pairs; if so, we create a group from them or update an existing one.
target_intersection = [
sorted((k1,k2))
for k1,v1 in targets.items()
for k2, v2 in targets.items()
if k1 != k2 and bounding_box_intersection(v1, v2)
]
groups = []
for pair in target_intersection:
found_group = False
for group in groups:
if set(pair).intersection(group):
group.update(pair)
found_group = True
break
if not found_group:
groups.append(set(pair))Result
In the picture, each detected person is marked with a bounding box, the class name and the class probability. The program returns:
groups = find_groups(persons)
print(f"Result = {groups}")
Result = [4,3]Result = [4, 3] means two groups of people: four in the first and three in the second. An empty result, [], means no groups were detected in the image.
Conclusion
The approach is simple: a pre-trained detector finds people, and overlapping bounding boxes link them into groups. Accuracy depends on the detector and on how close people stand, since people near each other whose boxes don't overlap are not grouped. The same pipeline can count groups in CCTV footage or flag crowded moments for video editing.
Share and subscribe to our blog
How can we help you ?







