OpenCV Filtering, Thresholding, and Image Segmentation

Real images are noisy, unevenly lit, and full of detail a computer vision system does not need. Filtering removes some of that noise, and thresholding and segmentation separate the regions that matter from everything else. Together they make up the preprocessing stage of many classical OpenCV pipelines and of many deep learning data pipelines as well.
Part 1 covers color space conversions and four common smoothing filters: mean, Gaussian, median, and bilateral. Part 2 covers thresholding and segmentation, from simple and adaptive thresholds to Otsu binarization, GrabCut, connected components labeling, and the watershed algorithm. If you are new to OpenCV, start with the image manipulation guide, which introduces the transforms and masks used here.
Setup: imports, sample images, and helpers
The examples use OpenCV, NumPy, and Matplotlib, and load sample images from our public marketing-resources repository. The helper function displays several images side by side so each result can be compared with the original.
Imports and sample images
import cv2
import numpy as np
import urllib.request
import matplotlib.pyplot as plt# Collecting the sample image
image_url = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_3/Parrot_in_space.png"
resp = urllib.request.urlopen(image_url)
image_bytes = np.asarray(bytearray(resp.read()), dtype=np.uint8)Helper functions
# Function for the creation of flexible MatPlotLib figures
def create_mpl_figure(w,h,images,titles="Image",axis="off"):
plt.figure(figsize=[w,h])
for i, image in enumerate(images):
plt.subplot(1,len(images),i+1); plt.imshow(image); plt.title(titles[i]); plt.axis(axis);
# Function for the addition of Gaussian Noise to an Image
def add_gaussian_noise(image, mean=0, std=25):
noise = np.random.normal(mean, std, image.shape).astype(np.uint8)
noisy_image = cv2.add(image, noise)
return noisy_image
# Function for the addition of Salt and Pepper noise to an Image
def add_salt_and_pepper_noise(image, noise_ratio=0.02):
noisy_image = image.copy()
h,w,c = image.shape
noisy_pixels = int(h*w*noise_ratio)
for _ in range(noisy_pixels):
row, col = np.random.randint(0,h), np.random.randint(0,w)
if np.random.rand() < 0.5:
noisy_image[row, col] = [0, 0, 0]
else:
noisy_image[row, col] = [255, 255, 255]
return noisy_imagePart 1: Color spaces, filtering, and blurring
Smoothing filters replace each pixel with a combination of its neighbors. The choice of filter decides what is removed and what survives: some average everything, some ignore outliers, and some preserve edges while smoothing flat regions.
Where these techniques are used
Of the techniques in this part, things like color space conversions and how they allow for swifter image processing and image thresholding (used to create masks in the image manipulation guide) or various different types of filtering to allow for removal of noise from images or edge detection, these techniques include a number of common pre-processing steps for other techniques.
Color Space Conversions
What it does: As the name suggests, color space conversions convert an image of one color space (Grayscale, BGR, RGB, HSLuv, YCrCb) and converts it to another color space.
Why it matters: Color space conversions can be exceedingly useful. If you want to increase processing speed, you can change images to the grayscale color space, and cut down on image processing times up to ~66%. If you want to work on color based image masking, changing to another color space can assist considerably in the task.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color space conversions
rgb_image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
gray_image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2GRAY)
hsv_image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2HSV)
hls_image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2HLS)
lab_image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2LAB)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [bgr_image, rgb_image, gray_image, hsv_image, hls_image, lab_image], ["BGR", "RGB", "Gray Scale", "HSV", "HLS", "Lab"])
Mean Filtering
What it does: Mean filtering evaluates every pixel, and according to a box filter around the given pixel of _n_ x _n_ pixels, the pixel weight is averaged by taking all pixel values, adding them together, and then dividing by the number of pixels evaluated.
Why it matters: Mean filtering can be used to assist in the reduction of noise within images, as well as smoothing images. It is a common pre-processing step for many different image manipulation techniques.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Mean Filtering
box_filter = (15,15)
mean_image = cv2.blur(image, box_filter)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, mean_image], ["Original", "Mean Filtered"])
Gaussian Filtering
What it does: Gaussian filtering filters images by taking a kernel of _n_ x _n_ pixels from around the targeted pixel. According to the Gaussian curve, a bell-shaped curve that weighs those values towards the center highly and those further from the center lower, each pixel is once again filtered according to the values of the surrounding pixels, weighing more highly those closer to it. This is different than mean filtering, where each pixel within the box filter is weighed identically.
Why it matters: Gaussian filtering is extremely effective at handling gaussian noise, pre-processing images for edge detectors so that signal noise is not mis-interpreted as edges, or in real-time graphics creating depth of field or heat blur.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Adding Gaussian Noise to the Image
noisy_image = add_gaussian_noise(image, mean=0, std=1)
# Gaussian Filtering
gaussian_kernel = (25,25)
gaussian_image = cv2.GaussianBlur(noisy_image, gaussian_kernel, 0)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, noisy_image, gaussian_image], ["Original", "Gaussian Noise added to Image", "Gaussian Filtered"])
Median Blurring
What it does: Median filtering, like the gaussian filtering before it takes a kernel that evaluates a given pixel and the pixels surrounding it according to the kernel _n_ x _n_ pixels. This kernel is a square denoted by a single positive, odd integer. From these values, median blurring will take the middle value, and apply it to the targeted pixel.
Why it matters: Median filtering is very effective for handling salt and pepper noise. Whether it be handling visual anomalies in CCTV or dashcam footage, median filtering is able to remove the salt and pepper noise without damaging detail to a very noticeable degree.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Adding Salt and Pepper Noise to the Image
noisy_image = add_salt_and_pepper_noise(image, noise_ratio=0.20)
# Median Filtering
median_kernel = 3
median_image = cv2.medianBlur(noisy_image, median_kernel)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, noisy_image, median_image], ["Original", "Salt and Pepper Noise added to Image", "Median Filtered"])
Bilateral Filtering
What it does: Bilateral filtering smooths images while retaining sharp edges. It does this by averaging each pixel according to its surroundings if neighboring pixels are nearby to the evaluated pixel, and similar in intensity (or color dependent). It is similar in result to Gaussian filtering, but in this case, this technique is far more effective at retaining edges, though it less performance friendly.
Why it matters: Bilateral filtering is the most advanced filtering technique covered thus far and is exceptionally useful because of its ability to retain sharp details in images while blurring less important ones.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Add Gaussian noise to the Image
noisy_image = add_gaussian_noise(image, mean=0, std=1)
# Bilateral Filtering
bilateral_image = cv2.bilateralFilter(noisy_image, 15, 200, 200)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, noisy_image, bilateral_image], ["Original", "Gaussian Noise added to Image", "Bilateral Filtered"])
```
Part 2: Thresholding and image segmentation
Segmentation divides an image into meaningful regions. The techniques below range from a single global cutoff to marker-based algorithms that can separate objects touching each other.
Sample images for this part
# Collecting the sample image
image_url = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_4/Parrot_on_snowmobile.png"
resp = urllib.request.urlopen(image_url)
image_bytes = np.asarray(bytearray(resp.read()), dtype=np.uint8)
image_url_watershed = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_4/Parrots_watershed.png"
resp = urllib.request.urlopen(image_url_watershed)
image_bytes_watershed = np.asarray(bytearray(resp.read()), dtype=np.uint8)
image_url_ccl = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_4/Parrot_simple.png"
resp = urllib.request.urlopen(image_url_ccl)
image_bytes_ccl = np.asarray(bytearray(resp.read()), dtype=np.uint8)The segmentation examples display grayscale and label images, so this part uses an extended version of the helper that accepts Matplotlib color maps.
# Function for the creation of flexible MatPlotLib figures
def create_mpl_figure(w,h,images,titles="Image",axis="off",color_maps=None):
plt.figure(figsize=[w,h])
for i, image in enumerate(images):
plt.subplot(1,len(images),i+1);
if color_maps is None:
plt.imshow(image);
elif len(color_maps) > 1:
plt.imshow(image, cmap=f"{color_maps[i]}");
else:
plt.imshow(image, cmap=f"{color_maps[0]}")
plt.title(titles[i]);
plt.axis(axis);Where these techniques are used
Whether it be presence sensing for 3D printers using simple thresholding, warehouse shelf label reading using adaptive thresholding, or automatic pill counters using otsu's binarization, each of the image segmentation techniques in this part has a variety of different use cases that are quite common.
Simple Thresholding
What it does: Simple thresholding alters the target image by taking a user-inputted threshold and evaluating every pixel on the image by that value. If the evaluated pixel is less than the threshold, it will be set to 0, and if greater than, it will be set to the maximum intensity.
Why it matters: Simple thresholding allows users to isolate objects in images & videos, do basic edge detection, simplify images for more efficient processing, and as the masking examples in the image manipulation guide show, it is quite useful for creating masks on images.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color space conversions
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Splitting the channels for thresholding an individual color
r,g,b = cv2.split(image)
# Simple Thresholding
ret, thresh_g = cv2.threshold(g, 255, 255, cv2.THRESH_BINARY)
green_thresholded = cv2.merge([r,thresh_g,b])
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, green_thresholded], ["Original", "Green Thresholded"])
Image Binarization
What it does: Image binarization is a technique by which a given image is converted from its original coloration to a binary intensity set based upon a given threshold.
Why it matters: Image binarization has a number of different applications, but two of the most common would be in feature & edge detection, as well as in the reduction of data sizes, so that training and image consumption of a given model that is being trained is much, much faster, as the evaluation of the image is only happening on one channel, rather than on three or more.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes, cv2.IMREAD_GRAYSCALE)
# Image Binarization
ret, binarized_image = cv2.threshold(image, 100, 255, cv2.THRESH_BINARY)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, binarized_image], ["Grayscale Original", "Binarized"], "off", ["gray"])
Adaptive Thresholding
What it does: Adaptive Thresholding, rather than applying a global threshold to the entire image like in simple thresholding, it applies a dynamic threshold that is dependent upon the neighborhood of a given pixel.
Why it matters: Adaptive thresholding is very effective for the processing of documents for instance. Where a global threshold might be able to make some of the content of the scanned document clearer in one part of the image, it may end up obscuring detail if the lighting changes at points in the image. This is a situation where a technique like adaptive thresholding shines, as it is dynamic, and can produce consistent results regardless of original pixel intensity.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes, cv2.IMREAD_GRAYSCALE)
# Adaptive Thresholding
thresh = cv2.adaptiveThreshold(image, 255, cv2.ADAPTIVE_THRESH_MEAN_C, cv2.THRESH_BINARY, 25, 1)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, thresh], ["Original", "Adaptive Thresholding"], "off", ["gray"])
Otsu's Binarization
What it does: Otsu's Binarization serves a similar purpose to simple thresholding, but where in simple thresholding you must determine the threshold yourself, otsu's binarization determines the optimal threshold automatically. In this technique, the threshold passed is arbitrary, as it is overwritten automatically.
Why it matters: Otsu's binarization and its ability to find the optimal threshold value based upon the image's histogram is very effective for document analysis, even in some less than optimal lighting conditions. Additionally, it is often used within medical imaging, as well as license plate recognition.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes, cv2.IMREAD_GRAYSCALE)
# Otsu's Binarization
ret, otsu_image = cv2.threshold(image, 0, 255, cv2.THRESH_BINARY + cv2.THRESH_OTSU)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, otsu_image], ["Original", "Otsu's Binarization"], "off", ["gray"])
Grabcut Algorithm
What it does: Grabcut Algorithm is an image segmentation technique that separates the foreground from the background. It requires user interaction, typically in the form of a user drawing a rectangle around the given subject.
Why it matters: The Grabcut algorithm is first and foremost useful for foreground isolation from the background. It is not the fastest technique, to be sure, and while it may occasionally have artifacts surrounding your subject (as can be seen in the example), it is still a very useful technique for image segmentation.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Image Preprocessing
mask = np.zeros(image.shape[:2], np.uint8)
bgdModel = np.zeros((1, 65), np.float64)
fgdModel = np.zeros((1, 65), np.float64)
rectangle = (400, 250, 500, 450)
# Grabcut Algorithm
cv2.grabCut(image, mask, rectangle, bgdModel, fgdModel, 5, cv2.GC_INIT_WITH_RECT)
mask2 = np.where((mask == 2)|(mask == 0), 0, 1).astype('uint8')
image_segmented = image * mask2[:, :, np.newaxis]
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, image_segmented], ["Original", "Grabcut"], "on")
Connected Components Labeling
What it does: Connected components labeling, as the name suggests, seeks to determine the connectivity of blobs (distinct regions within an image with regard to color or intensity) within a given image.
Why it matters: Connected components labeling is useful for the counting of objects within an image, the identification and tracking of those same objects, as well as in other applications such as defect detection in manufacturing.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes_ccl, cv2.IMREAD_GRAYSCALE)
# Image Preprocessing
ret, thresh = cv2.threshold(image, 0, 255, cv2.THRESH_BINARY | cv2.THRESH_OTSU)
# Connected Components Labeling
connectivity = 4
output = cv2.connectedComponentsWithStats(thresh, connectivity, cv2.CV_32S)
(numLabels, labels, stats, centroids) = output
images = []
images.append(image)
for i in range(0, numLabels):
# Printing the component information as each is evaluated
# The first component is ALWAYS the background, so we add a suffix to indicate this
suffix = " (background)" if i == 0 else ""
text = f"examining component {i + 1}/{numLabels}{suffix}"
print(f"[INFO] {text}")
# Extracting the component information
x = stats[i, cv2.CC_STAT_LEFT]
y = stats[i, cv2.CC_STAT_TOP]
w = stats[i, cv2.CC_STAT_WIDTH]
h = stats[i, cv2.CC_STAT_HEIGHT]
area = stats[i, cv2.CC_STAT_AREA]
(cX, cY) = centroids[i]
# Creating a copy of the original image on which to draw the component information
output = image.copy()
cv2.rectangle(output, (x, y), (x + w, y + h), (0, 255, 0), 3)
cv2.circle(output, (int(cX), int(cY)), 4, (0, 0, 255), -1)
# Creating the mask for each component
componentMask = (labels == i).astype("uint8") * 255
# Appending the images to the list for display
images.append(output)
images.append(componentMask)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, images, ["Original", "Background Component", "Background Mask", "Parrot Component", "Parrot Mask"], "off", ["gray"])
Watershed Algorithm
What it does: The watershed algorithm treats each image evaluated as a topographical map where areas of high intensity denote peaks, and those of low intensity denote valleys. Assuming one were to punch holes at the bottom of each valley, and the image started to flood from these locations equally, each valley would begin to fill with different colored water (labels). When two of these bodies of water are getting ready to merge, you would will draw a damn. After the whole image is flooded, the damns would delineate your object boundaries.
Consider it in two steps. In step 1, you create hints (sure background and foreground). In step 2, you feed those hints to Watershed, which makes sense of these hints, defining precise boundaries.
Why it matters: The watershed algorithm finds consistent use in medical fields, specifically in MRIs and CAT scans, as well as in other applications like traffic analysis and object recognition. It is a very flexible object detection technique that is quite flexible, once one understands how to implement it.
The code and output
# Reading the sample image
original_img = cv2.imdecode(image_bytes_watershed, cv2.IMREAD_COLOR)
img = cv2.imdecode(image_bytes_watershed, cv2.IMREAD_COLOR)
g_image = cv2.imdecode(image_bytes_watershed, cv2.IMREAD_GRAYSCALE)
# Image pre-processing
# Thresholding the image, converting it to an intensity space
ret, thresh = cv2.threshold(g_image, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)
# Cleaning up any artifacting in the image
kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3, 3))
refined = cv2.morphologyEx(thresh, cv2.MORPH_OPEN, kernel, iterations=2)
# Our sure background. This will inform the algorithm where not to place markers (water)
sure_bg = cv2.dilate(refined, kernel, iterations=3)
# Distance transform. For every foreground pixel, this computes euclidean distance to nearest background pixel.
dist = cv2.distanceTransform(refined, cv2.DIST_L2, 5)
# Foreground Area. Shrinks the foreground objects to create sure internal markers of each valley.
ret, sure_fg = cv2.threshold(dist, 0.5 * dist.max(), 255, cv2.THRESH_BINARY)
sure_fg = sure_fg.astype(np.uint8)
# Unknown area. This are makes up pixels that are background AND not part of the shrunken foreground markers.
# They are ambiguous, and marking them unknown allows the watershed algorithm to decide on what they are later.
unknown = cv2.subtract(sure_bg, sure_fg)
# Labeling
# Scans the sure foreground, assigning unique integers to each connected blob in the binary mask.
ret, markers = cv2.connectedComponents(sure_fg)
# connectedComponents starts at 0 (background), and we already have one, so increment each marker by 1
markers += 1
# Mark the region of unknown with zero to be considered as valleys during the implementation of the algorithm.
markers[unknown == 255] = 0
# Watershed Algorithm
markers = cv2.watershed(img, markers)
labels = np.unique(markers)
parrots = []
for label in labels[2:]:
# Create a binary image in which only the area of the label is in the foreground
# And the rest of the image is in the background
target = np.where(markers == label, 255, 0).astype(np.uint8)
# Perform contour extraction on the created binary image
contours, hierarchy = cv2.findContours(
target, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE
)
parrots.append(contours[0])
# Draw the outline
watershed_img = cv2.drawContours(img, parrots, -1, color=(0, 254, 0), thickness=2)
# Color conversion for proper display
o_img = cv2.cvtColor(original_img, cv2.COLOR_BGR2RGB)
w_img = cv2.cvtColor(watershed_img, cv2.COLOR_BGR2RGB)
# Creation of the MatPlotLib figure for comparison of images
plt.figure(figsize=[30,10])
plt.subplot(1, 2, 1)
plt.imshow(o_img, vmin=0, vmax=255)
plt.title("Original")
plt.axis('off')
plt.subplot(1, 2, 2)
plt.imshow(w_img, vmin=0, vmax=255)
plt.title("Watershed Algorithm")
plt.axis('off')
plt.show()
Where to go next
Filtering and segmentation decide what the rest of a vision pipeline gets to see. A well-chosen blur keeps edges intact while removing noise, and a good threshold or marker strategy turns a photograph into regions you can count, measure, or pass to a model.
Continue with edge, shape, and feature detection with OpenCV, which builds on these preprocessing steps. To revisit the basics, see the OpenCV image manipulation guide. For applied examples, see face detection with OpenCV and object detection using YOLO.
If you are taking a vision pipeline beyond a notebook, our computer vision development services cover data and labeling, model selection, evaluation on your own images, and deployment to cloud or edge hardware.
Frequently Asked Questions
When should I use a median blur instead of a Gaussian blur?
A median blur is usually the better choice for salt-and-pepper noise, because it ignores isolated outliers instead of averaging them in. A Gaussian blur suits general sensor noise. A bilateral filter preserves edges better than either, at a higher computational cost.
What does Otsu's binarization do?
Otsu's method chooses a global threshold automatically from the image histogram, selecting the value that best separates the pixels into two classes. It works best when the histogram has two clear peaks, such as a dark object on a light background.
What is the difference between global and adaptive thresholding?
Global thresholding applies one cutoff to the whole image. Adaptive thresholding computes a separate threshold for each neighborhood, which handles uneven lighting such as shadows across a scanned document.
When should I use watershed instead of GrabCut?
GrabCut extracts a foreground object from a rough rectangle or mask using iterative graph cuts. Watershed is better suited to separating multiple touching objects, such as coins or cells, once you have markers for sure foreground and sure background.