OpenCV Image Manipulation: Transforms, Arithmetic, and Masks

OpenCV is one of the most widely used open-source libraries for computer vision. Before a system can detect objects, read documents, or segment an image, it needs dependable ways to move, resize, combine, and isolate pixels. This guide covers those foundations in Python, with the code and the output for every technique.
Part 1 covers drawing and geometric transforms: annotating images, resizing, translation, rotation, flipping, cropping, and image pyramids. Part 2 moves to pixel-level control: image arithmetic, blending, bitwise operations, masking, and working with individual color channels. Most of these operations reappear later as preprocessing steps, including in the filtering, thresholding, and segmentation techniques covered in the next guide.
Setup: imports, sample images, and helpers
The examples use OpenCV, NumPy, and Matplotlib, and load sample images from our public marketing-resources repository. The helper function displays several images side by side so each result can be compared with the original.
Imports and sample images
import cv2
import numpy as np
import urllib.request
import matplotlib.pyplot as plt#Collecting the sample image
image_url = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_1/Parrot-on-branch.png"
resp = urllib.request.urlopen(image_url)
image_bytes = np.asarray(bytearray(resp.read()), dtype=np.uint8)Helper function
# Function for the creation of flexible MatPlotLib figures
def create_mpl_figure(w,h,images,titles="Image",axis="off"):
plt.figure(figsize=[w,h])
for i, image in enumerate(images):
plt.subplot(1,len(images),i+1); plt.imshow(image); plt.title(titles[i]); plt.axis(axis);Part 1: Drawing and geometric transforms
Geometric operations change where pixels are, not what they are. They are the first tools most computer vision projects reach for, whether the goal is to annotate a detection, normalize input sizes for a model, or create additional training examples.
Where these techniques are used
This part covers a variety of basic image manipulation techniques, all of which have a number of very useful implementations. Whether it be using rectangles, circles, resizing or cropping in order to track objects in images or videos, or using a combination of cropping, rotation, flipping in order to train a machine learning model to better recognize a given subject, each of these techniques has a variety of different real world applications which makes them useful.
Drawing on Images
What it does: Adds visual elements (shapes and text) to an image
Why it matters: Drawing on images is useful for annotation, visualization, or debugging during processing.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
image_edit = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Image Transformations
cv2.line(image_edit, (550, 870), (1000, 820), (0, 255, 0), 2)
cv2.rectangle(image_edit, (650, 225), (1000, 500), (255, 255, 255), 3)
cv2.circle(image_edit, (307, 550), 60, (0, 0, 255), -1)
cv2.putText(image_edit, "OpenCV", (650, 200), cv2.FONT_HERSHEY_SIMPLEX, 3, (0, 255, 255), 2)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image, image_edit], ["Original", "Edited"])
Resizing
What it does: Resizing scales an image according to desired size of an image, scale factors fx and fy, and an interpolation method.
Why it matters: Resizing, and more broadly image scaling as a whole can be very useful when it comes to image processing in the process of training a machine learning model. By reducing the number of pixels in an image, it can reduce the amount of time spent training for a given model by presenting it with less complex, albeit less accurate training data.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Getting the dimensions of the image
height, width = image.shape[:2]
# Creating the resized image
resized_image = cv2.resize(image, (60,30), fx=0.1, fy=0.1, interpolation=cv2.INTER_LINEAR)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image, resized_image], ["Original", "Resized"])
Translation
What it does: Shifts a given image by a specified number of pixels along the x and y axes. These pixel shifts can be represented by the following notations: tx & ty
Why it matters: Image translations are often used for object tracking, image alignment, and augmentation of data used for machine learning.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Image Translation
# Getting the dimensions of the image
height, width = image.shape[:2]
# Create translation values
tx, ty = width / 4, height / 4
# Create translation matrix
translation_matrix = np.array([
[1, 0, tx],
[0, 1, ty]
], dtype=np.float32)
# Apply matrix to image
translated_image = cv2.warpAffine(src=image, M=translation_matrix, dsize=(width, height))
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image, translated_image], ["Original", "Translated"])
Rotation
What it does: Allows users to rotate a given image about a certain point in the image by a specified number of degrees using a rotation matrix, or by the center of the image in 90 degree increments by using the getRotationMatrix2D() & warpAffine() methods or the rotate() method accordingly.
Why it matters: Rotation can be used to automate rotation of important physical documents submitted electronically, increasing accuracy of other methods for recognizing text and images on scanned / photographically captured documents.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Specific Degree Rotation
# Getting the dimensions of the image and the centerpoint
height, width = image.shape[:2]
center = (width/2, height/2)
# Creating rotation matrix and applying it to the image, retaining the same dimensions
rotate_matrix = cv2.getRotationMatrix2D(center=center, angle=45, scale=1)
rotated_image = cv2.warpAffine(src=image, M=rotate_matrix, dsize=(width, height))
# 90-Degree Increment Rotation
rotated_image_2 = cv2.rotate(image, cv2.ROTATE_180) # Also try ROTATE_90_CLOCKWISE, ROTATE_90_COUNTERCLOCKWISE
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image, rotated_image, rotated_image_2], ["Original", "Rotated 45 Degrees", "Rotated 180 Degrees"])
Flipping
What it does: Flips a given image about either the x axis, the y axis, or the x and y axis using the according flip codes 0, 1, and -1.
Why it matters: Flipping images can create more reliable machine learning models by providing them with new data samples of existing data on which they have been trained. It can be also used to correct orientation of camera feeds for surveillance purposes, among many other uses.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Flipping the sample image
flipped_image = cv2.flip(image, -1)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image, flipped_image], ["Original", "Flipped"])
Cropping / Zooming
What it does: Displays a certain section of an image, defined by slicing a given image.
Why it matters: Cropping can be very useful for object detection and recognition as a pre-processing step by cropping out the relevant portion of an image, allowing faster and more accurate recognition. Additionally, it can be used as a step of image of segmentation, and other techniques for image analysis.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Get image shape (width = 1024, height = 1536, channel = 3)
print(image.shape)
# Crop image (image[min_y:max_y, min_x:max_x])
cropped_image = image[220:500, 650:1000]
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image, cropped_image], ["Original", "Cropped"])
Image Pyramids
What it does: Image Pyramids upsample or downsample a given image.
Why it matters: Upsampling can be used to make smaller images more visible by making them larger, allowing them to be more accurately processed while increasing their size, whereas downsampling can decrease image sizes, enabling more images to be stored, as well as increasing the performance of image processing.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Upscaling and downscaling the image
downscaled_image_1 = cv2.pyrDown(image)
downscaled_image_2 = cv2.pyrDown(downscaled_image_1)
downscaled_image = cv2.pyrDown(downscaled_image_2)
upscaled_image = cv2.pyrUp(downscaled_image)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [downscaled_image, image, upscaled_image], ["Downscaled Image", "Original", "Upscaled Image created from downscaled_image"])
Part 2: Pixel-level control with arithmetic, bitwise operations, and masks
With the basics in place, this part works directly with pixel values. Brightness, contrast, blending, and bitwise logic are useful on their own, and they are also common preprocessing steps for thresholding, segmentation, and feature detection.
Sample images for this part
# Collecting the sample images
image_url = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_2/Parrot-on-pirate-ship.png"
resp = urllib.request.urlopen(image_url)
image_bytes = np.asarray(bytearray(resp.read()), dtype=np.uint8)
image_url_2 = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_1/Parrot-on-branch.png"
resp_2 = urllib.request.urlopen(image_url_2)
image_bytes_2 = np.asarray(bytearray(resp_2.read()), dtype=np.uint8)
image_url_3 = "https://raw.githubusercontent.com/SoftwareSushi/marketing-
resources/main/images/opencv/fundamentals/part_2/bitwise_parrot.png"
resp_3 = urllib.request.urlopen(image_url_3)
image_bytes_3 = np.asarray(bytearray(resp_3.read()), dtype=np.uint8)
image_url_4 = "https://raw.githubusercontent.com/SoftwareSushi/marketing-resources/main/images/opencv/fundamentals/part_2/bitwise_comparisons.png"
resp_4 = urllib.request.urlopen(image_url_4)
image_bytes_4 = np.asarray(bytearray(resp_4.read()), dtype=np.uint8)Where these techniques are used
Whether it is splitting different color channels to allow for the creation of color blind filters on videos / images, or using bitwise XOR operations to detect changes in image for video surveillance applications, the following techniques have a variety of different applications in the world that make them very useful to learn.
Image Addition, Subtraction, Blending
What it does: Image addition takes two images, and adds them together. This can result on overlays of images onto others, or increasing / decreasing brightness of one image by adding or subtracting a duplicate image of greater or lesser brightness.
Why it matters: Whether it be adding watermarks to copyrighted images, subtracting noise from an image, or detecting changes in between two similar frames of an image or video, image addition and subtraction have many practical uses.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
bgr_image2 = cv2.imdecode(image_bytes_2, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image_sample_1 = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
image_sample_2 = cv2.cvtColor(bgr_image2, cv2.COLOR_BGR2RGB)
# Image Addition and Subtraction
image_brighter = cv2.add(image_sample_1, 100)
image_darker = cv2.subtract(image_sample_1, 100)
image_blend = cv2.addWeighted(image_sample_1, 0.5, image_sample_2, 0.5, 0)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image_darker, image_sample_1, image_brighter], ['Darker', 'Original', 'Brighter'])
create_mpl_figure(30,10, [image_sample_1, image_blend, image_sample_2], ['Sample 1', 'Blended', 'Sample 2'])

Image Multiplication
What it does: Image multiplication increases or decreases the intensity values of pixels within an image, increasing or decreasing contrast within the given image. It multiplies either two images together, or an image and a scalar value. Scalars > 1 will result in higher contrast, and scalars < 1 will result in lower contrast.
Why it matters: One use of image multiplication can assist in the training of ML models, allowing them to learn to recognize a given object in lower and higher contrast situations.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Image multiplication
image_higher_contrast = cv2.multiply(image, 2.5)
image_lower_contrast = cv2.multiply(image, 0.5)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30,10, [image_lower_contrast, image, image_higher_contrast], ['Lower Contrast', 'Original', 'Higher Contrast'])
Bitwise AND
What it does: Bitwise AND evaluates either two images, or an image and a scalar. The result is another image where a pixel is set to 1 if the corresponding pixels in both inputs are 1.
Why it matters: Bitwise AND operations are very commonly used for masking, a technique covered later in this part.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes_3, cv2.IMREAD_COLOR)
image_2 = cv2.imdecode(image_bytes_4, cv2.IMREAD_COLOR)
# Bitwise AND operation
bitwise_and = cv2.bitwise_and(image, image_2)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, bitwise_and, image_2 ], ['Sample 1', 'Bitwise AND result of both Samples', 'Sample 2'],'on')
Bitwise OR
What it does: Bitwise OR evaluates either two images, or an image and a scalar. The result is another image where a pixel is set to 1 if, between the two images, at least of of them contains a corresponding pixel set to 1.
Why it matters: Moving beyond simple image combination, Bitwise OR can be useful for creating composite images where information from multiple different sources is stored. Should you ever have need of combining features and elements of different images into a distinct output image, Bitwise OR will be the best method.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes_3, cv2.IMREAD_COLOR)
image_2 = cv2.imdecode(image_bytes_4, cv2.IMREAD_COLOR)
# Bitwise OR operation
bitwise_or = cv2.bitwise_or(image, image_2)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, bitwise_or, image_2], ['Sample 1', 'Bitwise OR result of both Samples', 'Sample 2'], 'on')
Bitwise NOT
What it does: Bitwise NOT takes a single image or scalar, and inverts the bits of each pixel value. In simple terms, it takes the binary representation of every pixel (assuming 8-bit: 00000000 - 11111111) and inverts each bit.
Why it matters: Bitwise NOT has a variety of uses, but some of them are as follows: inversion of images, creating negative images, or highlighting important features of a given image.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes_3, cv2.IMREAD_COLOR)
# Bitwise OR operation
bitwise_not = cv2.bitwise_not(image)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, bitwise_not], ['Sample', 'Bitwise NOT result of the Sample'], 'on')
Bitwise XOR
What it does: Bitwise XOR evaluates either two images, or an image and a scalar. The result is an image where each pixel is set to 1 if the corresponding pixels from each image differ from one another. If they do not differ from one another, they remain the same.
Why it matters: Bitwise XOR is very useful for noting differences between two images, because pixels that are the same in each image compared will be 0 in the output, and images that are different will be set to non-0 values. This is exceedingly useful for change-detection, a common practice for video surveillance cameras to automatically detect movement.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes_3, cv2.IMREAD_COLOR)
image_2 = cv2.imdecode(image_bytes_4, cv2.IMREAD_COLOR)
# Bitwise XOR operation
bitwise_xor = cv2.bitwise_xor(image, image_2)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, bitwise_xor, image_2], ['Sample 1', 'Bitwise XOR result of both Samples', 'Sample 2'], 'on')
Masking
What it does: Masking is a technique where the user will define a ROI (region of interest) in the image, and then perform operations only within that region.
Why it matters: The use of masks allows for selective processing of images, that is, applying filters, color adjustments or other image manipulations to certain parts of an image without changing the rest of the image. Additionally, it enables a number of other techniques, such as object isolation, image composition, or region-based analysis, among others.
The code and output
# Reading the sample image
image = cv2.imdecode(image_bytes_3, cv2.IMREAD_COLOR)
image_2 = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image_foreground = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)
image_background = cv2.cvtColor(image_2, cv2.COLOR_BGR2RGB)
# Resizing the background image for consistency
mask_w = image_foreground.shape[0]
mask_h = image_foreground.shape[1]
image_background_resized = cv2.resize(image_background, (mask_h, mask_w))
# Masking Operation
gray_foreground = cv2.cvtColor(image_foreground, cv2.COLOR_BGR2GRAY)
# Create mask and inverted mask of the foreground
retval, mask = cv2.threshold(gray_foreground, 10, 255, cv2.THRESH_BINARY)
mask_inv = cv2.bitwise_not(mask)
# Apply foreground image to background
foreground = cv2.bitwise_and(image_background_resized, image_background_resized, mask=mask)
# Apply background image to foreground
background = cv2.bitwise_and(image_foreground, image_foreground, mask=mask_inv)
# Combine the two
masked_image = cv2.bitwise_or(foreground, background)
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image_foreground, masked_image, image_background_resized], ['Foreground', 'Masked Image', 'Background'])
Splitting & Merging Color Channels
What it does: Splitting and merging color channels of images, like the title suggests, splits a multi-channel array into several single-channel arrays and merges several arrays to make a single multi-channel array.
Why it matters: Splitting and Merging color channels allows users to make more precise adjustments to color when it comes to individually editing color channels, whether that be concerning their vibrancy, brightness, contrast, or some other characteristic. It can also allow for the creation of color-blind filters.
The code and output
# Reading the sample image
bgr_image = cv2.imdecode(image_bytes, cv2.IMREAD_COLOR)
# Color conversion to ensure proper display of images
image = cv2.cvtColor(bgr_image, cv2.COLOR_BGR2RGB)
# Color Splitting and Merging
r,g,b = cv2.split(image)
merged_channels = cv2.merge([r,g,b])
# Creation of the MatPlotLib figure for comparison of images
create_mpl_figure(30, 10, [image, r, g, b, merged_channels], ['Original', 'Red Channel', 'Green Channel', 'Blue Channel', 'Merged Channels'])
```
Where to go next
Geometric transforms and pixel operations are the building blocks of most OpenCV pipelines. Resizing and cropping control what a downstream model sees, flipping and rotation expand a training set, and masks let you process one region of an image without touching the rest.
Continue with OpenCV filtering, thresholding, and image segmentation, then edge, shape, and feature detection with OpenCV. For applied examples, see face detection with OpenCV and object detection using YOLO.
If you are taking a vision pipeline beyond a notebook, our computer vision development services cover data and labeling, model selection, evaluation on your own images, and deployment to cloud or edge hardware.
Frequently Asked Questions
Why do OpenCV images look blue in Matplotlib?
OpenCV loads color images in BGR channel order, while Matplotlib expects RGB. Convert with cv2.cvtColor(image, cv2.COLOR_BGR2RGB) before displaying the image.
What is the difference between cv2.add and adding NumPy arrays?
For 8-bit images, cv2.add saturates, so values above 255 are clipped to 255. Adding two uint8 NumPy arrays wraps around instead, which can turn bright pixels dark. Use the OpenCV functions when you want predictable brightness changes.
Which interpolation method should I use when resizing?
OpenCV recommends cv2.INTER_AREA when shrinking an image, and cv2.INTER_CUBIC (slower) or cv2.INTER_LINEAR (faster) when enlarging it.
What is a mask in OpenCV?
A mask is a single-channel image the same size as the input, where non-zero pixels mark the region of interest. Functions such as cv2.bitwise_and accept a mask so an operation only applies inside that region.