RawZipper lossy and lossless performance benchmarks on NVIDIA GPU
Updated:
The adaptive local tone mapping algorithm is supposed to improve image quality. We need local tone mapping because our eyes perceive the world locally and with dynamic adaptation, while standard digital sensors and displays have severe limitations in handling the vast range of brightness found in the real world.
The global tone mapping algorithm is pretty simple. It uses the same curve or function for every single pixel in the image. It does this based on the image's global statistics, like the average luminance. When we do it that way, we lose detail in both the brightest (highlights) and darkest (shadows) parts of the image at the same time. The image looks flat and unnatural. Actually, the local tone mapping algorithm adjusts the brightness pixel by pixel and region by region based on the local context of the surrounding pixels.
Key Advantages of RawZipper
- Lossless and Lossy encoding
- Compression ratio control for lossy encoding
- Real time raw encoding on NVIDIA GPUs, including Jetson
We need local tone mapping because it is an essential computational bridge between the physical reality of light and the limitations of our digital displays. It allows us to faithfully reproduce the rich visual experience of the real world with all its bright highlights and deep shadows on a standard screen, preserving detail, enhancing contrast, and creating a more natural and compelling image.
Here's some info about how fast that algorithm performs on mobile and desktop GPUs from NVIDIA. This will help you to get an understanding about its processing speed.
Images for evaluation
- Image file format – PGM, 1-channel, 12-bit, RGGB or other bayer patterns
- Standard image resolutions for testing: 1920 × 1080 (2K), 3840 × 2160 (4K), 7680 × 4320 (8K)
- Additional high resolutions: 9344 × 7000 (63 MPix), 11272 × 9200 (99 MPix), 14204 × 10652 (148 MPix)
Hardware and software for NVIDIA Jetson configurations
- Jetson Orin NX 8GB (Ampere, 32 SMMs, 1024 cores) with 8-core ARM, Jetpack 6.2, CUDA Toolkit 12.6
- Jetson Orin AGX 64GB (Ampere, 64 SMMs, 2048 cores) with 12-core ARM, Jetpack 6.2, CUDA Toolkit 12.6
Hardware and software for desktop configuration
- GPU NVIDIA GeForce RTX 4090 (Ada Lovelace, 128 SMMs, 16384 cores) with CPU AMD Ryzen9 7950X (16 cores)
- OS Windows 11 Pro (x64), version 23H2, CUDA Toolkit 12.6
How we've done performance measurements
We've done time and performance measurements for the RawZipper codec for 12-bit raw bayer images with 2K, 4K, and 8K resolutions, and more. All the results we've got here don't include any host I/O latency. That's the time it takes to load an image into RAM from a hard disk or solid-state drive, and also to save it back. We also didn't include the time it takes to transfer data between devices. We've made an assumption about how to use the encoder in our standard image processing pipeline, when the initial data is stored in GPU memory. Check out the table below for our averaged measurement results for the best series of 1000 processed frames. We've been using 12-bit PGM images for evaluation because this is the main use case for raw encoding.
| GPU \ Image resolution | 1920 × 1080 (2K) | 3840 × 2160 (4K) | 7680 × 4320 (8K) | 9344 × 7000 (63 MPix) | 11272 × 9200 (99 MPix) | 14204 × 10652 (148 MPix) |
| Jetson Orin NX 8GB | 24 ms | 44 ms | 72 ms | -- | -- | -- |
| Jetson Orin AGX 64GB | 12 ms | 22 ms | 36 ms | -- | -- | -- |
| GeForce RTX 4090 | 0.5 ms | 1.2 ms | 2.2 ms | 4 ms | 4.5 ms | -- |
Table 1 - Performance benchmarks for RawZipper lossy encoding algorithm on Jetson Orin NX/AGX and GeForce RTX 4090
As usualy, the higher performance could be achieved with bigger image resolutions. This is actually the case for how well software designed with CUDA is performing when it comes to ISP processing. For a 63 MPix image, we're getting over 3 GPix/s on RTX 4090, which is a very promising result. The above benchmarks show that we can use RawZipper for real-time camera applications, particularly with Jetson Orin hardware.
How to verify these numbers: they were measured on the NVIDIA GPU stated above and are reproducible on your own GPU — download the demo apps and measure the speed on your own images.