GPU server for video analytics: how many cameras per GPU on the L4, L40S and RTX PRO cards
Eurokommerz, Vienna, since 2006: Private AI/ML · IT Managed Services · Enterprise Training · AI Hardware & Software
- A GPU for video analytics is limited first by its hardware video decoders (NVDEC) and then by inference per analysed frame; GPU memory is rarely the first limit for CNN detectors
- The L4 has 4 NVDEC engines and 4 JPEG decoders at 72 W in a low-profile single-slot card, the L40S 3 decoders at 350 W, the RTX PRO 6000 Server Edition 4 Blackwell decoders at up to 600 W
- In NVIDIA’s DeepStream 7.1 documentation one L4 ran 68 H.264 or 81 H.265 streams at 1080p and 30 frames per second, with a ResNet18 detector on every frame, a tracker and two classifiers
- NVIDIA’s DeepStream 9.1 documentation (updated 28 July 2026) gives 1,524 frames per second for PeopleNet 2.6.3 with a tracker on the RTX PRO 6000 Server Edition and 1,114 on the L40S; analysing 10 of 30 frames allows up to about three times the cameras
- By our estimate for 1080p30 cameras, a light detector on every frame needs one L4 for 16 cameras, two for 64 and six for 256; PeopleNet on 10 of 30 frames needs one RTX PRO 6000 Server Edition for 64 and three for 256
Supplied by Eurokommerz: AI servers, built to order Request a configuration →
GPU for video analytics: decoders first, then inference
A GPU for video analytics is limited first by its hardware video decoders, the NVDEC engines that turn each H.264 or H.265 stream back into frames, and then by the inference work per analysed frame. For the CNN detectors most camera pipelines run, GPU memory is rarely the first limit. NVIDIA publishes the decode rate per NVDEC engine and the engines per card, a full DeepStream measurement for the L4, and frames per second for larger models on the L40S and the RTX PRO 6000.
From those figures, 16 cameras need one card, 64 cameras one or two cards and 256 cameras three to six, depending on the model and the frames per second it analyses. The camera counts are our estimates from NVIDIA’s figures, with sources and method under each table.
NVDEC, NVENC and JPEG decoders on the L4, L40S and RTX PRO cards
In DeepStream’s default mode the decoder plug-in decodes every frame of every stream, whether or not the model analyses it. Decode load therefore follows the cameras’ frame rate, not the analysis rate.
| GPU | NVDEC AND NVENC | POWER, COOLING | FORM FACTOR | 1080P30 DECODE LIMIT |
|---|---|---|---|---|
| NVIDIA L4 | 4 and 2, plus 4 JPEG decoders | 72 W, passive | low profile, single slot | 120 H.264, 218 H.265 |
| NVIDIA L40S | 3 and 3 | 350 W, passive | full height, dual slot | 90 H.264, 164 H.265 |
| RTX PRO 6000 Server Edition | 4 and 4 | up to 600 W, passive | full height, dual slot (air) | 289 H.264, 249 H.265 |
| RTX PRO 5000 | 3 and 3 | 300 W, active | full height, dual slot | 217 H.264, 187 H.265 |
| RTX PRO 4500 | 2 and 2 | 200 W, active | full height, dual slot | 144 H.264, 124 H.265 |
| RTX PRO 4000 | 2 and 2 | 145 W, active | full height, single slot | 144 H.264, 124 H.265 |
| RTX PRO 4000 SFF | 2 and 2 | 70 W, active | low profile, dual slot | 144 H.264, 124 H.265 |
| RTX PRO 2000 | 1 and 1 | 70 W, active | 2.7 by 6.6 in, dual slot | 72 H.264, 62 H.265 |
NVIDIA product pages, read on 9 October 2026; NVENC counts as in NVIDIA’s Video Codec SDK support matrix, read on the same day. Only the L4 page lists JPEG decoders. NVIDIA also lists a liquid-cooled, single-slot Server Edition. Decode limit: our arithmetic, engines × NVIDIA’s indicative rate per engine at 1920 × 1080 (NVDEC Application Note, Video Codec SDK 13.1) ÷ 30 frames per second, a ceiling, not a stream count.
The NVDEC Application Note for Video Codec SDK 13.1 gives indicative rates per engine at 1920 × 1080 in 4:2:0: 903 frames per second for H.264 and 1,641 for H.265 on Ada, and 2,172 and 1,872 on Blackwell. NVIDIA measured them at the highest video clock, 2,160 MHz on Ada and 2,362 MHz on Blackwell, and writes that performance “should scale according to the video clocks as reported by nvidia-smi on target GPU” and varies across GPU classes. The last column is therefore an upper bound. On the L4, NVIDIA’s own DeepStream run reached 68 H.264 streams against the 120 of the arithmetic.
All eight cards encode AV1 according to the support matrix, and the L40S page states that its engines include AV1 encode and decode. The L4’s four JPEG decoders serve cameras that send MJPEG or still images, formats the DeepStream decoder plug-in supports alongside H.264, H.265 and AV1. Our comparison of the 70 W cards covers the L4 and the RTX PRO 4000 SFF in detail.
Camera stream counts NVIDIA publishes for DeepStream
Of the cards in this article, we found a published stream count only for the L4. The performance page of the DeepStream 7.1 documentation, last updated on 15 September 2025, ran 1080p files at 30 frames per second through a pruned ResNet18 TrafficCamNet detector on every frame, an IOU tracker at 960 × 544 and two ResNet18 vehicle classifiers. On an AMD EPYC 7763 system with NVIDIA’s INT8 sample configuration, one L4 handled 68 H.264 streams at 75.74 per cent GPU and 8.06 per cent CPU utilisation, and 81 H.265 streams at 100 per cent GPU and 46.1 per cent CPU. With H.264 the GPU still had headroom, which suggests the decoders were the limit; with H.265 the inference work filled the GPU first. NVIDIA does not say which unit set either limit.
The DeepStream 9.1 documentation, last updated on 28 July 2026, gives frames per second per GPU instead, measured end to end, “considering video capture and decode, pre-processing, batching, inference, and post-processing”, with the 1080p H.265 sample file and rendering turned off. It does not state how many streams produced these figures.
| MODEL, TRACKER | INPUT TO MODEL | PRO 6000 SE | PRO 6000 WS | L40S |
|---|---|---|---|---|
| PeopleNet 2.6.3, MV3DT | 640 × 640, FP16 | 1,524 | 1,899 | 1,114 |
| RT-DETR, no tracker | 640 × 640, FP16 | 1,037 | 1,063 | 643 |
| Traffic | 544 × 960, FP16 | 920 | 994 | 643 |
| Grounding-DINO, no tracker | 544 × 960, FP16 | 159 | 178 | 107 |
NVIDIA DeepStream 9.1 documentation, Performance, dGPU pretrained models, frames per second per GPU with TensorRT, last updated 28 July 2026. TL is TrafficCamNet Transformer Lite. NVIDIA lists no L4 or RTX PRO 4000 figures in this table.
Divided by 30, the PeopleNet figure equals about 50 cameras analysed at full frame rate on one RTX PRO 6000 Server Edition and 37 on an L40S. At 10 analysed frames per second the counts rise by up to about three times, as decoding and the tracker still handle every frame. Open-vocabulary detectors such as Grounding-DINO run at about a tenth of that rate. The footnote to the “1,040 concurrent AV1 video streams at 720p30” on NVIDIA’s L4 page describes an AV1 encode run on a server with eight L4 cards, so the figure does not apply to analytics.
How to size a video analytics server
Size decode and inference separately and take the larger requirement.
- Decode load: cameras × the frame rate they send, per codec and resolution. Compare it with the decode limit in the first table and with NVIDIA’s measured L4 figure, and keep it well below the ceiling.
- Inference load: cameras × the frames per second the model analyses. In DeepStream the inference plug-in’s
intervalproperty sets the “number of consecutive batches to be skipped for inference”, sointerval=2analyses one frame in three, 10 of 30. The tracker still receives the skipped frames and tracks objects on them without new detections. - Divide the inference load by the frames per second NVIDIA publishes for the same or a similar model on the card, and plan to about 70 per cent of it, leaving room for peaks and growth.
- Check the CPU, system memory, network and power of the server, covered below.
Two decoder settings reduce work further. The decoder’s drop-frame-interval passes on only every n-th frame, “a value of 5 means the decoder outputs every fifth frame”, which saves the work after the decoder but not the decoding, so size the decoders for the full frame rate. skip-frames set to decode_key decodes key frames only, so the analysed rate follows the cameras’ key-frame interval; it suits scenes that change slowly. By pixel count, one 4K stream takes about four times the decode capacity of a 1080p stream, our assumption, as the application note gives 1080p figures only.
Cameras per GPU: configurations for 16, 64 and 256 cameras
The table applies this method to 1080p cameras at 30 frames per second, with two pipelines: a light CNN detector with a tracker and classifiers on every frame, as in NVIDIA’s L4 measurement, and PeopleNet 2.6.3 with MV3DT, NVIDIA’s multi-view 3D tracker, on 10 of 30 frames.
| CAMERAS, 1080P30 | FRAMES PER SECOND | CNN, ALL FRAMES | PEOPLENET, 10 FPS |
|---|---|---|---|
| 16 cameras | 480 decoded, 480 or 160 analysed | 1 L4 | 1 L40S, or an L4 after a test with your model |
| 64 cameras | 1,920 decoded, 1,920 or 640 analysed | 2 L4 | 1 RTX PRO 6000 Server Edition, or 1 L40S with H.265 cameras |
| 256 cameras | 7,680 decoded, 7,680 or 2,560 analysed | 6 L4 | 3 RTX PRO 6000 Server Edition, or 4 L40S with H.265 cameras |
Our estimates, not measurements: L4 counts from NVIDIA’s 68 H.264 streams per card (DeepStream 7.1), the other cards from DeepStream 9.1 frames per second and the decode limits in the first table, each planned to about 70 per cent, assuming the inference work scales with the analysed frames. NVIDIA publishes no PeopleNet figure for the L4.
The PeopleNet column names larger cards because NVIDIA publishes that model’s rate only for them. With H.264 cameras, 64 streams would use about 71 per cent of an L40S’s decode ceiling, so the table names the L40S for H.265 cameras only; the RTX PRO 6000 Server Edition, with four Blackwell decoders, handles either codec. For a small site without server airflow, the RTX PRO 4000 SFF has two Blackwell decoders at 70 W in a card with its own fan, but NVIDIA publishes no DeepStream figure for it, so test your model on it first.
We build GPU servers to order for these camera counts and check the rack, power and airflow before we quote. Send us your camera count, codecs and models through the form below for a configuration within one business day.
CPU, system memory, network and power around the cards
The CPU receives and parses the RTSP streams. In NVIDIA’s L4 run on an AMD EPYC 7763, 81 H.265 streams kept the CPU at 46.1 per cent and 68 H.264 streams at 8.06 per cent, so the codec changes the CPU load considerably. We found no NVIDIA figure for cores per stream, so plan the processor from a test run. Each decoded 1080p frame in NV12 takes about 3.1 MB (1920 × 1080 × 1.5 bytes), and the pipeline keeps several of them per stream, so GPU memory grows with the camera count; NVIDIA’s L4 runs with 68 and 81 streams fitted in the card’s 24 GB.
Plan the network from the camera bitrate times the camera count, plus the uplink to the video management system. For power, six L4 draw 432 W, four L40S 1,400 W and three RTX PRO 6000 Server Edition cards up to 1,800 W, before processors and fans. The L4, the L40S and the air-cooled Server Edition are passive and need the server’s airflow. Our article on how many GPUs fit in one server covers PCIe lanes, power supplies and airflow, and our L4 vs L40S comparison lists L4 counts per server model.
Video search and summarisation with vision language models
Search and summaries over recorded video use vision language models and an LLM, which move the limit from decoders to memory. NVIDIA’s Video Search and Summarization (VSS) blueprint, on its prerequisites page updated on 16 July 2026, says it was “validated and tested” on the H100, the RTX PRO 6000 Blackwell, the L40S, DGX Spark, IGX Thor and AGX Thor, and lists the RTX PRO 4500 Blackwell with “limited support to alerts profile”. Its base development profile runs on one RTX PRO 6000 Blackwell with the models sharing the GPU, or on two with dedicated GPUs; the L40S needs two. The search profile needs three to four RTX PRO 6000 cards, or four L40S, when all models run locally. NVIDIA also lists a minimum of an 18-core x86 CPU, 128 GB of RAM, a 1 TB SSD and a 1 Gbps network interface. On DGX Spark and the Thor platforms, the page lists only configurations with the LLM running remotely. Deploying open and commercial models on-premise is part of our Private AI/ML service.
We build AI servers to order with RTX PRO 6000 Server Edition or L40S cards for these profiles. Write to us with the hours of video and the questions your teams would ask of it.
Privacy and security of a video analytics server
Video in which people can be identified is personal data under the GDPR, and whether and how it may be analysed is an assessment for the company’s legal department. On the technical side, keep the cameras on their own network segment, give the analytics server only the access it needs and harden its management controller, drivers and container runtime as described in our guide to securing a GPU server.
What we supply
We supply the NVIDIA L4 and L40S, the RTX PRO 6000 Server Edition and the RTX PRO 2000, 4000, 4000 SFF, 4500 and 5000 as GPUs for servers you already run, or in AI servers built to order, assembled and burn-in tested, on one EU contract and invoice with manufacturer warranty. Before you order, we confirm that the card fits the chassis and slot and that the rack can power and cool it. Operating system, drivers, CUDA and a container runtime are installed on request, and models on top are our Private AI/ML service, with engineering by our partner Vixen.UNO.
FAQ
How many cameras can one GPU handle for video analytics?
Is the NVIDIA L4 suitable for video analytics?
What hardware does NVIDIA DeepStream need for many cameras?
L4 or L40S for video analytics?
How much GPU memory does video analytics need?
Which GPU for an NVR with AI analytics?
Send us the number of cameras, their resolution, frame rate and codec, the models you run, how many frames per second they analyse and where the server will stand. We reply within one business day with a configuration and a written quote, and we check the rack, power and airflow before we quote.
Talk to an expertWe reply within one business day