GPU Comparison

NVIDIA L40S vs L4

Compare NVIDIA L40S and L4 across architecture, GPU memory, bandwidth, AI inference, training, infrastructure requirements and practical workload fit.

Daya ShankarLast verified: August 11, 2026Research methodology

NVIDIA Ada Lovelace

L40S

GPU Memory

48 GB GDDR6 ECC

Bandwidth

864 GB/s

Ada data-center GPU for generative AI inference, graphics, rendering and media workloads.

Explore NVIDIA L40S
VS

NVIDIA Ada Lovelace

L4

GPU Memory

24 GB GDDR6

Bandwidth

300 GB/s

Low-power Ada accelerator optimized for inference, video, graphics and mainstream deployment.

Explore NVIDIA L4

L40S vs L4 at a Glance

Start with workload fit, then validate the choice against the exact cloud configuration and pricing available to you.

Choose L40S when

You need more than 24 GB.

Choose L4 when

24 GB is sufficient.

Compare the full workload

Memory, bandwidth, precision, interconnects, media features, power and cloud price can all change the right answer.

NVIDIA L40S vs NVIDIA L4 Specifications

H100, H200 and A100 use SXM figures. B200 uses current NVIDIA HGX B200 specifications. L40S and L4 use their native PCIe card specifications, so each product is represented in its primary deployment form.

SpecificationNVIDIA L40SNVIDIA L4
ArchitectureNVIDIA Ada LovelaceNVIDIA Ada Lovelace
GPU Memory48 GB GDDR6 ECC24 GB GDDR6
Memory Bandwidth864 GB/s300 GB/s
FP3291.6 TFLOPS30.3 TFLOPS
TF32 Tensor Core366 TFLOPS*120 TFLOPS*
FP16 / BF16 Tensor Core733 TFLOPS*242 TFLOPS*
FP8 Tensor Core1,466 TFLOPS*485 TFLOPS*
FP4 Tensor CoreNot natively supportedNot natively supported
NVLinkNot supportedNot supported
MIGNot supportedNot supported
Maximum Power350 W72 W
Form FactorDual-slot PCIeLow-profile single-slot PCIe

* Tensor Core values marked with an asterisk are NVIDIA sparse specifications where applicable; dense performance is lower.

What Is the Main Difference Between NVIDIA L40S and NVIDIA L4?

L40S and L4 share the Ada Lovelace architecture and fourth-generation Tensor Cores, but they target different performance envelopes.

L40S doubles memory capacity to 48 GB versus 24 GB on L4, increases bandwidth from 300 GB/s to 864 GB/s and lists roughly three times the FP32 and sparse FP8 Tensor Core throughput. It also carries a much higher 350 W maximum power rating.

L4's strength is efficiency: 72 W, single-slot low-profile deployment and strong media support. L40S is the higher-performance choice for larger inference models, rendering and heavier generative AI.

Memory Capacity and Bandwidth

These specifications affect model fit, KV-cache headroom, batch size and memory-bound workloads.

GPU Memory

48 GB GDDR6 ECC vs 24 GB GDDR6
L40SL4

Memory Bandwidth

864 GB/s vs 300 GB/s
L40SL4

L40S vs L4 for LLM Inference

Both support Ada FP8, but L40S has twice the memory and far more Tensor Core throughput. That makes it better for larger models, bigger batches and higher per-GPU throughput. L4 is well suited to smaller or quantized models where 24 GB is enough.

L40S vs L4 for Graphics and Generative AI

Both include Ada RT Cores, but L40S is the substantially more powerful graphics GPU and carries 48 GB of memory. It is better suited to rendering, virtual workstations, 3D workflows and heavier image-generation workloads.

L40S vs L4 for Power and Server Density

L4 uses only 72 W and fits a low-profile single-slot PCIe form factor. L40S can consume up to 350 W and uses a dual-slot card. For dense inference fleets, L4 can therefore be operationally easier even when L40S is much faster per GPU.

Which GPU Fits Your Workload?

Use this as directional guidance. Benchmark your own model and software stack before making a large infrastructure commitment.

WorkloadL40SL4Direction
LLM inference up to 24 GBBest performanceBest efficiencyDepends on throughput target
LLM inference 24–48 GBBest fitDoes not fitL40S
Image generationBest fitGoodL40S
RTX rendering / visualizationBest fitGoodL40S
Video AI / transcodingExcellentBest efficiency fitDepends on density
Power-constrained serversHigh powerBest fitL4
Low-profile single-slot deploymentNoBest fitL4

Cloud Pricing

Compare Current Provider Pricing

There is no single cloud price for either GPU. Rates vary by provider, region, server configuration, billing model and commitment. Compare current provider offers after you know which hardware class fits the workload.

Final Decision

Should You Choose NVIDIA L40S or NVIDIA L4?

Choose NVIDIA L40S if:

  • You need more than 24 GB.
  • Per-GPU inference throughput is more important than watts.
  • You run heavier rendering, visualization or generative AI.
  • A dual-slot 350 W PCIe GPU fits your server design.

Choose NVIDIA L4 if:

  • 24 GB is sufficient.
  • Power efficiency is a priority.
  • You need low-profile single-slot density.
  • Your workload is mainstream inference or video serving.

L40S vs L4 FAQs

Common questions about choosing between these NVIDIA GPUs.

Which is better, NVIDIA L40S or L4?
Choose L40S when you need substantially more AI/graphics performance and 48 GB of memory. Choose L4 when 24 GB is enough and power efficiency, low-profile density and mainstream inference or video serving matter more.
What is the main difference between L40S and L4?
L40S uses Ada Lovelace with 48 GB GDDR6 ECC and 864 GB/s memory bandwidth, while L4 uses Ada Lovelace with 24 GB GDDR6 and 300 GB/s. Tensor Core generation, precision support, interconnects and power can also differ.
Is L40S or L4 better for LLM inference?
Both support Ada FP8, but L40S has twice the memory and far more Tensor Core throughput. That makes it better for larger models, bigger batches and higher per-GPU throughput. L4 is well suited to smaller or quantized models where 24 GB is enough.
Which GPU is better for AI training, L40S or L4?
Both include Ada RT Cores, but L40S is the substantially more powerful graphics GPU and carries 48 GB of memory. It is better suited to rendering, virtual workstations, 3D workflows and heavier image-generation workloads.
How much memory do L40S and L4 have?
NVIDIA L40S provides 48 GB GDDR6 ECC, while NVIDIA L4 provides 24 GB GDDR6. Memory capacity alone does not determine performance, so bandwidth, precision support and workload behavior should also be considered.
Which is cheaper to rent, L40S or L4?
Cloud rental pricing for L40S and L4 varies by provider, region, configuration and billing model. Check current provider pricing rather than assuming one GPU is always cheaper.

Sources & Verification

Hardware specifications were checked against official NVIDIA product pages and documentation.

Last verified: August 11, 2026 · View our data source standards