NVIDIA L40S Cloud GPU
Compare NVIDIA L40S specifications, AI performance, graphics capabilities and workload fit before choosing a cloud GPU provider.

NVIDIA Ada Lovelace data-center GPU
One GPU for AI, graphics and media
NVIDIA L40S is a data-center GPU built for generative AI, LLM inference and training, rendering, virtual production, video processing and NVIDIA Omniverse workloads. It combines 48 GB of ECC memory with Tensor Cores, RT Cores and dedicated media engines.
- Ada LovelaceGPU architecture
- 48 GBGDDR6 memory with ECC
- 864 GB/sMemory bandwidth
- 3× / 3×NVENC / NVDEC engines
Who should choose an L40S cloud GPU?
Choose L40S when you need a versatile GPU for AI inference, fine-tuning, image generation, rendering and video. It is less suitable when your workload needs more than 48 GB of memory, direct NVLink scaling or MIG-based partitioning.
NVIDIA L40S specifications
L40S is a passive, dual-slot PCIe data-center GPU. Its specification set combines AI compute, ray tracing, display output, media engines and virtual GPU support in one accelerator.
| Specification | NVIDIA L40S |
|---|---|
| Architecture | NVIDIA Ada Lovelace |
| GPU memory | 48 GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s |
| CUDA cores | 18,176 |
| Tensor Cores | 568 fourth-generation Tensor Cores |
| RT Cores | 142 third-generation RT Cores |
| FP32 | 91.6 TFLOPS |
| TF32 Tensor Core | 183 TFLOPS dense / 366 TFLOPS sparse |
| BF16 / FP16 Tensor Core | 362.05 TFLOPS dense / 733 TFLOPS sparse |
| FP8 Tensor Core | 733 TFLOPS dense / 1,466 TFLOPS sparse |
| RT Core performance | 212 TFLOPS |
| Interconnect | PCIe Gen4 x16, 64 GB/s bidirectional |
| NVLink | Not supported |
| MIG | Not supported |
| NVENC / NVDEC | 3× / 3×, including AV1 encode and decode |
| Display outputs | 4× DisplayPort 1.4a |
| Form factor | Dual-slot, passive PCIe card |
| Maximum power consumption | 350 W |
Sparse Tensor Core figures require supported structured-sparsity workloads. Available displays, vGPU profiles, codecs and virtualization features can depend on the server, cloud plan and software licence.
L40S advantage
Choose L40S for mixed AI and visual workloads
L40S combines FP8 Tensor performance, ray tracing, DisplayPort output, vGPU support and multiple media engines. It can replace separate AI inference, rendering and video accelerators in mixed workloads.
Important limitations
L40S does not support MIG or NVLink
L40S cannot be partitioned with Multi-Instance GPU and does not have a direct NVLink connection. Workloads requiring isolated hardware slices or tightly coupled multi-GPU scaling should consider another data-center GPU.
What makes L40S different?
Open each capability for a practical explanation of its role in cloud AI, graphics and media workloads.
Fourth-generation Tensor CoresFP8 acceleration for AI workloads
L40S supports TF32, BF16, FP16, FP8, INT8 and INT4 Tensor Core workloads. NVIDIA publishes up to 1,466 TFLOPS of FP8 Tensor performance with sparsity.
Third-generation RT CoresReal-time ray tracing and rendering
L40S includes 142 third-generation RT Cores and delivers 212 TFLOPS of RT Core performance for rendering, virtual production, simulation and 3D visualization.
Dedicated media enginesThree encoders and three decoders
L40S includes three NVENC and three NVDEC engines with AV1 encode and decode support, making it suitable for video processing, streaming and media AI pipelines.
DLSS 3 and graphics accelerationHigher frame rates for visual workloads
DLSS 3 combines deep learning, Tensor Cores and the Optical Flow Accelerator to improve rendering performance and responsiveness in supported applications.
Virtual GPU supportDesigned for virtual workstations
NVIDIA lists virtual GPU software support for L40S, allowing qualified environments to deliver accelerated virtual workstations and shared enterprise graphics workloads.
Data-center security and reliabilityBuilt for continuous enterprise operation
L40S supports secure boot with root of trust, is NEBS Level 3 ready and is designed for passive cooling in qualified data-center servers.
Best workloads for an L40S cloud GPU
Generative AI and LLM inference
A strong fit for inference, retrieval-augmented generation and selected fine-tuning workloads that fit within 48 GB of GPU memory.
Image and multimodal generation
Well suited to diffusion models, image generation, computer vision and multimodal pipelines that combine Tensor Core and graphics performance.
Rendering and 3D graphics
Designed for interactive rendering, virtual production, product design, architecture, engineering and other RTX-accelerated visual workloads.
Video processing and streaming
Three NVENC and three NVDEC engines, including AV1 support, make L40S useful for transcoding, streaming and AI-enhanced media workflows.
L40S is a strong fit when…
- You need one GPU for AI, rendering and video workloads.
- Your models and batches fit within 48 GB of GPU memory.
- You need FP8 support for compatible AI inference or training.
- You need RTX, ray tracing or virtual workstation capabilities.
- You need multiple hardware video encode and decode engines.
- You do not require NVLink or Multi-Instance GPU partitioning.
Consider another GPU when…
L40S may not be the best choice when your workload needs more than 48 GB, high FP64 throughput, direct GPU interconnects or MIG.
- Compare H100 for larger AI training and high-end inference.
- Compare H200 when memory capacity and bandwidth are the bottleneck.
- Compare A100 for MIG, FP64 HPC and NVLink-based configurations.
Provider comparison
Cloud providers offering NVIDIA L40S
Compare the listed L40S instance resources, deployment region, starting price and billing options. Check whether video engines, RTX features, display functions and virtual GPU licensing are exposed in the selected cloud configuration.
| Provider | L40S offering | GPU memory | Example configuration | Regions | Starting price | Billing options | Details |
|---|---|---|---|---|---|---|---|
| AceCloud | 1× NVIDIA L40S | 48 GB GDDR6 ECC | 16 vCPU · 64 GB RAM | India and United States | From ₹83,000/month | Monthly · 6/12-month plans · selected spot | View provider |
| E2E Networks | 1× NVIDIA L40S | 48 GB GDDR6 ECC | 60 vCPU · 220 GB RAM | India | ₹102/hour · ₹53,900/month | On-demand · monthly · 6 months · annual · spot subject to capacity | View provider |
| Cyfuture AI | 1× NVIDIA L40S | 48 GB GDDR6 | 16 vCPU · 256 GB RAM | India, United States and Europe | ₹124/hour on-demand · ₹61/hour for 12 months | On-demand · 1/6/12-month reserved | View provider |
| DigitalOcean | 1× NVIDIA L40S GPU Droplet | 48 GB GDDR6 | 8 vCPU · 64 GiB RAM | Toronto | $1.57/GPU-hour | Per-second on-demand · Spot not publicly listed | View provider |
| AWS | EC2 G6e | 48 GB GDDR6 per GPU | 1–8 GPUs · up to 192 vCPU and 1.5 TiB RAM | Selected AWS regions | Use AWS Pricing Calculator | On-Demand · Reserved Instances · Savings Plans · Spot | View provider |
Compare the complete instance rather than only the GPU price. Providers can differ in vCPU, RAM, storage, networking, region, commitment period, taxes, vGPU licensing and access to NVENC, NVDEC, RTX or display capabilities. Spot capacity can be interrupted and may not be available for every configuration.
NVIDIA L40S FAQs
How much GPU memory does NVIDIA L40S have?
L40S has 48 GB of GDDR6 memory with ECC and 864 GB/s of memory bandwidth.
Is NVIDIA L40S suitable for LLM inference?
Yes. L40S supports FP8 and provides 48 GB of memory, making it suitable for many LLM inference and selected fine-tuning workloads that fit within its memory capacity.
Can NVIDIA L40S be used for LLM training?
Yes, for models and training configurations that fit within 48 GB. Larger distributed training workloads may be better suited to GPUs with more memory and high-speed GPU interconnects.
Does NVIDIA L40S support MIG?
No. NVIDIA lists Multi-Instance GPU support as unavailable on L40S.
Does NVIDIA L40S support NVLink?
No. NVIDIA lists no NVLink support for L40S. Multi-GPU communication therefore uses the host system and PCIe rather than a direct NVLink connection.
Is L40S good for rendering and video?
Yes. L40S combines 142 third-generation RT Cores, 212 TFLOPS of RT performance, four DisplayPort 1.4a outputs and three NVENC plus three NVDEC engines.
Compare L40S configurations before choosing a provider
Compare the complete instance, not only the GPU name. Check CPU, RAM, storage, networking, region, vGPU licensing, video-engine access and whether the plan is optimized for AI, rendering or virtual workstations.