NVIDIA L40S Cloud GPU

NVIDIA L40S Cloud GPU

Compare NVIDIA L40S specifications, AI performance, graphics capabilities and workload fit before choosing a cloud GPU provider.

NVIDIA L40S cloud GPU

NVIDIA Ada Lovelace data-center GPU

One GPU for AI, graphics and media

NVIDIA L40S is a data-center GPU built for generative AI, LLM inference and training, rendering, virtual production, video processing and NVIDIA Omniverse workloads. It combines 48 GB of ECC memory with Tensor Cores, RT Cores and dedicated media engines.

  • Ada LovelaceGPU architecture
  • 48 GBGDDR6 memory with ECC
  • 864 GB/sMemory bandwidth
  • 3× / 3×NVENC / NVDEC engines

Who should choose an L40S cloud GPU?

Choose L40S when you need a versatile GPU for AI inference, fine-tuning, image generation, rendering and video. It is less suitable when your workload needs more than 48 GB of memory, direct NVLink scaling or MIG-based partitioning.

NVIDIA L40S specifications

L40S is a passive, dual-slot PCIe data-center GPU. Its specification set combines AI compute, ray tracing, display output, media engines and virtual GPU support in one accelerator.

SpecificationNVIDIA L40S
ArchitectureNVIDIA Ada Lovelace
GPU memory48 GB GDDR6 with ECC
Memory bandwidth864 GB/s
CUDA cores18,176
Tensor Cores568 fourth-generation Tensor Cores
RT Cores142 third-generation RT Cores
FP3291.6 TFLOPS
TF32 Tensor Core183 TFLOPS dense / 366 TFLOPS sparse
BF16 / FP16 Tensor Core362.05 TFLOPS dense / 733 TFLOPS sparse
FP8 Tensor Core733 TFLOPS dense / 1,466 TFLOPS sparse
RT Core performance212 TFLOPS
InterconnectPCIe Gen4 x16, 64 GB/s bidirectional
NVLinkNot supported
MIGNot supported
NVENC / NVDEC3× / 3×, including AV1 encode and decode
Display outputs4× DisplayPort 1.4a
Form factorDual-slot, passive PCIe card
Maximum power consumption350 W

Sparse Tensor Core figures require supported structured-sparsity workloads. Available displays, vGPU profiles, codecs and virtualization features can depend on the server, cloud plan and software licence.

L40S advantage

Choose L40S for mixed AI and visual workloads

L40S combines FP8 Tensor performance, ray tracing, DisplayPort output, vGPU support and multiple media engines. It can replace separate AI inference, rendering and video accelerators in mixed workloads.

Important limitations

L40S does not support MIG or NVLink

L40S cannot be partitioned with Multi-Instance GPU and does not have a direct NVLink connection. Workloads requiring isolated hardware slices or tightly coupled multi-GPU scaling should consider another data-center GPU.

What makes L40S different?

Open each capability for a practical explanation of its role in cloud AI, graphics and media workloads.

  • Fourth-generation Tensor CoresFP8 acceleration for AI workloads

    L40S supports TF32, BF16, FP16, FP8, INT8 and INT4 Tensor Core workloads. NVIDIA publishes up to 1,466 TFLOPS of FP8 Tensor performance with sparsity.

  • Third-generation RT CoresReal-time ray tracing and rendering

    L40S includes 142 third-generation RT Cores and delivers 212 TFLOPS of RT Core performance for rendering, virtual production, simulation and 3D visualization.

  • Dedicated media enginesThree encoders and three decoders

    L40S includes three NVENC and three NVDEC engines with AV1 encode and decode support, making it suitable for video processing, streaming and media AI pipelines.

  • DLSS 3 and graphics accelerationHigher frame rates for visual workloads

    DLSS 3 combines deep learning, Tensor Cores and the Optical Flow Accelerator to improve rendering performance and responsiveness in supported applications.

  • Virtual GPU supportDesigned for virtual workstations

    NVIDIA lists virtual GPU software support for L40S, allowing qualified environments to deliver accelerated virtual workstations and shared enterprise graphics workloads.

  • Data-center security and reliabilityBuilt for continuous enterprise operation

    L40S supports secure boot with root of trust, is NEBS Level 3 ready and is designed for passive cooling in qualified data-center servers.

Best workloads for an L40S cloud GPU

  • Generative AI and LLM inference

    A strong fit for inference, retrieval-augmented generation and selected fine-tuning workloads that fit within 48 GB of GPU memory.

  • Image and multimodal generation

    Well suited to diffusion models, image generation, computer vision and multimodal pipelines that combine Tensor Core and graphics performance.

  • Rendering and 3D graphics

    Designed for interactive rendering, virtual production, product design, architecture, engineering and other RTX-accelerated visual workloads.

  • Video processing and streaming

    Three NVENC and three NVDEC engines, including AV1 support, make L40S useful for transcoding, streaming and AI-enhanced media workflows.

L40S is a strong fit when…

  • You need one GPU for AI, rendering and video workloads.
  • Your models and batches fit within 48 GB of GPU memory.
  • You need FP8 support for compatible AI inference or training.
  • You need RTX, ray tracing or virtual workstation capabilities.
  • You need multiple hardware video encode and decode engines.
  • You do not require NVLink or Multi-Instance GPU partitioning.

Consider another GPU when…

L40S may not be the best choice when your workload needs more than 48 GB, high FP64 throughput, direct GPU interconnects or MIG.

  • Compare H100 for larger AI training and high-end inference.
  • Compare H200 when memory capacity and bandwidth are the bottleneck.
  • Compare A100 for MIG, FP64 HPC and NVLink-based configurations.

Provider comparison

Cloud providers offering NVIDIA L40S

Compare the listed L40S instance resources, deployment region, starting price and billing options. Check whether video engines, RTX features, display functions and virtual GPU licensing are exposed in the selected cloud configuration.

View all providers
ProviderL40S offeringGPU memoryExample configurationRegionsStarting priceBilling optionsDetails
AceCloud1× NVIDIA L40S48 GB GDDR6 ECC16 vCPU · 64 GB RAMIndia and United StatesFrom ₹83,000/monthMonthly · 6/12-month plans · selected spotView provider
E2E Networks1× NVIDIA L40S48 GB GDDR6 ECC60 vCPU · 220 GB RAMIndia₹102/hour · ₹53,900/monthOn-demand · monthly · 6 months · annual · spot subject to capacityView provider
Cyfuture AI1× NVIDIA L40S48 GB GDDR616 vCPU · 256 GB RAMIndia, United States and Europe₹124/hour on-demand · ₹61/hour for 12 monthsOn-demand · 1/6/12-month reservedView provider
DigitalOcean1× NVIDIA L40S GPU Droplet48 GB GDDR68 vCPU · 64 GiB RAMToronto$1.57/GPU-hourPer-second on-demand · Spot not publicly listedView provider
AWSEC2 G6e48 GB GDDR6 per GPU1–8 GPUs · up to 192 vCPU and 1.5 TiB RAMSelected AWS regionsUse AWS Pricing CalculatorOn-Demand · Reserved Instances · Savings Plans · SpotView provider

Compare the complete instance rather than only the GPU price. Providers can differ in vCPU, RAM, storage, networking, region, commitment period, taxes, vGPU licensing and access to NVENC, NVDEC, RTX or display capabilities. Spot capacity can be interrupted and may not be available for every configuration.

NVIDIA L40S FAQs

How much GPU memory does NVIDIA L40S have?

L40S has 48 GB of GDDR6 memory with ECC and 864 GB/s of memory bandwidth.

Is NVIDIA L40S suitable for LLM inference?

Yes. L40S supports FP8 and provides 48 GB of memory, making it suitable for many LLM inference and selected fine-tuning workloads that fit within its memory capacity.

Can NVIDIA L40S be used for LLM training?

Yes, for models and training configurations that fit within 48 GB. Larger distributed training workloads may be better suited to GPUs with more memory and high-speed GPU interconnects.

Does NVIDIA L40S support MIG?

No. NVIDIA lists Multi-Instance GPU support as unavailable on L40S.

Does NVIDIA L40S support NVLink?

No. NVIDIA lists no NVLink support for L40S. Multi-GPU communication therefore uses the host system and PCIe rather than a direct NVLink connection.

Is L40S good for rendering and video?

Yes. L40S combines 142 third-generation RT Cores, 212 TFLOPS of RT performance, four DisplayPort 1.4a outputs and three NVENC plus three NVDEC engines.

Compare L40S configurations before choosing a provider

Compare the complete instance, not only the GPU name. Check CPU, RAM, storage, networking, region, vGPU licensing, video-engine access and whether the plan is optimized for AI, rendering or virtual workstations.