~VRAM-calculator
MODEL7B MODEINFERENCE PRECISION16-BIT FIT24 GB

AI Deployment Calculator

Estimate the GPU VRAM and hardware tier needed to deploy an AI model's workload.

Presets
Model

Deployment

Advanced assumptions

On-disk weight size in GB. Overrides the parameter-based weight estimate when set.

Fraction (0–1, not a percentage) of the known file kept in VRAM. Only applies when Known Model File Size is set.

Estimated VRAM Required

18.8 GB

Recommended Example

24 GB hardware tier

e.g. RTX 4090

Usage Capacity

    Why this recommendation

    • Minimum GPU VRAM Capacity
    • Usable VRAM Target
    • Usable VRAM on Recommended Hardware Tier
    • Fit Headroom
    • Estimated Speed
    Values Used In Calculations
      Formula used

      Assumptions used
        How VRAM is calculated
        Required_GB = (Weights + Working_Memory + Training_State + Runtime_Overhead) x Buffer

        This LLM VRAM calculator estimates model weights, working memory for context or media activations, training state, runtime overhead, and a safety buffer before selecting a hardware tier. Use it to work out how much VRAM a 70B model needs, the GPU memory for LoRA or QLoRA fine-tuning, or the requirements for full training, multimodal models, diffusion, video, audio, tabular, and custom workloads.

        How much VRAM do I need for a 70B model?

        At default inference settings a 70B model needs about 160.8 GB at FP16, 87.7 GB at 8-bit, or 51.1 GB at 4-bit. Context length, batch size, and runtime overhead move that number; the table below lists common model sizes at each precision.

        Quick VRAM reference: GPU memory per model size and precision
        Model FP16 8-bit 4-bit
        7B 18.8 GB 11.5 GB 7.8 GB
        13B 32.3 GB 18.7 GB 11.9 GB
        70B 161.1 GB 88.0 GB 51.4 GB

        These rows use the same default text-generation inference inputs as the calculator above: server/cloud runtime, 8k context, one concurrent request, and no manual model file size override.

          Hardware tier

          single accelerator

          ≤ 24 Common consumer / small inference
          32–48 High-end consumer, workstation, larger local inference
          64–96 Large local systems and datacenter accelerators
          141–192 High-memory datacenter / frontier accelerators
          > 192 Does not fit on a single accelerator
          • 160 GB 2x 80 GB GPUs with tensor/model parallelism
          • 320 GB 4x 80 GB GPUs with tensor/model parallelism