Skip to content

[Feature] DeepSeek V3 optimization #2591

Description

@zhyncs

Checklist

Adoption

SGLang adoption for DeepSeek V3 and R1

Usage

User Guide for Existing System (Installation & Launch)

https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3

Please use the latest version v0.4.2.post4. Please prefer to use docker image. docker pull lmsysorg/sglang:latest

For running on AMD MI300X, use this as a reference. Running DeepSeek-R1 on a single NDv5 MI300X VM

Features

More things (e.g., PD disaggregation, cache) are tracked at #4042

Activity

  1. pinned this issue on Dec 26, 2024
  2. libratiger commented on Dec 26, 2024

    @libratiger
    Contributor

    Very quick response !
    I understand that the overlap scheduler is model-independent and is a general optimization that should be supported by default.
    At least some special optimizations are needed?

  3. merrymercy commented on Dec 26, 2024

    @merrymercy
    Contributor

    The overlap scheduler is model-independent but has not been supported when using dp attention. We have a private branch for this and will upstream it soon.

  4. fengyang95 commented on Dec 26, 2024

    @fengyang95

    Is the memory sufficient for an 8 gpus instance? This model size is too large.

  5. zhyncs commented on Dec 26, 2024

    @zhyncs
    ContributorAuthor

    Is the memory sufficient for an 8 gpus instance? This model size is too large.

    671B works on H200 * 8 with FP8 (671 < 141 * 8)

  6. zhyncs commented on Dec 26, 2024

    @zhyncs
    ContributorAuthor

    Hi @fengyang95 You can also consider multi node.

    If you do not have GPUs with large enough memory, please try multi-node tensor parallelism (help 1 help 2).

  7. 35 remaining items

  8. zhyncs commented on Feb 17, 2025

    @zhyncs
    ContributorAuthor

    ref
    MTP support: #3582
    v0.4.3.post1 release: #3638

    SGLang supports MTP (nextn) in the Triton backend, achieving a speed of 77 tokens/s, twice as fast as other OSS LLM engines.

  9. ShivamB25 commented on Feb 17, 2025

    @ShivamB25

    @zhyncs can you please recommed the command for long prompt (lots of queries) since we are working with large documents. we want to run our own.
    . system has 8 x mi300x , 2TB ram and dual amd process epyc 9654 (96cores). if that helps.
    i am using this as of now

    ➜  ~ docker run -d \      
      --name sglang \
      --device=/dev/kfd \
      --device=/dev/dri \
      --security-opt seccomp=unconfined \
      --cap-add=SYS_PTRACE \
      --group-add video \
      --privileged \
      --shm-size 256g \
      --ipc=host \
      -p 3000:3000 \
      -v ~/.cache/huggingface:/root/.cache/huggingface \
      -e HSA_NO_SCRATCH_RECLAIM=1 \
      jesselopezmicrosoft/sglang:mi300x \
      python3 -m sglang.launch_server \
        --model unsloth/DeepSeek-R1 \
        --tp 8 \
        --trust-remote-code \
        --host 0.0.0.0 \
        --port 3000 \
        --enable-hierarchical-cache \
        --enable-cache-report \
        --chunked-prefill-size 8192 \
        --enable-metrics \
        --show-time-cost \
        --stream-output \
        --quantization fp8 \
        --attention-backend triton \
        --sampling-backend flashinfer \
        --grammar-backend xgrammar```
    
  10. xutizhou commented on Feb 19, 2025

    @xutizhou
    Collaborator

    @zhyncs Is there any plan for optimizing the FP8 GEMM kernel? Recently, I’ve been working on some optimizations for DeepSeekV3 and achieved promising results. I’d love to contribute to the community if possible. Let me know how I can help!

  11. YangQun1 commented on Feb 19, 2025

    @YangQun1
    Contributor

    @zhyncs Is there any plan to support ep moe for deepseek-v3/r1? there are related issues: #3371
    Ok, I find the feature issue. #2740

  12. zhyncs commented on Feb 19, 2025

    @zhyncs
    ContributorAuthor

    @zhyncs Is there any plan for optimizing the FP8 GEMM kernel? Recently, I’ve been working on some optimizations for DeepSeekV3 and achieved promising results. I’d love to contribute to the community if possible. Let me know how I can help!

    @xutizhou sgl-kernel already supports the CUTLASS block wise fp8 https://github.com/sgl-project/sglang/blob/main/sgl-kernel/src/sgl-kernel/csrc/fp8_blockwise_gemm_kernel.cu

    Please join the slack channel https://slack.sglang.ai

  13. zhyncs commented on Feb 19, 2025

    @zhyncs
    ContributorAuthor

    @zhyncs Is there any plan to support ep moe for deepseek-v3/r1? there are related issues: #3371 Ok, I find the feature issue. #2740

    ref #3602

  14. xutizhou commented on Feb 20, 2025

    @xutizhou
    Collaborator

    @zhyncs Is there any plan for optimizing the FP8 GEMM kernel? Recently, I’ve been working on some optimizations for DeepSeekV3 and achieved promising results. I’d love to contribute to the community if possible. Let me know how I can help!

    @xutizhou sgl-kernel already supports the CUTLASS block wise fp8 https://github.com/sgl-project/sglang/blob/main/sgl-kernel/src/sgl-kernel/csrc/fp8_blockwise_gemm_kernel.cu

    Please join the slack channel https://slack.sglang.ai

    Thank you for the reference provided.

  15. zqqQiu commented on Feb 21, 2025

    @zqqQiu

    thks for reply,I solved this problem by check model file compeletion (some safetensor missing , but sglang or vllm don't report error ) @yinfan98 @zhyncs

    Can you share the command on 8*h20?

  16. unpinned this issue on Feb 21, 2025
  17. YangZeyu95 commented on Feb 24, 2025

    @YangZeyu95

    Hi, is there any plan to support pipeline parallelism? As I don't have IB for connection between 2 servers, pp will be very helpful for performance

  18. tutu329 commented on Mar 21, 2025

    @tutu329

    Hi, is there any plan to support pipeline parallelism? As I don't have IB for connection between 2 servers, pp will be very helpful for performance

    the same problem

  19. groklab commented on Mar 24, 2025

    @groklab

    Hi SGLang Team - should we expect a smooth transition to deepseek v3 0324 checkpoint (without any changes) or do you need some backend changes before we can pull the 0324 weights?

    https://huggingface.co/deepseek-ai/DeepSeek-V3-0324

  20. zhyncs commented on Mar 25, 2025

    @zhyncs
    ContributorAuthor

    Hi SGLang Team - should we expect a smooth transition to deepseek v3 0324 checkpoint (without any changes) or do you need some backend changes before we can pull the 0324 weights?

    https://huggingface.co/deepseek-ai/DeepSeek-V3-0324

    Hi @groklab Thanks to let me know. I will upload the NextN weights for this new checkpoint.

  21. parambole commented on Jul 9, 2025

    @parambole

    Hey @zhyncs I am currently working on AI-Hypercomputer/maxtext#1837 Deepseek's Multi-Token Prediction into Maxtext.

    Question:

    As part of the MTP implementation, has the team been able to load the open MTP module Deepseek V3 weights and analyze the implementation during pre-training and fine-tuning? I am curious to know about any observed behavior?

    Specifically I am observing: deepseek-ai/DeepSeek-V3#928

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions