Repository navigation
[Feature] DeepSeek V3 optimization #2591
Description
Activity
- addedenhancementNew feature or requestNew feature or requestquantLLM QuantizationLLM Quantization
on Dec 26, 2024 - pinned this issue
on Dec 26, 2024 Very quick response !
I understand that the overlap scheduler is model-independent and is a general optimization that should be supported by default.
At least some special optimizations are needed?The overlap scheduler is model-independent but has not been supported when using dp attention. We have a private branch for this and will upstream it soon.
Reacted by libra, agiping, Diogo and Jueon ParkIs the memory sufficient for an 8 gpus instance? This model size is too large.
Is the memory sufficient for an 8 gpus instance? This model size is too large.
671B works on H200 * 8 with FP8 (671 < 141 * 8)
Reacted by zengqingfu1442, 懒人旭, ZGP and Yiqian-LiuHi @fengyang95 You can also consider multi node.
If you do not have GPUs with large enough memory, please try multi-node tensor parallelism (help 1 help 2).
Reacted by Yuanpeng Li35 remaining items
@zhyncs can you please recommed the command for long prompt (lots of queries) since we are working with large documents. we want to run our own.
. system has 8 x mi300x , 2TB ram and dual amd process epyc 9654 (96cores). if that helps.
i am using this as of now➜ ~ docker run -d \ --name sglang \ --device=/dev/kfd \ --device=/dev/dri \ --security-opt seccomp=unconfined \ --cap-add=SYS_PTRACE \ --group-add video \ --privileged \ --shm-size 256g \ --ipc=host \ -p 3000:3000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ -e HSA_NO_SCRATCH_RECLAIM=1 \ jesselopezmicrosoft/sglang:mi300x \ python3 -m sglang.launch_server \ --model unsloth/DeepSeek-R1 \ --tp 8 \ --trust-remote-code \ --host 0.0.0.0 \ --port 3000 \ --enable-hierarchical-cache \ --enable-cache-report \ --chunked-prefill-size 8192 \ --enable-metrics \ --show-time-cost \ --stream-output \ --quantization fp8 \ --attention-backend triton \ --sampling-backend flashinfer \ --grammar-backend xgrammar```@zhyncs Is there any plan for optimizing the FP8 GEMM kernel? Recently, I’ve been working on some optimizations for DeepSeekV3 and achieved promising results. I’d love to contribute to the community if possible. Let me know how I can help!
@zhyncs Is there any plan for optimizing the FP8 GEMM kernel? Recently, I’ve been working on some optimizations for DeepSeekV3 and achieved promising results. I’d love to contribute to the community if possible. Let me know how I can help!
@xutizhou sgl-kernel already supports the CUTLASS block wise fp8 https://github.com/sgl-project/sglang/blob/main/sgl-kernel/src/sgl-kernel/csrc/fp8_blockwise_gemm_kernel.cu
Please join the slack channel https://slack.sglang.ai
@zhyncs Is there any plan for optimizing the FP8 GEMM kernel? Recently, I’ve been working on some optimizations for DeepSeekV3 and achieved promising results. I’d love to contribute to the community if possible. Let me know how I can help!
@xutizhou sgl-kernel already supports the CUTLASS block wise fp8 https://github.com/sgl-project/sglang/blob/main/sgl-kernel/src/sgl-kernel/csrc/fp8_blockwise_gemm_kernel.cu
Please join the slack channel https://slack.sglang.ai
Thank you for the reference provided.
- unpinned this issue
on Feb 21, 2025 Hi, is there any plan to support pipeline parallelism? As I don't have IB for connection between 2 servers, pp will be very helpful for performance
Hi, is there any plan to support pipeline parallelism? As I don't have IB for connection between 2 servers, pp will be very helpful for performance
the same problem
Hi SGLang Team - should we expect a smooth transition to deepseek v3 0324 checkpoint (without any changes) or do you need some backend changes before we can pull the 0324 weights?
Hi SGLang Team - should we expect a smooth transition to deepseek v3 0324 checkpoint (without any changes) or do you need some backend changes before we can pull the 0324 weights?
Hi @groklab Thanks to let me know. I will upload the NextN weights for this new checkpoint.
Hey @zhyncs I am currently working on AI-Hypercomputer/maxtext#1837 Deepseek's Multi-Token Prediction into Maxtext.
Question:
As part of the MTP implementation, has the team been able to load the open MTP module Deepseek V3 weights and analyze the implementation during pre-training and fine-tuning? I am curious to know about any observed behavior?
Specifically I am observing: deepseek-ai/DeepSeek-V3#928
Checklist
Adoption
SGLang adoption for DeepSeek V3 and R1
Usage
User Guide for Existing System (Installation & Launch)
https://github.com/sgl-project/sglang/tree/main/benchmark/deepseek_v3
Please use the latest version v0.4.2.post4. Please prefer to use docker image.
docker pull lmsysorg/sglang:latestFor running on AMD MI300X, use this as a reference. Running DeepSeek-R1 on a single NDv5 MI300X VM
Features
moe_align_block_size@HandH1998 @zhyncs @BBufE=256,N=256,device_name=NVIDIA_H200,dtype=fp8_w8a8.json@BBufnextnspeculative decoding @ispobock [Track] DeepSeek V3/R1 nextn progress #3472More things (e.g., PD disaggregation, cache) are tracked at #4042