Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
cuda
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
I wrote 21 CUDA kernels on a free GPU. Here's what the profiler taught me.
Musharaf Khan Pathan
Musharaf Khan Pathan
Musharaf Khan Pathan
Follow
Oct 10
I wrote 21 CUDA kernels on a free GPU. Here's what the profiler taught me.
#
cuda
#
gpu
#
performance
#
cpp
Comments
Add Comment
5 min read
Running vLLM natively on Windows
aivrar
aivrar
aivrar
Follow
Oct 9
Running vLLM natively on Windows
#
vllm
#
windows
#
cuda
#
llm
Comments
Add Comment
6 min read
Running the 510 GB DeepSeek-V4.1-Flash on an 8 GB GPU — and three bugs that never raise an error
Helgard
Helgard
Helgard
Follow
Oct 3
Running the 510 GB DeepSeek-V4.1-Flash on an 8 GB GPU — and three bugs that never raise an error
#
ai
#
llm
#
cuda
#
python
Comments
Add Comment
8 min read
Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
Helgard
Helgard
Helgard
Follow
Oct 3
Running a 133 GB MoE model on an 8 GB GPU at 11 tokens/s by streaming experts from NVMe
#
ai
#
llm
#
cuda
#
python
Comments
Add Comment
6 min read
How vLLM CUDA Kernels Write and Read the Paged KV Cache
yuan lei
yuan lei
yuan lei
Follow
Oct 9
How vLLM CUDA Kernels Write and Read the Paged KV Cache
#
vllm
#
cuda
#
llm
#
ai
Comments
1
 comment
5 min read
How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs
Bhushan Kinge
Bhushan Kinge
Bhushan Kinge
Follow
Sep 24
How Fast Can a 421M-Parameter Decision Model Run? I Benchmarked Laya Across NVIDIA GPUs
#
machinelearning
#
cuda
#
performance
#
opensource
Comments
1
 comment
6 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
ComradePenguin
ComradePenguin
ComradePenguin
Follow
Sep 18
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
#
cuda
#
nvidia
#
gpu
1
 reaction
Comments
Add Comment
10 min read
Nvidia Just Let Rust Into CUDA. Here's Why That's a Bigger Deal Than It Sounds
Ashraf
Ashraf
Ashraf
Follow
Sep 18
Nvidia Just Let Rust Into CUDA. Here's Why That's a Bigger Deal Than It Sounds
#
rust
#
ai
#
cuda
#
programming
1
 reaction
Comments
Add Comment
5 min read
CUDA Rust: two native tracks, not a wrapper over C++
Juan Torchia
Juan Torchia
Juan Torchia
Follow
Sep 17
CUDA Rust: two native tracks, not a wrapper over C++
#
english
#
rust
#
cuda
#
gpuprogramming
1
 reaction
Comments
Add Comment
5 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 23
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x
#
gemma
#
llamacpp
#
cuda
#
benchmarking
10
 reactions
Comments
2
 comments
11 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary
Stjepan
Stjepan
Stjepan
Follow
Sep 6
Reverse-Engineering NVIDIA: Modifying a CUDA binary
#
nvidia
#
gpu
#
cuda
#
hex
Comments
Add Comment
4 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 22
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It
#
gemma
#
vllm
#
gcp
#
cuda
13
 reactions
Comments
Add Comment
13 min read
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step
xbill
xbill
xbill
Follow
for
Google Developer Experts
Oct 2
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step
#
gemma
#
llamacpp
#
cuda
#
huggingface
9
 reactions
Comments
2
 comments
12 min read
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16
xbill
xbill
xbill
Follow
for
Google Developer Experts
Sep 30
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16
#
gemma
#
vllm
#
cuda
#
machinelearning
9
 reactions
Comments
2
 comments
11 min read
Finding a Random Island with Geometry and CUDA
Muhammad Adil
Muhammad Adil
Muhammad Adil
Follow
Aug 19
Finding a Random Island with Geometry and CUDA
#
gpucomputing
#
geospatial
#
cuda
#
geometry
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account