Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
gpu
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Same Source, Two Cost Curves: Local Preview vs GPU Final
Nidheeshdas Thavorath
Nidheeshdas Thavorath
Nidheeshdas Thavorath
Follow
Sep 23
Same Source, Two Cost Curves: Local Preview vs GPU Final
#
video
#
gpu
#
performance
#
devops
Comments
Add Comment
5 min read
Let your AI agent rent a GPU: llms.txt, --json and --budget
Lium
Lium
Lium
Follow
for
Lium
Sep 24
Let your AI agent rent a GPU: llms.txt, --json and --budget
#
ai
#
agents
#
gpu
#
devops
1
 reaction
Comments
1
 comment
4 min read
Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI
ram.mehta.1899@gmail.com
ram.mehta.1899@gmail.com
ram.mehta.1899@gmail.com
Follow
Sep 22
Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI
#
llm
#
platformengineering
#
infrastructure
#
gpu
Comments
Add Comment
1 min read
How Much VRAM Do You Really Need to Run a 70B LLM?
Peter Gedeon
Peter Gedeon
Peter Gedeon
Follow
Sep 23
How Much VRAM Do You Really Need to Run a 70B LLM?
#
ai
#
llm
#
gpu
#
ram
Comments
Add Comment
7 min read
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
Fyodor Serzhenko
Fyodor Serzhenko
Fyodor Serzhenko
Follow
Sep 18
Same nvJPEG2000, different numbers: timer boundaries and frames in flight
#
cuda
#
gpu
#
performance
#
cpp
Comments
Add Comment
13 min read
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
ComradePenguin
ComradePenguin
ComradePenguin
Follow
Sep 18
Exploring Result Visibility of Fixed-Latency Instructions on the SM120 Architecture
#
cuda
#
nvidia
#
gpu
1
 reaction
Comments
Add Comment
10 min read
Why GPU Availability Is Still the Biggest Bottleneck in ML Infra
Sumukh Shenoy
Sumukh Shenoy
Sumukh Shenoy
Follow
Sep 22
Why GPU Availability Is Still the Biggest Bottleneck in ML Infra
#
machinelearning
#
infrastructure
#
gpu
#
devops
1
 reaction
Comments
1
 comment
3 min read
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)
Sho Tanaka (tsho)
Sho Tanaka (tsho)
Sho Tanaka (tsho)
Follow
Sep 17
DeepSpeed ZeRO vs PyTorch FSDP2, Measured: 1.8x on NVLink, Barely Finishing on PCIe (Plus How to Actually Get GPUs on GKE)
#
machinelearning
#
pytorch
#
deepspeed
#
gpu
Comments
Add Comment
10 min read
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix
Yehor Cherednichenko
Yehor Cherednichenko
Yehor Cherednichenko
Follow
Sep 17
Half of my FP8 convolutions were silently running in f32 – a one-line XLA fix
#
machinelearning
#
performance
#
compilers
#
gpu
2
 reactions
Comments
1
 comment
6 min read
What Happens When You Ask an LLM a Question
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Vishnu Hari Dadhich
Follow
Sep 14
What Happens When You Ask an LLM a Question
#
ai
#
localllm
#
llmbasics
#
gpu
Comments
Add Comment
8 min read
CUDA Cores vs Tensor Cores Explained
Sumukh Shenoy
Sumukh Shenoy
Sumukh Shenoy
Follow
Sep 14
CUDA Cores vs Tensor Cores Explained
#
gpu
#
machinelearning
#
deeplearning
#
beginners
Comments
1
 comment
2 min read
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators
stmanst
stmanst
stmanst
Follow
Sep 9
Inside SFPU Overflow Bugs: How a 40-Year-Old Rounding Trick Breaks on Modern AI Accelerators
#
ai
#
machinelearning
#
gpu
#
debugging
Comments
Add Comment
3 min read
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill
erniou86
erniou86
erniou86
Follow
Sep 9
I Put a Paid AI Video Generator on My Own Gaming GPU — No Cloud Bill
#
ai
#
gpu
#
selfhosting
#
machinelearning
Comments
1
 comment
3 min read
Reverse-Engineering NVIDIA: Modifying a CUDA binary
Stjepan
Stjepan
Stjepan
Follow
Sep 6
Reverse-Engineering NVIDIA: Modifying a CUDA binary
#
nvidia
#
gpu
#
cuda
#
hex
Comments
Add Comment
4 min read
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Dinesh Kumar Ramasamy
Follow
Sep 5
From API to GPU, Week 6 (Part 1): A Model That Predicts, and How Wrong It Is
#
ai
#
llm
#
gpu
#
machinelearning
Comments
Add Comment
13 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account