Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
gpu
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
From Naive CUDA to Performance Engineering: My First GPU Matmul Journey
Karthik Unnikrishnan
Karthik Unnikrishnan
Karthik Unnikrishnan
Follow
Sep 29
From Naive CUDA to Performance Engineering: My First GPU Matmul Journey
#
ai
#
gpu
#
amd
#
programming
Comments
Add Comment
3 min read
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
AI Explore
AI Explore
AI Explore
Follow
Sep 29
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer
#
gpu
#
performance
#
apple
#
datascience
Comments
Add Comment
10 min read
Before You Rent a GPU Server, Check These 10 Things
Arthur
Arthur
Arthur
Follow
Sep 29
Before You Rent a GPU Server, Check These 10 Things
#
webdev
#
gpu
#
devops
#
machinelearning
Comments
Add Comment
3 min read
Distributed Training & Inference: From CPUs and GPUs to a Cluster
Aleksei Romanov
Aleksei Romanov
Aleksei Romanov
Follow
for
g factor
Sep 29
Distributed Training & Inference: From CPUs and GPUs to a Cluster
#
machinelearning
#
pytorch
#
gpu
#
distributedsystems
1
 reaction
Comments
Add Comment
14 min read
When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM
orca_forge
orca_forge
orca_forge
Follow
Sep 29
When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM
#
gpu
#
vram
#
claudecode
Comments
Add Comment
3 min read
How to Use Your NVIDIA GPU for Local AI on Linux in 2026
Mohammed Ali Chherawalla
Mohammed Ali Chherawalla
Mohammed Ali Chherawalla
Follow
Sep 29
How to Use Your NVIDIA GPU for Local AI on Linux in 2026
#
ai
#
linux
#
gpu
#
tutorial
Comments
Add Comment
5 min read
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
Donald Lee
Donald Lee
Donald Lee
Follow
Sep 29
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs
#
llm
#
llamacpp
#
localai
#
gpu
1
 reaction
Comments
1
 comment
3 min read
Docker Model Runner vs Ollama in 2026: Workflow Trade-offs and Benchmarking Limits
Jangwook Kim
Jangwook Kim
Jangwook Kim
Follow
Sep 29
Docker Model Runner vs Ollama in 2026: Workflow Trade-offs and Benchmarking Limits
#
docker
#
ollama
#
localllm
#
gpu
Comments
2
 comments
9 min read
What it costs to rent an H100, B200 or RTX 4090 in September 2026: live prices from 28 GPU clouds
FastGPU
FastGPU
FastGPU
Follow
Sep 27
What it costs to rent an H100, B200 or RTX 4090 in September 2026: live prices from 28 GPU clouds
#
ai
#
machinelearning
#
cloud
#
gpu
Comments
1
 comment
7 min read
AsyncGRPO: Eliminating GPU Idle Bubbles in Environment-Heavy RL Post-Training
Aleksei Romanov
Aleksei Romanov
Aleksei Romanov
Follow
for
g factor
Sep 27
AsyncGRPO: Eliminating GPU Idle Bubbles in Environment-Heavy RL Post-Training
#
machinelearning
#
reinforcementlearning
#
llm
#
gpu
Comments
1
 comment
9 min read
Same Source, Two Cost Curves: Local Preview vs GPU Final
Nidheeshdas Thavorath
Nidheeshdas Thavorath
Nidheeshdas Thavorath
Follow
Sep 23
Same Source, Two Cost Curves: Local Preview vs GPU Final
#
video
#
gpu
#
performance
#
devops
Comments
Add Comment
5 min read
Let your AI agent rent a GPU: llms.txt, --json and --budget
Lium
Lium
Lium
Follow
for
Lium
Sep 24
Let your AI agent rent a GPU: llms.txt, --json and --budget
#
ai
#
agents
#
gpu
#
devops
1
 reaction
Comments
1
 comment
4 min read
Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI
[email protected]
[email protected]
[email protected]
Follow
Sep 22
Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI
#
llm
#
platformengineering
#
infrastructure
#
gpu
Comments
Add Comment
1 min read
Why Does Your Local Model Crash at 32k Tokens?
Eryk Kubiak
Eryk Kubiak
Eryk Kubiak
Follow
Sep 25
Why Does Your Local Model Crash at 32k Tokens?
#
llm
#
gpu
#
quantization
#
machinelearning
1
 reaction
Comments
2
 comments
12 min read
How Much VRAM Do You Really Need to Run a 70B LLM?
Peter Gedeon
Peter Gedeon
Peter Gedeon
Follow
Sep 23
How Much VRAM Do You Really Need to Run a 70B LLM?
#
ai
#
llm
#
gpu
#
ram
Comments
Add Comment
7 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account