DEV Community

#gpu

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
From Naive CUDA to Performance Engineering: My First GPU Matmul Journey

From Naive CUDA to Performance Engineering: My First GPU Matmul Journey

Comments
3 min read
Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Your Apple Silicon GPU Loses to One CPU Core Until a Million Rows. I Measured 111 Operations, Then Rebuilt ArrowMetal 0.2.0 Around the Answer

Comments
10 min read
Before You Rent a GPU Server, Check These 10 Things

Before You Rent a GPU Server, Check These 10 Things

Comments
3 min read
Distributed Training & Inference: From CPUs and GPUs to a Cluster

Distributed Training & Inference: From CPUs and GPUs to a Cluster

1
Comments
14 min read
When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM

When GPU Resources Run Dry, It Still Looks Like Everything Is Working: A Tale of Two AI Competing for VRAM

Comments
3 min read
How to Use Your NVIDIA GPU for Local AI on Linux in 2026

How to Use Your NVIDIA GPU for Local AI on Linux in 2026

Comments
5 min read
How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

How to pick --n-cpu-moe in llama.cpp: Qwen3.6 35B-A3B on 12, 16 and 24 GB GPUs

1
Comments 1
3 min read
Docker Model Runner vs Ollama in 2026: Workflow Trade-offs and Benchmarking Limits

Docker Model Runner vs Ollama in 2026: Workflow Trade-offs and Benchmarking Limits

Comments 2
9 min read
What it costs to rent an H100, B200 or RTX 4090 in September 2026: live prices from 28 GPU clouds

What it costs to rent an H100, B200 or RTX 4090 in September 2026: live prices from 28 GPU clouds

Comments 1
7 min read
AsyncGRPO: Eliminating GPU Idle Bubbles in Environment-Heavy RL Post-Training

AsyncGRPO: Eliminating GPU Idle Bubbles in Environment-Heavy RL Post-Training

Comments 1
9 min read
Same Source, Two Cost Curves: Local Preview vs GPU Final

Same Source, Two Cost Curves: Local Preview vs GPU Final

Comments
5 min read
Let your AI agent rent a GPU: llms.txt, --json and --budget

Let your AI agent rent a GPU: llms.txt, --json and --budget

1
Comments 1
4 min read
Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI

Evaluating Multi-Node LLM Orchestrators: Exo, GPUStack, and LocalAI

Comments
1 min read
Why Does Your Local Model Crash at 32k Tokens?

Why Does Your Local Model Crash at 32k Tokens?

1
Comments 2
12 min read
How Much VRAM Do You Really Need to Run a 70B LLM?

How Much VRAM Do You Really Need to Run a 70B LLM?

Comments
7 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.