DEV Community

#benchmark

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
We recompute TypeSafe's 444x claim — here's what we found

We recompute TypeSafe's 444x claim — here's what we found

Comments
1 min read
Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First

Verification as Protocol: We Test AI Agents' Memory — Our Grader Failed First

Comments
1 min read
PQC-Bench Part 1 — measuring X25519MLKEM768 costs on a $0.04/hour Graviton3

PQC-Bench Part 1 — measuring X25519MLKEM768 costs on a $0.04/hour Graviton3

Comments 1
5 min read
One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.

One Hung API Call Used to Kill My 1,000-Run Benchmark. Here's the Fix.

7
Comments
3 min read
Claude Opus 5.5: 40% cheaper, frontier-grade performance

Claude Opus 5.5: 40% cheaper, frontier-grade performance

Comments
5 min read
Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

Glasshouse v0.1 Is Out: A Memory Benchmark for AI Systems

7
Comments 1
2 min read
Onboarding benchmarks from real data across 464 SaaS products: median tour completion is 29%, 1-2 step tours complete at 73%, 9+ step tours at 8%

Onboarding benchmarks from real data across 464 SaaS products: median tour completion is 29%, 1-2 step tours complete at 73%, 9+ step tours at 8%

Comments 1
2 min read
netcup VPS 1000 G12 benchmarked: how fast is it really?

netcup VPS 1000 G12 benchmarked: how fast is it really?

Comments
7 min read
Enterprise Vector Database 2026: Qdrant vs Milvus vs pgvector vs Pinecone

Enterprise Vector Database 2026: Qdrant vs Milvus vs pgvector vs Pinecone

Comments
23 min read
TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper?

TypeSafe’s JEV Model: Is It Really 193x Faster and 444x Cheaper?

2
Comments 1
6 min read
I Made a Memory Benchmark as Fair as I Could. 15 Stars, 0 Runs. What's Wrong With It?

I Made a Memory Benchmark as Fair as I Could. 15 Stars, 0 Runs. What's Wrong With It?

7
Comments 4
2 min read
JetBrains Ranked AI Agents on Real Kotlin Projects. The Token Column Is the Real Story.

JetBrains Ranked AI Agents on Real Kotlin Projects. The Token Column Is the Real Story.

1
Comments
7 min read
AI Agent Evaluation End to End: How the LFORLA Drone Build Benchmark Scores Planning, Sourcing, and Assembly

AI Agent Evaluation End to End: How the LFORLA Drone Build Benchmark Scores Planning, Sourcing, and Assembly

2
Comments 2
4 min read
How Fast Can Neovim Start? Benchmarking Popular Distros

How Fast Can Neovim Start? Benchmarking Popular Distros

Comments
6 min read
Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)

Why I’m Building a New AI Memory Benchmark (And Why the Existing Ones Fall Short)

2
Comments
2 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.