DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

Kaggle Benchmarking Challenge Submission

Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

13
Comments 4
10 min read
1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

Kaggle Benchmarking Challenge Submission

1:30 AM Happens Twice on November 1. Most AI Models Picked One and Moved On.

Comments 1
5 min read
Only Gemini Failed My False-Premise Benchmark — 7 Models Tested

Kaggle Benchmarking Challenge Submission

Only Gemini Failed My False-Premise Benchmark — 7 Models Tested

Comments
4 min read
ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM

Kaggle Benchmarking Challenge Submission

ProofSec: Benchmarking Epistemic Robustness and Evidence-Grounded Vulnerability Reasoning in Frontier LLM

Comments
14 min read
Hindsight-Free Temporal Evidence Eligibility

Kaggle Benchmarking Challenge Submission

Hindsight-Free Temporal Evidence Eligibility

Comments
4 min read
I Spun a Wheel of Fortune at 13 AI Models. Here's Who Took the Bait.

Kaggle Benchmarking Challenge Submission

I Spun a Wheel of Fortune at 13 AI Models. Here's Who Took the Bait.

Comments 1
13 min read
It knows you changed jobs. It still writes to your old manager.

Kaggle Benchmarking Challenge Submission

It knows you changed jobs. It still writes to your old manager.

Comments 1
8 min read
I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.

Kaggle Benchmarking Challenge Submission

I Benchmarked Whether AI Models Forget Corrections. The Answer Surprised Me.

Comments
4 min read
Building a C++ Logical Bug Detection Benchmark on Kaggle: Testing 3 AI Models

Building a C++ Logical Bug Detection Benchmark on Kaggle: Testing 3 AI Models

Comments 2
2 min read
Would you ship this Italian? A benchmark built from real website bugs.

Kaggle Benchmarking Challenge Submission

Would you ship this Italian? A benchmark built from real website bugs.

Comments
4 min read
I Tested 5 AI Models on Hinglish — Here's Who Won

I Tested 5 AI Models on Hinglish — Here's Who Won

3
Comments 1
1 min read
AI models catch bad code, then cry wolf on the good code

Kaggle Benchmarking Challenge Submission

AI models catch bad code, then cry wolf on the good code

Comments
5 min read
Do LLMs Actually Fix Tricky React Hooks, or Do They Just Cheat?

Do LLMs Actually Fix Tricky React Hooks, or Do They Just Cheat?

1
Comments
3 min read
A Paused AI Workflow: Retry, Resume, or Keep Holding?

Kaggle Benchmarking Challenge Submission

A Paused AI Workflow: Retry, Resume, or Keep Holding?

Comments
4 min read
MY Edit or Abstain: AI Code-Repair Decision Benchmark

MY Edit or Abstain: AI Code-Repair Decision Benchmark

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.