DEV Community

#benchmarking

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
We Wanted a Fast Offline AI. First, We Had to Fix Our Benchmark.

We Wanted a Fast Offline AI. First, We Had to Fix Our Benchmark.

Comments
22 min read
I Built a Benchmark That Catches AI Models Cheating (And They All Failed)

I Built a Benchmark That Catches AI Models Cheating (And They All Failed)

Comments
3 min read
Noise Sources: Affinity, Frequency, and Background Load — How to isolate and quantify non‑compiler variance with csperf

Noise Sources: Affinity, Frequency, and Background Load — How to isolate and quantify non‑compiler variance with csperf

Comments
3 min read
Kev gives the same answer every time. Until you batch it.

Kev gives the same answer every time. Until you batch it.

Comments
8 min read
I Tested Whether Frontier AI Can Do a Texas Real Estate Agent's Desk Work — Here's What 3 Models Taught Me (and What the 4th Taught Me About Benchmarks)

Kaggle Benchmarking Challenge Submission

I Tested Whether Frontier AI Can Do a Texas Real Estate Agent's Desk Work — Here's What 3 Models Taught Me (and What the 4th Taught Me About Benchmarks)

Comments
6 min read
Warmup Deep Dive: What You’re Throwing Away and Why

Warmup Deep Dive: What You’re Throwing Away and Why

Comments
4 min read
Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens

Stability Metrics: Min, Max, and Standard Deviation as First-Class Citizens

Comments
3 min read
Benchmarking regex patterns properly: P95/P99, equivalence checks, real engines

Benchmarking regex patterns properly: P95/P99, equivalence checks, real engines

Comments
2 min read
MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Kaggle Benchmarking Challenge Submission

MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?

Comments
3 min read
LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement

LiteLLM Rust Gateway Benchmarked: Fast and Tiny, but Not Yet a Python Proxy Replacement

Comments
11 min read
I benchmarked what frontier models actually know about 2026 — most of it, they don't

Kaggle Benchmarking Challenge Submission

I benchmarked what frontier models actually know about 2026 — most of it, they don't

Comments
3 min read
Measuring the Rust Storage Engine: Lioran S3 PUT Timing from Network to RocksDB

Measuring the Rust Storage Engine: Lioran S3 PUT Timing from Network to RocksDB

1
Comments
2 min read
Kaniko's Performance Gap Narrows: Outdated 2018 Benchmarks Mislead, Recent Updates Close BuildKit Speed Divide

Kaniko's Performance Gap Narrows: Outdated 2018 Benchmarks Mislead, Recent Updates Close BuildKit Speed Divide

Comments
13 min read
Jev is the best decision model. Here's what to run when you can't use it.

Jev is the best decision model. Here's what to run when you can't use it.

Comments
10 min read
NeMo Guardrails vs Guardrails AI: The Production Latency Benchmark We Could Not Honestly Complete

NeMo Guardrails vs Guardrails AI: The Production Latency Benchmark We Could Not Honestly Complete

Comments
11 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.