DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
오픈 퀀텀 챌린지(OQC): 재현 가능한 하드웨어 프리 양자 벤치마크

오픈 퀀텀 챌린지(OQC): 재현 가능한 하드웨어 프리 양자 벤치마크

Comments
1 min read
Open Quantum Challenge (OQC): a reproducible, hardware-free quantum benchmark

Open Quantum Challenge (OQC): a reproducible, hardware-free quantum benchmark

Comments
1 min read
Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions

Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions

Comments
2 min read
How LLM Evaluation Actually Works: Inside a Benchmark That Produces Comparable Numbers

How LLM Evaluation Actually Works: Inside a Benchmark That Produces Comparable Numbers

Comments
4 min read
HumanEval Passes. Production Burns. The Real Story of AI Code Generation in the Vibe Coding Era

HumanEval Passes. Production Burns. The Real Story of AI Code Generation in the Vibe Coding Era

Comments
8 min read
I benchmarked seven hotel price APIs on the same Rome room, and the prices were 16% apart because of tax

I benchmarked seven hotel price APIs on the same Rome room, and the prices were 16% apart because of tax

Comments
4 min read
I benchmarked four Google Flights APIs on the same searches, and the fares matched to the dollar

I benchmarked four Google Flights APIs on the same searches, and the fares matched to the dollar

Comments
4 min read
What Independent Benchmarks Say About Opus 5.5

What Independent Benchmarks Say About Opus 5.5

Comments
8 min read
LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers

LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers

Comments
4 min read
Parakeet Redux: is a 178 MB speech model useful in the real world?

Parakeet Redux: is a 178 MB speech model useful in the real world?

Comments
16 min read
How we use Jev to answer quiz questions (a 90-question benchmark)

How we use Jev to answer quiz questions (a 90-question benchmark)

Comments
3 min read
When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician

When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician

Comments
7 min read
How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides

How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides

Comments 1
10 min read
Half the MCP servers that answer you don't actually work

Half the MCP servers that answer you don't actually work

Comments 2
8 min read
Hugging Face now has 48 official benchmarks. Here is what the map looks like

Hugging Face now has 48 official benchmarks. Here is what the map looks like

Comments
3 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.