Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmark
Follow
Hide
Posts
Left menu
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
오픈 퀀텀 챌린지(OQC): 재현 가능한 하드웨어 프리 양자 벤치마크
hugginf_expert
hugginf_expert
hugginf_expert
Follow
Oct 11
오픈 퀀텀 챌린지(OQC): 재현 가능한 하드웨어 프리 양자 벤치마크
#
quantum
#
benchmark
#
opensource
#
ai
Comments
Add Comment
1 min read
Open Quantum Challenge (OQC): a reproducible, hardware-free quantum benchmark
GINIGEN AI
GINIGEN AI
GINIGEN AI
Follow
Oct 11
Open Quantum Challenge (OQC): a reproducible, hardware-free quantum benchmark
#
quantum
#
benchmark
#
opensource
#
machinelearning
Comments
Add Comment
1 min read
Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions
QuantID
QuantID
QuantID
Follow
Oct 11
Eleven Number-One Records: Measuring a Model That Swept Math, Science, Law and Decisions
#
datascience
#
benchmark
#
llm
#
statistics
Comments
Add Comment
2 min read
How LLM Evaluation Actually Works: Inside a Benchmark That Produces Comparable Numbers
RESK
RESK
RESK
Follow
Oct 10
How LLM Evaluation Actually Works: Inside a Benchmark That Produces Comparable Numbers
#
llm
#
evaluation
#
benchmark
#
ai
Comments
Add Comment
4 min read
HumanEval Passes. Production Burns. The Real Story of AI Code Generation in the Vibe Coding Era
Emmanuel R
Emmanuel R
Emmanuel R
Follow
for
CobuildX AI
Oct 8
HumanEval Passes. Production Burns. The Real Story of AI Code Generation in the Vibe Coding Era
#
ai
#
llm
#
programming
#
benchmark
Comments
Add Comment
8 min read
I benchmarked seven hotel price APIs on the same Rome room, and the prices were 16% apart because of tax
Matan Rabi
Matan Rabi
Matan Rabi
Follow
Oct 7
I benchmarked seven hotel price APIs on the same Rome room, and the prices were 16% apart because of tax
#
api
#
webdev
#
benchmark
#
travel
Comments
Add Comment
4 min read
I benchmarked four Google Flights APIs on the same searches, and the fares matched to the dollar
Matan Rabi
Matan Rabi
Matan Rabi
Follow
Oct 7
I benchmarked four Google Flights APIs on the same searches, and the fares matched to the dollar
#
api
#
webdev
#
benchmark
#
travel
Comments
Add Comment
4 min read
What Independent Benchmarks Say About Opus 5.5
Nomad
Nomad
Nomad
Follow
Oct 7
What Independent Benchmarks Say About Opus 5.5
#
ai
#
claude
#
llm
#
benchmark
Comments
Add Comment
8 min read
LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers
RESK
RESK
RESK
Follow
Oct 7
LLM Evaluation: How a Benchmark Turns Raw Answers Into Comparable Numbers
#
llm
#
evaluation
#
benchmark
#
ai
Comments
Add Comment
4 min read
Parakeet Redux: is a 178 MB speech model useful in the real world?
Abhash Chakraborty
Abhash Chakraborty
Abhash Chakraborty
Follow
Oct 6
Parakeet Redux: is a 178 MB speech model useful in the real world?
#
machinelearning
#
speechrecognition
#
opensource
#
benchmark
Comments
Add Comment
16 min read
How we use Jev to answer quiz questions (a 90-question benchmark)
jack
jack
jack
Follow
Oct 6
How we use Jev to answer quiz questions (a 90-question benchmark)
#
ai
#
llm
#
benchmark
#
chrome
Comments
Add Comment
3 min read
When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician
Ward Ed
Ward Ed
Ward Ed
Follow
Oct 6
When a 0.4-Point Lead Means Nothing: Reading Open-Model Leaderboard Margins Like a Statistician
#
huggingface
#
llm
#
benchmark
#
machinelearning
Comments
Add Comment
7 min read
How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides
Ward Ed
Ward Ed
Ward Ed
Follow
Oct 5
How Hugging Face Official Benchmark Leaderboards Actually Work: .eval_results YAML, the base_model Filter, and the 30% the Default View Hides
#
huggingface
#
llm
#
benchmark
#
machinelearning
Comments
1
comment
10 min read
Half the MCP servers that answer you don't actually work
Pennyforge
Pennyforge
Pennyforge
Follow
Oct 6
Half the MCP servers that answer you don't actually work
#
mcp
#
modelcontextprotocol
#
devtools
#
benchmark
Comments
2
comments
8 min read
Hugging Face now has 48 official benchmarks. Here is what the map looks like
ai maya
ai maya
ai maya
Follow
Oct 4
Hugging Face now has 48 official benchmarks. Here is what the map looks like
#
ai
#
llm
#
machinelearning
#
benchmark
Comments
Add Comment
3 min read
👋
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account