Benchmark coding LLMs on Project Euler–style problems under the same machine, same time budget, and the same verification rules. Each model’s solution is executed in isolation, runtime is measured, correctness is verified against a ground truth, a score is computed (by difficulty & latency), and a leaderboard is printed.
- Discovers problems in
Questions/<problem_id>/. - Runs each
*.pysolution (one per model) in a fresh subprocess. - Measures wall-clock time (seconds).
- Verifies stdout against the expected single value.
- Applies scoring rules (difficulty points, per-minute late penalty, timeout = zero).
- Writes a detailed CSV (
results.csv) and prints a ranked leaderboard.
Snapshot source: the repository's current results.csv on April 8, 2026.
- Coverage: 20 problems x 8 models = 160 runs
- Current model set:
ChatGPT-5,ChatGPT-5.4,Claude-Haiku-4.5,Claude-Opus-4.6,DeepSeek-2025-10,DeepSeek-V3.2,Gemini-3.1-Pro,Gemini-Flash-2.5-pro
| Rank | Model | Score |
|---|---|---|
| 1 | Gemini-3.1-Pro |
828.75 |
| 2 | Claude-Opus-4.6 |
753.75 |
| 3 | ChatGPT-5.4 |
653.75 |
| 4 | DeepSeek-V3.2 |
595.00 |
| 5 | Claude-Haiku-4.5 |
520.00 |
| 6 | Gemini-Flash-2.5-pro |
498.75 |
| 7 | ChatGPT-5 |
493.75 |
| 8 | DeepSeek-2025-10 |
347.50 |
| Rank | Model | Score |
|---|---|---|
| 1 | Gemini-3.1-Pro |
1750.00 |
| 2 | Claude-Opus-4.6 |
1600.00 |
| 3 | ChatGPT-5.4 |
1500.00 |
| 4 | DeepSeek-V3.2 |
1400.00 |
| 5 | Claude-Haiku-4.5 |
1270.00 |
| 6 | Gemini-Flash-2.5-pro |
1180.00 |
| 7 | ChatGPT-5 |
1100.00 |
| 8 | DeepSeek-2025-10 |
800.00 |
| Model | Correct / 20 | Accuracy |
|---|---|---|
Gemini-3.1-Pro |
18 / 20 |
90% |
Claude-Opus-4.6 |
16 / 20 |
80% |
ChatGPT-5.4 |
15 / 20 |
75% |
DeepSeek-V3.2 |
14 / 20 |
70% |
Claude-Haiku-4.5 |
13 / 20 |
65% |
Gemini-Flash-2.5-pro |
12 / 20 |
60% |
ChatGPT-5 |
11 / 20 |
55% |
DeepSeek-2025-10 |
8 / 20 |
40% |
- Solved by all 8 models:
74,112,172,190,301,357,493 - Solved by exactly 1 model:
54,439 - Solved by no model yet:
505 - Toughest partially-solved problems in this snapshot:
399(3 / 8),502(2 / 8),81(3 / 8)
For full per-run detail, see results.csv.
LLMHackathon/
├── runner.py # main runner (executes, verifies, scores, ranks)
├── Questions/ # problems live here (directory per problem id)
│ ├── 54/
│ │ ├── GPT-5.py
│ │ └── Claude-Haiku-4.5.py
│ ├── 81/
│ │ └── SomeModel.py
│ └── ...
└── README.md
Note: The directory name is
Questions.
You can change it inrunner.pyby editingQuestions Directory.
- Python 3.10+ (tested with Python 3.11)
- Same machine for all runs (so timing is comparable)
- Model solutions must print only the final answer via
print(...) - No interactive input; no network access
- Standard library is preferred;
numpyis available when it materially helps
-
Put model solutions under
Questions/<problem_id>/<ModelName>.py.Example:
Questions/54/GPT-5.py Questions/54/Claude-Haiku-4.5.py -
Ensure each script prints one single line with the final numeric answer:
Questions/54/GPT-5.py Questions/54/Claude-Haiku-4.5.py # ... your computation ... print(376)
-
Update
EXPECTED_OUTPUTSinrunner.pywith the expected(answer, difficulty):EXPECTED_OUTPUTS = { "54": ("376", 10), "81": ("427337", 10), "99": ("709", 10), # ... }
-
Run:
python3 runner.py
-
Inspect:
- Console table (per-run results + leaderboard)
results.csvfor archival & post-analysis
- Each problem has a difficulty score (e.g., 10, 25, 70, …).
- If the model prints the correct answer within 60 seconds → full difficulty points.
- If the answer is correct but takes longer than 60 seconds, subtract 10 points per extra full minute:
- 61–120 s → −10
- 121–180 s → −20
- …
- The minimum is 0 points (no negative scores).
- If the answer is wrong, timeout, or error → 0 points.
- Final leaderboard = sum of points across all problems, highest to lowest.
runner.pyusesTIMEOUT = 600seconds for execution safety.
The 60 s rule above applies to scoring (full points if≤ 60 s, otherwise per-minute penalties).
Each row = one run of <problem_id>/<model>.py.
| Column | Meaning |
|---|---|
Question No |
Problem id (e.g., 54) |
LLM |
Model name (derived from filename) |
Execution Time (s) |
Runtime in seconds, or timeout / error |
Is Correct |
✅ for correct, ❌ (…reason…) otherwise |
At the end of a run, the script:
- writes the CSV,
- re-reads it,
- computes scores,
- prints the leaderboard sorted from highest to lowest.
Use a disciplined prompt so models produce computational code rather than hard-coded answers:
In the attached image, you are given an algorithmic problem in the style of Project Euler.
This is a hackathon challenge. Your code will be executed in an independent evaluation system, and only the value printed via print(...) will be compared with the actual correct answer.
Your code is expected to compute the answer by itself by solving the given problem.
The execution environment provides the following resources:
• A 16-core multi-processor CPU
• 64 GB of RAM
• An NVIDIA GPU with 8 GB of VRAM
• A Linux environment with Python 3.11 installed
• You may use NumPy, multiprocessing, and other standard Python libraries
• You may not use third-party libraries (e.g., sympy, numba, gmpy2, tensorflow)
Please:
1. Analyze the problem and choose an appropriate algorithm.
2. Consider using parallel processing (e.g., multiprocessing) or efficient memory handling to speed up computation.
3. Use fast numerical libraries like numpy when necessary.
4. Write Python 3 code that runs correctly and outputs only the final result using print(...).
5. Do not hardcode the answer in your code; compute it programmatically.
6. Use print(...) only for the final result — no debug or intermediate outputs.
7. The code must finish execution within 60 seconds and produce the correct result.
Notes:
• You are encouraged to utilize multi-core capabilities using tools such as multiprocessing.Pool or concurrent.futures.ProcessPoolExecutor.
• GPU acceleration may indirectly help via NumPy, but direct CUDA programming is not allowed.
• Please do not calculate or provide the final answer yourself — just write the code that computes it.
-
Floating answers (e.g., Euler #493)
The judge uses string equality. If a problem expects a decimal string, print that exact string (e.g.,6.818741802).
If you prefer tolerance-based checking, extendrunner.pyaccordingly. -
Parallelism
Solutions may leveragemultiprocessingto use multiple cores. Each solution runs in a separate process, so memory leaks don’t accumulate across runs. -
Determinism
Scripts must not depend on randomness unless they seed and deterministically converge to the exact expected value.
- Create a folder:
Questions/<problem_id>/ - Add one or more model files:
*.py - Add the expected
(answer, difficulty)toEXPECTED_OUTPUTSinrunner.py - Run
python3 runner.py
- Submit PRs that:
- Add new problems/folders
- Add new model baselines
- Improve scoring/verification
- Document best prompts for fair comparison
Please keep solutions free of external dependencies (standard library only).
GNU GENERAL PUBLIC LICENSE