- Paris, France
- https://numberly.com/
- in/gabriel-l-2850a9a0
Stars
- All languages
- Assembly
- Astro
- Awk
- C
- C#
- C++
- CSS
- Clojure
- CodeQL
- CoffeeScript
- Common Lisp
- Cuda
- Cython
- Dart
- Dockerfile
- Elixir
- Elm
- Emacs Lisp
- Erlang
- F#
- F*
- Gherkin
- Go
- Go Template
- HCL
- HTML
- Handlebars
- Haskell
- Java
- JavaScript
- Jinja
- Julia
- Jupyter Notebook
- Kotlin
- Lean
- Lua
- MDX
- MLIR
- Makefile
- Meson
- Mustache
- Nim
- OCaml
- Objective-C
- PHP
- PLpgSQL
- Perl
- PostScript
- PowerShell
- Python
- QML
- R
- Ruby
- Rust
- SCSS
- SQL
- Scala
- Shell
- Starlark
- Svelte
- Swift
- Tcl
- TeX
- Tree-sitter Query
- TypeScript
- Verilog
- Vim Script
- YARA
- Zig
- jq
Give Claude Code eyes 👁️ — a camera skill so it can SEE real hardware, displays and wiring: verify a rendered panel, check wiring before power-on, and catch bugs that live on the glass, not the logs.
The HTML5 Creation Engine: Create beautiful digital content with the fastest, most flexible 2D WebGL renderer.
Qwen-AgentWorld: Language World Models for General Agents
YC Bench: a Live Benchmark for Forecasting Startup Outperformance in Y Combinator Batches arXiv:2604.02378
Open-source benchmark for evaluating LLMs on 220 real professional tasks across 9 sectors and 44 occupations. Reproducible experiments, artifact validation, grading, and a live evidence dashboard.
Frontier Models playing the board game Diplomacy.
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
DFloat11 [NeurIPS '25]: Lossless Compression of LLMs and DiTs for Efficient GPU Inference
Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining"
An agent benchmark with tasks in a simulated software company.
Official repository for "DynaSaur: Large Language Agents Beyond Predefined Actions"
👩⚖️ Agent-as-a-Judge: The Magic for Open-Endedness
SWE-bench: Can Language Models Resolve Real-world Github Issues?
An SRE agent benchmark inspired by SWE-bench, focused on real-world Site Reliability Engineering tasks: incident response, infra changes, observability triage, and reliability improvements across K…
MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation
A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.
[NeurIPS 2025 Spotlight] Scaling Computer-Use Grounding via UI Decomposition and Synthesis
[ACL'26] Official Code for "ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning"
Cargo subcommand `release`: everything about releasing a rust crate.
A batteries-included framework for building web apps
The elegant bundler for libraries powered by Rolldown
Easily export your Rust types to other languages
A fully-featured team and group chat application that you can easily selfhost.