- Nyland, Finland
-
17:20
(UTC +03:00) - https://musayev.me
- in/mehtimusayev
- https://huggingface.co/mehti
- All languages
- ActionScript
- Awk
- Batchfile
- BitBake
- C
- C#
- C++
- CSS
- Clojure
- CoffeeScript
- Cuda
- Cython
- Dart
- Dockerfile
- EJS
- Flix
- Fortran
- FreeMarker
- Go
- HTML
- Java
- JavaScript
- Julia
- Jupyter Notebook
- Kotlin
- Lean
- MATLAB
- MDX
- Makefile
- OCaml
- OpenEdge ABL
- PHP
- PowerShell
- Python
- R
- Ruby
- Rust
- SCSS
- Scala
- Shell
- Swift
- TeX
- Terra
- Twig
- TypeScript
- Vim Script
- YARA
Starred repositories
Agent memory for LLMs: 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, episodic and semantic memory, MemGPT, Mem0, Letta, Zep, Graphiti, LoCoMo benchma…
The open-source AI workbench for scientific research
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing.
DFlash: Block Diffusion for Flash Speculative Decoding
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
Perplexity open source garden for inference technology
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization
A simple SWE style browser agent framework that achieves SOTA results on long horizon web tasks.
1 place to call all your agents - OpenCode, Hermes, Claude Managed Agents, Cursor Agents API, DeepAgents.
Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.
Incremental engine for long horizon agents 🌟 Star if you like it!
cuda-oxide is an experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign languag…
TokenSpeed is a speed-of-light LLM inference engine.
Polymarket Data Retriever that fetches, processes, and structures Polymarket data including markets, order events and trades.
1K resolution vision transformers pretrained on 1B human images.
TriAttention — Efficient long reasoning with trigonometric KV cache compression. Enables OpenClaw local deployment on memory-constrained GPUs.
NVIDIA AITune is an inference toolkit designed for tuning and deploying Deep Learning models with a focus on NVIDIA GPUs.
Dimensional is the agentic operating system for physical space. Command humanoids, quadrupeds, drones, and other hardware platforms in natural language and build multi-agent systems that work seaml…
Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized experts across multiple task domains.
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
The best-benchmarked open-source AI memory system. And it's free.
Multi-agent systems, memory, planning, reasoning loops