-
float16-cloud
- Bangkok, Thailand
- @KMatiDev1
Stars
AlphaFold 3 inference pipeline.
A high-throughput and memory-efficient inference and serving engine for LLMs
HeartMuLa Official Repo: The Most Powerful Open-Source Music Generation Model of 2026
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
An open-source AI coding agent that lives in your terminal.
The state-of-the-art image restoration model without nonlinear activation functions.
N_Body_Simulation_CPU_versus_GPU
JamePeng / llama-cpp-python
Forked from abetlen/llama-cpp-pythonEfficiency Python bindings for llama.cpp
CV-CUDA™ is an open-source, GPU accelerated library for cloud-scale image processing and computer vision.
PyTorch compiler that accelerates training and inference. Get built-in optimizations for performance, memory, parallelism, and easily write your own.
NumPy and SciPy on Multi-Node Multi-GPU systems
Code I wrote for my AI & LLM workshops
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
🌈 React for interactive command-line apps
🦄 Record your terminal and generate animated gif images or share a web player
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…
Experimental projects related to TensorRT
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
Together Mixture-Of-Agents (MoA) – 65.1% on AlpacaEval with OSS models
An AI search engine inspired by Perplexity