Skip to content
View Ottovonxu's full-sized avatar

Block or report Ottovonxu

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Cuda kernels for leveraging LLM sparsity to improve throughput and decrease the memory requirements during inference and training.

Cuda 254 25 Updated Jun 29, 2026

An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.

Rust 194,923 109,480 Updated Jun 26, 2026

A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power.

JavaScript 1,039 110 Updated Jul 16, 2026

Code for the NPJ AI paper "How Large Language Models Encode Theory-of-Mind: A Study on Sparse Parameter Patterns"

Python 11 4 Updated Nov 9, 2025

A PyTorch native platform for training generative AI models

Python 5,567 914 Updated Jul 26, 2026

Train transformer language models with reinforcement learning.

Python 18,936 2,869 Updated Jul 26, 2026

A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.

Python 2,252 247 Updated Oct 16, 2025

[EMNLP'25] Code for the paper "DEL-ToM: Inference-Time Scaling for Theory-of-Mind Reasoning via Dynamic Epistemic Logic"

Python 6 2 Updated Oct 14, 2025

A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.

Python 4,632 755 Updated May 17, 2026

PubMedQA: A Dataset for Biomedical Research Question Answering

Python 433 59 Updated Apr 18, 2023

Vortex: Programmable Sparse Attention for Agents as Algorithm Designers

Python 67 11 Updated Jun 24, 2026

🤖 MLE-Agent: Your intelligent companion for seamless AI engineering and research. 🔍 Integrate with arxiv and paper with code to provide better code/research plans 🧰 OpenAI, Anthropic, Gemini, Ollam…

Python 1,564 107 Updated Jul 10, 2026

Optimize prompts, code, and more with AI-powered Reflective Optimization

Jupyter Notebook 5,856 487 Updated Jul 26, 2026

Open-source implementation of AlphaEvolve

Python 6,799 1,092 Updated Jul 18, 2026

[ICML 2026] Decoding Tree Sketching (DTS): a training-free & model agonistic & plug-in framework for LLM parallel reasoning.

Python 71 12 Updated May 12, 2026

This guide provides a setup for running `vLLM` with `Triton` on NVIDIA GH200.

4 Updated Apr 6, 2025

RapidIn: Scalable Influence Estimation for Large Language Models (LLMs). The implementation for paper "Token-wise Influential Training Data Retrieval for Large Language Models" (Accepted on ACL 2024).

Python 22 5 Updated Mar 10, 2026

A simple and effective LLM pruning approach.

Python 868 131 Updated Aug 9, 2024

The fastest, litest AI Gateway. Rust core with Python SDK. Call 100+ LLM APIs in OpenAI (or native) format with cost tracking, guardrails, load balancing, and logging [Bedrock, Azure, OpenAI, Anthr…

Python 54,743 10,087 Updated Jul 26, 2026

H-Net: Hierarchical Network with Dynamic Chunking

Python 869 102 Updated Nov 20, 2025

A lightweight python-only library for reading and writing SMILES strings

Python 164 25 Updated Jul 6, 2026

Code repository for ICLR 2025 paper "LeanQuant: Accurate and Scalable Large Language Model Quantization with Loss-error-aware Grid"

Python 29 3 Updated Mar 2, 2025
Python 23 3 Updated May 23, 2025

Code Repo for EMNLP paper: Do LLMs Know to Respect Copyright Notice

Python 5 Updated Nov 18, 2024

A game theoretic approach to explain the output of any machine learning model.

Jupyter Notebook 25,645 3,733 Updated Jul 21, 2026

[NeurIPS 2024] Simple and Effective Masked Diffusion Language Model

Python 702 102 Updated Sep 29, 2025

Production-tested AI infrastructure tools for efficient AGI development and community-driven innovation

8,032 291 Updated May 15, 2025

A novel medical large language model family with 13/70B parameters, which have SOTA performances on various medical tasks

Python 166 20 Updated Jan 15, 2025

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,779 1,104 Updated Apr 30, 2026

MoBA: Mixture of Block Attention for Long-Context LLMs

Python 2,153 157 Updated Apr 3, 2025
Next