-
AWS
- NYC
-
17:52
(UTC -04:00) - in/austinmwelch
Lists (2)
Sort Name ascending (A-Z)
Stars
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台…
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
Open-source framework for the research and development of foundation models.
Our library for RL environments + evals
Warp is an agentic development environment, born out of the terminal.
Benchmark LLMs by fighting in Street Fighter 3! The new way to evaluate the quality of an LLM
Qwen3.6-35B-A3B-heretic NVFP4 + DFlash speculative decoding on DGX Spark (GB10/sm_121a). Source-built vLLM image + 7 patches + comprehensive deployment guide.
🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
🦀🌡️ Real-time system monitor for Apple Silicon Macs (M1–M5). No sudo. TUI, JSON/Prometheus metrics server, and Rust library.
Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.
A framework for teaching AI to write like you. Not like a better version of you. Like you.
Atropos is a Language Model Reinforcement Learning Environments framework for collecting and evaluating LLM trajectories through diverse environments
🎨 NeMo Data Designer: Generate high-quality synthetic data from scratch or from seed data.
User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
Run Coding Agents in Sandboxes. Control Them Over HTTP. Supports Claude Code, Codex, OpenCode, and Amp.
World model reinforcement learning for multi-turn VLM agents. RL for vision framework (NeurIPS 2025).
Official repository of the 2025 paper, LLM Economist: Large Population Models and Mechanism Design in Multi-Agent Generative Simulacra.
JupyMD: Use Obsidian as a Jupyter notebook IDE
Every Eval Ever is a shared schema and crowdsourced eval database. It defines a standardized metadata format for storing AI evaluation results — from leaderboard scrapes and research papers to loca…
bf16 LoRA fine-tuning of [Qwen3.5-35B-A3B](https://huggingface.co/unsloth/Qwen3.5-35B-A3B) (a 35B-total / 3B-active Mixture-of-Experts vision-language model) on a single NVIDIA DGX Spark — without …
Trio – a friendly Python library for async concurrency and I/O
AI agents running research on single-GPU nanochat training automatically
Framework for evaluating and improving agents
Scripts for agents, shared between my repositories.
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.