-
NJU -> HKUST -> Intel
- Shanghai
Stars
"OpenHarness: Open Agent Harness with a Built-in Personal Agent--Ohmo!"
An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
A Claude Code plugin that shows what's happening - context usage, active tools, running agents, and todo progress
tomreinert / ghostty-sidegeist
Forked from ghostty-org/ghostty👻 Ghostty Sidegeist: a Ghostty fork with vertical tabs in a sidebar and a git panel.
🚀 Efficient implementations for emerging model architectures
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
Some common CUDA kernel implementations (Not the fastest).
Issues related to MLPerf® Inference policies, including rules and suggested changes
felipeagc / ollama
Forked from ollama/ollamaGet up and running with Llama 2, Mistral, and other large language models locally.
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. Tensor…
A high-throughput and memory-efficient inference and serving engine for LLMs
⚡ Build your chatbot within minutes on your favorite device; offer SOTA compression techniques for LLMs; run LLMs efficiently on Intel Platforms⚡
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
The official PyTorch implementation of Google's Gemma models
An Open-Source Distributed Deep Learning Framework
Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discr…
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
sgwhat / analytics-zoo
Forked from intel/BigDLDistributed Tensorflow, Keras and PyTorch on Apache Spark/Flink & Ray
Java 面试 & 后端通用面试指南,覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发
此项目是机器学习(Machine Learning)、深度学习(Deep Learning)、NLP面试中常考到的知识点和代码实现,也是作为一个算法工程师必会的理论基础知识。
🚀 fullstack tutorial 2022,后台技术栈/架构师之路/全栈开发社区,春招/秋招/校招/面试