Skip to content
View yangxinQA's full-sized avatar

Block or report yangxinQA

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

生成测试用例skill

Python 118 17 Updated Apr 3, 2026

Sample application that showcases Data Cloud, Agents and Prompts.

HTML 154 110 Updated Sep 1, 2026

The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

TypeScript 390,262 82,086 Updated Sep 22, 2026

The open source coding agent.

TypeScript 209,379 27,612 Updated Sep 22, 2026

The LLM Evaluation Framework

Python 18,394 1,967 Updated Sep 22, 2026

Awesome_Multimodel is a curated GitHub repository that provides a comprehensive collection of resources for Multimodal Large Language Models (MLLM). It covers datasets, tuning techniques, in-contex…

378 25 Updated Jul 3, 2026

Central repo to connect and document components/repos needed for IOS stf support

Go 161 64 Updated Dec 5, 2025

from vibe coding to agentic engineering - practice makes claude perfect

HTML 66,240 6,582 Updated Sep 22, 2026

Public repository for Agent Skills

Python 177,638 21,038 Updated Sep 22, 2026
Go 115 17 Updated Oct 6, 2025

Curated list of awesome Cursor Rules .mdc files

Python 3,573 447 Updated May 19, 2026

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

JavaScript 265,291 39,633 Updated Sep 21, 2026

✨✨Latest Advances on Multimodal Large Language Models

18,032 1,137 Updated Sep 18, 2026

Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1

Python 77,445 12,456 Updated Aug 26, 2026

Agentic Design Patterns: A Hands-On Guide to Building Intelligent Systems by Antonio Gulli

Jupyter Notebook 222 2,074 Updated Sep 7, 2025

谷歌新书Agent设计模式(agentic design patterns)最佳中文版,持续优化。附:在线阅读、pdf和epub电子书下载。

HTML 8,047 1,170 Updated Aug 30, 2026

A Step-by-Step Implementation of Qwen 3 MoE Architecture from Scratch

Jupyter Notebook 84 9 Updated Aug 5, 2025

AI Observability & Evaluation

Python 11,579 1,145 Updated Sep 22, 2026

[ACL 2024] MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

153 118 Updated Jul 24, 2024

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

Jupyter Notebook 105,401 16,170 Updated Sep 22, 2026

Codes for our paper "ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate"

Python 343 32 Updated Oct 19, 2024

Evaluate your LLM's response with Prometheus and GPT4 💯

Python 1,117 68 Updated Apr 25, 2025

Multi-Turn RAG Benchmark

Python 155 30 Updated Sep 4, 2026

An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.

Python 39,548 4,776 Updated May 1, 2026

The code and data for "MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark" [NeurIPS 2024]

Python 427 55 Updated Mar 18, 2026

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

Python 3,459 495 Updated Sep 21, 2026

Z-Bench 1.0 by 真格基金:一个麻瓜的大语言模型中文测试集。Z-Bench is a LLM prompt dataset for non-technical users, developed by an enthusiastic AI-focused team in Zhenfund.

504 41 Updated Jun 28, 2023
Next