Skip to content
View Mattdl's full-sized avatar

Block or report Mattdl

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

thinkingbox is a framework for defining tool as MCP servers, running LLM agents against them, and evaluating agent behavior — for offline training-data generation, reinforcement-learning training l…

Python 28 1 Updated Sep 1, 2026

Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…

Python 143,651 22,978 Updated Aug 31, 2026

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

Python 16,584 1,559 Updated Aug 29, 2026

#1 Persistent memory for AI coding agents based on real-world benchmarks

TypeScript 27,895 2,406 Updated Aug 31, 2026

The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.

Python 64,512 7,558 Updated Sep 1, 2026

Secure, cross-platform Git credential storage with authentication to GitHub, Azure Repos, and other popular Git hosting services.

C# 9,239 2,904 Updated Sep 1, 2026

An agentic evaluation framework

Python 23 5 Updated Feb 11, 2026
Python 1,635 187 Updated Aug 11, 2026

Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.

24,528 2,606 Updated Aug 23, 2026

Official Code of Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

Python 2,568 297 Updated Oct 5, 2025

Readymade evaluators for agent trajectories

Python 712 54 Updated Jul 14, 2026

A Docker sandbox template for running GitHub Copilot CLI in an isolated environment, similar to how Docker supports Claude Code and Gemini CLI via docker sandbox run

Shell 23 2 Updated Aug 27, 2026

Zotero MCP: Connects your Zotero research library with Claude and other AI assistants via the Model Context Protocol to discuss papers, get summaries, analyze citations, and more.

Python 4,861 387 Updated Aug 25, 2026

DSPy: The framework for programming—not prompting—language models

Python 37,707 3,279 Updated Sep 1, 2026

LLM Wiki is a cross-platform desktop application that turns your documents into an organized, interlinked knowledge base — automatically. Instead of traditional RAG (retrieve-and-answer from scratc…

TypeScript 17,218 2,043 Updated Aug 25, 2026

AI agents running research on single-GPU nanochat training automatically

Python 95,041 13,380 Updated Mar 26, 2026

TypeScript AI agent orchestration framework with dynamic workflows. Describe the goal, not the graph: a coordinator plans the task DAG at runtime and runs it on any LLM (Claude, ChatGPT, Gemini, De…

TypeScript 6,852 2,435 Updated Aug 31, 2026

Open-source Claude Code skills and Codex skills for AI-first work. Audit, re-engineer, and bootstrap projects with AI-first design principles.

TypeScript 97 3 Updated Jul 13, 2026

The Multilingual Entity Linking of Occupations (MELO) Benchmark

Python 5 2 Updated Jan 25, 2025

WorkRB: Work Research Benchmark

Python 41 7 Updated Jul 22, 2026

SKILLSPAN: Competences as Spans for Skill Extraction from Job Postings

Perl 69 16 Updated Feb 13, 2025

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python 164,697 34,413 Updated Sep 1, 2026

Late Interaction Models Training & Retrieval

Python 888 95 Updated Jul 23, 2026

GitHub Mirror of RecPack: Experimentation Toolkit for Top-N Recommendation (see https://gitlab.com/recpack-maintainers/recpack)

Python 22 3 Updated Dec 11, 2023

State-of-the-Art Embeddings, Retrieval, and Reranking

Python 19,056 2,870 Updated Sep 1, 2026

The code used to evaluate embedding models on the Massive Legal Embedding Benchmark (MLEB).

Python 39 8 Updated Feb 24, 2026

MTEB: State-of-the-art evaluation of embeddings across languages and modalities

Python 3,413 683 Updated Sep 1, 2026

🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL

Python 1,200 110 Updated Aug 17, 2026

In this codebase we establish a benchmark for egocentric user adaptation based on Ego4d.First, we start from a population model which has data from many users to learn user-agnostic representations…

Python 15 Updated Jul 24, 2026

PyTorch implementation of various methods for continual learning (XdG, EWC, SI, LwF, FROMP, DGR, BI-R, ER, A-GEM, iCaRL, Generative Classifier) in three different scenarios.

Jupyter Notebook 1,878 345 Updated Nov 5, 2025
Next