Lists (2)
Sort Name ascending (A-Z)
Stars
GitHub Actions jobs on AWS Lambda MicroVMs, one ephemeral runner per job. A CDK construct library.
Benchmark framework that measures cost, quality, and duration of coding agents across any AI Coding Assistant CLI, any model, and any use case — with pluggable verification and real-repo support.
Loom for AWS is an enterprise-grade platform for building, deploying, and operating AI agents on Amazon Bedrock AgentCore Runtime and AWS Strands Agents.
An LLM-as-a-judge HTTP proxy to secure agents in production
AAI Partner Basecamp exercises
Agentic AI security tool that applies proactive, attacker-first analysis directly to source code.
Innovation Sandbox on AWS enables cloud administrators to automate the management of temporary sandbox environments by implementing service control policies, spend controls, and account recycling m…
Reference code for the Meta-Harness paper.
Meta-Harness: 76.4% on Terminal-Bench 2.0 (Claude Opus 4.6)
lilianweng / claude-code
Forked from codeaashu/claude-codeClaude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflo…
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Use Garry Tan's exact Claude Code setup: 23 opinionated tools that serve as CEO, Designer, Eng Manager, Release Manager, Doc Engineer, and QA
Agentic AI assistant for proactive AWS Health Event management and impact analysis
Manages Unified Access to Generative AI Services built on Envoy Gateway
ERPAVal — autonomous software development. Six-phase Explore/Research/Plan/Act/Validate/Compound workflow with classifier-driven routing and a compounding lessons store.
A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
This workshop teaches systematic approaches to evaluating Generative AI workloads for production use. You'll learn to build evaluation frameworks that go beyond basic metrics to ensure reliable mod…
Next Generation Agentic Proxy for AI Agents and MCP servers
macOS app to create standard or customized configuration profiles.
Caveman Compression is a semantic compression method for LLM contexts. It removes predictable grammar while preserving the unpredictable, factual content that defines meaning.
Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.
Zero-dependency Chrome DevTools Protocol CLI for AI agents. 45+ commands, per-tab daemons, security hardened.