guardllm uses a defense-in-depth security pipeline designed to harden MCP servers and MCP clients against prompt injection, data exfiltration, replay attacks, and trust-boundary violations from unknown-provenance content sources such as web search results, emails, documents, calendar data, and other untrusted inputs.
| Layer | Name | Purpose | Primary Module |
|---|---|---|---|
| L0 | Input Sanitization | Strip hidden HTML, dangerous attributes/comments, invisible Unicode, and normalize content before further processing | guardllm.security.sanitizer |
| L1 | Content Isolation | Wrap untrusted input in <untrusted_content ...> tags with source attribution |
guardllm.security.isolation |
| L2 | Action Gate | Optional interactive confirmation gate for sensitive actions | guardllm.security.action_gate |
| L3 | Source Gate | Enforce provenance-based KG extraction policies (allow, quarantine, block) |
guardllm.security.source_gate |
| L4 | Canary Detection | Session canary generation/detection to flag leakage/exfiltration | guardllm.security.canary |
| L5 | Tool Firewall | Authorize tools by policy + explicit authorization events | guardllm.security.policy_engine |
| L6 | Outbound DLP | Block high-overlap egress and secret-like patterns | guardllm.security.outbound_dlp |
| L7 | Provenance Tracking | Track untrusted spans and block suspicious reuse across trust boundaries | guardllm.security.provenance |
| L8 | Rate Limiting | Per-context action throttling for abuse resistance | guardllm.security.rate_limiter |
| L9 | Request Binding | Bind tool execution to message hash + args hash + TTL | guardllm.security.request_binding |
| L10 | OAuth Scope Progression | Scope narrowing/escalation policy between auth/session states | Host application responsibility |
| L11 | Audit Logging | Structured security event logging for analysis and incident response | guardllm.security.audit |
Note: Layer numbering intentionally matches the parent hardening model. L10 is documented for completeness but remains outside the library boundary.
This section is the source of truth for what is wired through guardllm.Guard today.
| Layer | Status | Guard API Surface |
|---|---|---|
| L0 Input Sanitization | Implemented | process_inbound(...) |
| L1 Content Isolation | Implemented | process_inbound(...) |
| L2 Action Gate | Implemented | confirm_action(...), guard_tool_call(..., require_confirmation=True) |
| L3 Source Gate | Implemented | guardllm.security.source_gate.check_extraction_allowed(...) |
| L4 Canary Detection | Implemented | Guard(canary_session_id=...) + inbound/outbound checks |
| L5 Tool Firewall | Implemented | check_tool_call(...), guard_tool_call(...) |
| L6 Outbound DLP | Implemented | check_outbound(...) |
| L7 Provenance Tracking | Implemented | process_inbound(...), check_outbound(...) |
| L8 Rate Limiting | Implemented | check_tool_call(...), check_outbound(...), guard_tool_call(...) |
| L9 Request Binding | Implemented | bind_request(...), check_tool_call(...), guard_tool_call(...) |
| L10 OAuth Scope Progression | Not implemented in library | Host application responsibility |
| L11 Audit Logging | Implemented | Guard(audit_logger=...) emits security events |
| Validation (spec §12.2) | Implemented | validate_tool_args(...), guard_tool_call(validate=True) |
| Error Sanitization (spec §12.6) | Implemented | sanitize_exception(...) |
The central orchestrator is guardllm.security.pipeline.SecurityPipeline, exposed through the high-level Guard API.
Inbound path:
- Sanitize untrusted input (L0)
- Isolate by trust level (L1)
- Ingest for outbound DLP comparisons (L6 data prep)
- Record provenance spans (L7)
- Check canary presence in inbound payloads (L4)
Tool-call path:
- Tool authorization/firewall checks (L5)
- Rate limiting (L8)
- Optional request binding verification (L9)
Outbound path:
- DLP overlap/secret checks (L6)
- Provenance reuse guard (L7)
- Rate limiting (L8)
- Canary leakage detection (L4)
guardllm supports explicit security contexts for:
- MCP server responses (
context_mcp_server) - MCP client requests (
context_mcp_client) - Documents (
context_document) - Web results (
context_web)
You can also define custom SecurityContext values for sources like:
email_contentcalendar_contenttool_outputrag_content
These source types integrate directly with source-gate and provenance behavior.
- Trust defaults to
UNTRUSTED. - Destructive tool calls are blocked unless explicitly enabled via
PolicyConfig(enable_destructive=True). - Destructive tool calls require authorization events in client mode.
- Request binding is optional but recommended for all write-capable actions.
- OAuth/OIDC integration is supported via host-side scope-to-policy mapping (see
docs/oauth_integration.md).
- Hidden-instruction prompt injection in HTML/text payloads
- Unicode obfuscation attacks (zero-width/bidi controls)
- Exfiltration by copying untrusted spans into outbound content
- Replay/deferred tool execution after conversation state changes
- Over-privileged tool invocation and destructive action abuse
guardllm is an application-layer hardening library. It does not replace:
- network segmentation
- host/container isolation
- secret management systems
- transport-layer authN/authZ
- OAuth/OIDC token issuance, validation, and lifecycle management
Use guardllm as one layer in a full security architecture.