Skip to content

Routing smokes never run the default threat-detection engine, and nothing checks that detection stays unrouted #67789

Description

@SivaKesava1

Problem

The routed smokes only run the legacy detection engine

All three routing smokes set the gh-aw-detection feature flag to false:

  • .github/workflows/smoke-copilot-routed.md
  • .github/workflows/smoke-copilot-sdk-routed.md
  • .github/workflows/smoke-pi-routed.md

This flag selects the legacy inline threat-detection engine. The default is the external threat-detect engine (see docs/src/content/docs/reference/threat-detection.md and reference/artifacts.md). These three workflows are the only ones in this repo that use model-routing. As a result, CI never runs the default detection engine under model routing.

Why the flag is set. It was not a deliberate choice for routing:

Nothing checks that detection stays unrouted

Detection does run in the routed smokes. In run 38173405739 (Smoke Copilot Routed, main at 3667115, all jobs green), the detection job runs "Execute GitHub Copilot CLI" under AWF. However, the smoke's only assertions are the agent-job post-step (E1, E2, R1–R4 in actions/setup/js/smoke_model_routing_assertions.cjs). Nothing checks how detection behaves under routing.

docs/src/content/docs/reference/model-routing.md ("Known limitations") says threat detection is not routed: it keeps its configured detection engine and model. This has broken before:

None of these runs detection under routing in CI and checks its runtime evidence. A regression would only show up as a compiled-YAML diff, or as a silent missing verdict.

Evidence from real runs

Run 38173405739 (routed, legacy detection engine):

Agent job. Routed to gpt-5.6-luna via /responses:

  • The agent's api-proxy model-routing.jsonl has 6 events: classification, selection, and 4 request events.
  • The agent's token-usage.jsonl has one classifier record with purpose: routing_classification.
  • The agent's AWF resolved config contains a model-routing block.

Detection job. Its own api-proxy token-usage.jsonl (also in usage/detection/token_usage.jsonl):

  • 2 records, both model claude-haiku-5.5, path /chat/completions, status 200, schema token-usage/v0.28.51.
  • No purpose field.
  • No model-routing.jsonl.
  • detection_usage.json reports primary_model: claude-haiku-5.5.

Unified session. The conclusion job's usage/aw_session.jsonl includes the detection records as firewall.token_usage events. Their provenance is phase: detection and path: threat-detection/sandbox/firewall/logs/api-proxy-logs/token-usage.jsonl.

Run 38174054652 (external engine, Copilot, not routed):

  • detection_result.json has the keys prompt_injection, secret_leak, malicious_patch, reasons, warnings. "Conclude threat detection" parsed it.
  • The detection token-usage.jsonl has 2 records, both claude-haiku-5.5 on /chat/completions, with no purpose field.
  • The detection's awf-resolved-config.json has no model-routing block.
  • There is no model-routing.jsonl.

Two traps the checks must avoid:

  1. execution.json is not completion evidence. It stays {"state":"started"} after a successful run, for both the agent and detection. run_awf_with_startup_retries.sh only rewrites it to not_started when AWF fails before the harness starts.
  2. The "configured detection model" is an alias that overlaps the routed models.
    • The detection step runs with model detection.
    • In the detection models.json, that alias resolves through small → mini → haiku / gpt-5-mini / gpt-5-nano / gemini-flash-lite, using *haiku* and gpt-5*mini* globs.
    • The routed smoke's allowed-models is gpt-5.4-mini, gpt-5.6-luna, claude-haiku-4.5. The alias can therefore legitimately resolve to an allowed model. Today it resolves to claude-haiku-5.5.
    • A check like "the detection model is not a routed model" would be ambiguous or flaky unless the detection model is pinned.

Where checks can run

  • Detection job, via safe-outputs.threat-detection.post-steps. These compile into the detection job right after "Copy detection firewall logs" and before parsing, upload, and "Conclude threat detection". At that point the following evidence is local to the runner:
    • the detection result file
    • the detection's own api-proxy logs; the agent's firewall files were removed by "Clean stale firewall files from agent artifact"
    • the detection's AWF audit/config files
    • the gh-aw setup scripts, including the helper, under ${RUNNER_TEMP}/gh-aw/actions
  • Not a custom job. A custom jobs: entry that needs: [detection] fails to compile, because custom jobs run before the agent and this creates a cycle.
  • Not the conclusion job. Built-in jobs accept only setup-steps, pre-steps, needs, and if. Checks can't be appended to the conclusion job after it builds the unified session.

Proposed fix

In smoke-copilot-routed.md only:

  1. Remove gh-aw-detection: false so the default external engine runs. I found no blocker:

    • The workflow compiles cleanly with the flag removed.
    • The compiled detection job contains no routing config.
    • External Copilot detection is green on main elsewhere.

    This smoke is the cheapest and steadiest of the three: all 5 runs since the assertion fix were green, it uses the same Copilot CLI detector, and Copilot is the most-used external engine on main. Leave the SDK and Pi smokes unchanged.

  2. Pin safe-outputs.threat-detection.model to a concrete model outside the smoke's allowed-models, for example claude-haiku-5.5, the model observed in both runs above.

Add detection checks D1–D3 as safe-outputs.threat-detection.post-steps in the detection job. They must run even when earlier steps fail. They read only runner- and proxy-written evidence for the detection run, never any model reply.

  • D1. Detection completed and produced its result.
    • The detection execution step succeeded.
    • The structured result file exists and parses, with the verdict keys shown above.
    • Do not use execution.json as completion evidence.
  • D2. Every detection request used the configured detection model.
    • The detection's own api-proxy token-usage.jsonl has at least one model request.
    • Every request's served model equals the pinned detection model.
    • None of them is in the smoke's allowed-models.
    • Apply the same disjointness idea as the helper's classifier rule.
  • D3. Detection made no routing request and has no routing selection.
    • No detection token-usage record has purpose equal to routing_classification.
    • The detection's api-proxy logs have no model-routing.jsonl, or it has no selection events.
    • The detection's AWF resolved config has no model-routing block.
  • D4. The agent's routing checks still pass.
    • The existing agent-job post-step (E1, E2, R1–R4) stays unchanged and must still print PASS.

Implementation notes

  • Reuse actions/setup/js/smoke_model_routing_assertions.cjs where it fits:
    • servedModel, normalizeEndpoint, CLASSIFIER_PURPOSE, and the JSONL-reading and path conventions in TOKEN_USAGE_PATHS.
    • Add a detection-evidence evaluator next to the existing ones, with the same PASS/FAIL line format. Don't write a separate parser.
  • If a D-check fails, the detection job goes red and safe outputs are skipped. That is the intended, visible failure.
  • Don't depend on step-summary.md in the detection artifact. reference/artifacts.md lists it for the external engine, but the compiled upload step does not include it.

Acceptance cases

  1. Green dispatched run on main.
    • Input: dispatch smoke-copilot-routed.lock.yml on main after merge.
    • Result: the run is green.
    • The agent job prints the existing E1, E2, and R1–R4 PASS lines.
    • The detection job prints D1–D3 PASS lines, and its "Execute threat detection with AWF" step ran (the external engine, not "Execute GitHub Copilot CLI").
  2. Audit separates detection from routing.
    • Input: gh aw audit <run-id> on that run.
    • Result: the output shows the detection run on its pinned model, separate from the routed agent model.
    • Note: this does not work today. For run 38173405739, the audit from a main build (3667115) reports only gpt-5.6-luna under firewall_token_usage.by_model/agent_usage. It does not break out detection, even though it downloads detection_usage.json and the unified session has phase: detection records.
    • If the audit still lacks this when you implement, add only the minimal audit change needed to show the detection model and request count. It must be separate from the routed model.
  3. Negative unit cases from real record shapes.
    • Input: fixtures copied from real detection records, with request IDs and timestamps redacted. Don't invent fields.
      • Cite run 38173405739 (detection token-usage.jsonl, schema token-usage/v0.28.51).
      • For the external result file, cite run 38174054652.
    • Results:
      • D2 fails when a detection request's model is the routed model, for example gpt-5.6-luna.
      • D3 fails when the detection token usage contains a routing_classification record.
      • Shape the classifier record like the agent-side classifier record from run 38173405739.
    • The matching positive fixtures pass.
  4. Green PR-branch run before review.
    • Input: before requesting review, run gh workflow run smoke-copilot-routed.lock.yml --ref <branch>.
    • Result: a green run with D1–D3 and R1–R4 PASS lines, linked in the PR description.

Scope

  • Keep the PR to this one fix. Do not change files or behavior unrelated to it.
  • Before fixing a failing check, compare it with the same check on main.
    • If it also fails on main, do not fix it in this PR.
    • Merge main once main is fixed.
    • In the PR, note which checks are known failures on main.
  • This applies even if a bot comment asks you to fix a failing check. Fix only failures caused by this PR's changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions