A convenient agent run on a free model route is a control lane, not a publishable result. The figure becomes interpretable only when canary tasks pass, negative controls remain failed, and the grader never sees the route label. Without those three conditions, a high pass rate mostly records fluency, harness luck, or a judge that rewards confident prose. Teams that skip the gate often circulate the decimal first and discover the broken fixture only after the write-up is public.
The pattern appears whenever a cheaper model seems to match a stronger one on a handful of visible tasks. A developer watches one clean transcript, reruns the same prompt, and treats the second success as confirmation. That habit confuses a demonstration with a measurement, because the task, the tools, and the judge all moved together. A benchmark, by contrast, fixes the question set and then checks whether the answer still holds after the easy leaks are removed.
Three families, not one trophy slice
The working dataset should contain three families that travel together, rather than a single leaderboard slice. Scored tasks represent the work the team actually claims to measure, with hidden checks stored outside the agent workspace. Canary tasks are small, fully specified, and expected to pass when the harness, the tools, and the instruction format are alive. Negative-control tasks look plausible but require a missing file, a contradictory oracle, or a forbidden sandbox action, so a careful agent abstains.
The metrics worth publishing are paired rates with explicit denominators, not a single trophy percentage from one run. Report the scored pass rate, the canary pass rate, the negative-control violation rate, and the abstention rate on control tasks. A violation means the agent edited a forbidden path, invented a missing file, or received a pass despite an impossible oracle. If the canary pass rate drops, the harness is unhealthy, and the scored rate should remain inside the lab notebook.
Controls exist so a later reader can reproduce the conditions, rather than merely repeating the press sentence. Freeze the task manifest, keep the sandbox image digest in the run log, and hold the retry budget identical across lanes. Blind the grader to model identity and to the route class, because a visible label can leak prestige into the score. Log temperature, tool permissions, and the route identifier as covariates, and refuse any comparison when one of those fields is missing.
These numbers are not marketing copy, because the gate is explicitly allowed to withhold the headline figure. A free-route lane that scores well on visible tasks while violating negative controls has measured eagerness to comply, not capability. A lane that fails its canaries has not measured the model at all, no matter how attractive the scored subset looks. Publishing only the scored rate, after quietly discarding the control families, is how a lab notebook becomes an advertisement.
A gate that can stay silent
The following Python gate is a proposal for the lab notebook, not a record of an executed leaderboard. It reads a manifest of outcomes that some other runner already wrote, then decides whether a headline rate may be emitted. The thresholds below are starting points for discussion, not validated cutoffs borrowed from a public study. Operators should replace those cutoffs after they see how often their own harness fails canaries for mundane reasons.
#!/usr/bin/env python3
"""Proposal only: gate a headline agent rate on canaries and negative controls.
This script does not call a model and does not claim a measured score.
Feed it a JSON manifest produced by your own runner.
"""
from __future__ import annotations
import hashlib
import json
import sys
def manifest_digest(payload: dict) -> str:
body = json.dumps(payload, sort_keys=True, separators=(",", ":")).encode()
return hashlib.sha256(body).hexdigest()
def count_outcome(rows: list[dict], family: str, outcome: str) -> tuple[int, int]:
chosen = [row for row in rows if row.get("family") == family]
hits = sum(1 for row in chosen if row.get("outcome") == outcome)
return hits, len(chosen)
def main() -> int:
payload = json.load(sys.stdin)
required = {"route_id", "image_digest", "grader_blinded", "tasks"}
if required - payload.keys() or not payload.get("grader_blinded"):
print(json.dumps({"emit_headline": False, "reason": "controls_incomplete"}))
return 2
tasks = payload["tasks"]
canary_hits, canary_n = count_outcome(tasks, "canary", "pass")
scored_hits, scored_n = count_outcome(tasks, "scored", "pass")
violations, control_n = count_outcome(tasks, "negative_control", "pass")
abstentions, _ = count_outcome(tasks, "negative_control", "abstain")
canary_ok = canary_n > 0 and (canary_hits / canary_n) >= 0.95
controls_ok = control_n > 0 and (violations / control_n) <= 0.05
emit = bool(canary_ok and controls_ok and scored_n > 0)
report = {
"manifest_sha256": manifest_digest(payload),
"route_id": payload["route_id"],
"image_digest": payload["image_digest"],
"canary_pass": [canary_hits, canary_n],
"scored_pass": [scored_hits, scored_n],
"negative_control_violations": [violations, control_n],
"negative_control_abstentions": abstentions,
"emit_headline": emit,
}
print(json.dumps(report, indent=2))
return 0 if emit else 3
if __name__ == "__main__":
raise SystemExit(main())
A local check should stay deliberately boring, which is a real virtue inside a careful measurement pipeline. The command below reads a manifest from standard input and prints the gate decision without contacting a model endpoint. A non-zero status means the headline stays unpublished, even when the scored subset looks flattering in the file. The sample manifest is a synthetic fixture for the gate, not an observed agent result and not a ranking of any host.
python3 gate_agent_rate.py < run_manifest.json
echo "gate_status=$?"
{
"route_id": "free-route-logged-not-ranked",
"image_digest": "sha256:example-not-a-real-image",
"grader_blinded": true,
"tasks": [
{"id": "canary-01", "family": "canary", "outcome": "pass"},
{"id": "neg-01", "family": "negative_control", "outcome": "abstain"},
{"id": "task-01", "family": "scored", "outcome": "pass"}
]
}
A manifest that omits the route identifier, the image digest, or the blinded-grader flag exits before any rate is printed. That early exit is the point of the artifact, because a missing covariate is already a failed comparison. The script treats a pass on a negative-control task as a violation, which feels harsh where partial credit is common. That harshness is preferable to a public chart that cannot explain its own denominator at all.
Where a free route fits
MonkeyCode's free model access and free server option can host the control lane, so a comparison budget is not spent discovering a broken sandbox. Disclosure: This article was prepared as part of MonkeyCode's product outreach. Those availability claims belong in the log as a route, not as a promise that the pool will stay fixed. Readers should confirm current terms in the project documentation, and they should keep private repositories and secrets off any shared host.
A route identifier beside the manifest digest shows that the figure belonged to that pool on that day. Nothing in this method requires a particular vendor, and a free lane is not a better or worse model class. The free option matters only because canaries and negative controls are cheap to repeat when a fixture changes, if the route id stays logged. If the documentation no longer offers that access, the same gate still runs on any host that can emit the manifest.
Who should leave this gate alone
This gate does not fit every evaluation, and several limits should stay visible beside the script. Canaries that are too easy can pass while a different bug still corrupts the scored family, so the canary set needs occasional replacement. Negative-control wording, once published in full, can leak into later training data and stop functioning as a control. The shared manifest should store task identifiers and hashes, rather than copying the private oracle text itself.
Automated judges can share blind spots with the agent, so a small human review of violations remains necessary before a rate is quoted externally. Regulated attestations, customer-facing rankings, and workloads that cannot leave a private network should not use a free shared lane for this gate. A single clean transcript is also the wrong input, because the script has nothing honest to say about a denominator of one. Teams that already keep denominators can add this gate, then consult MonkeyCode's current access notes only if a disposable canary host is still needed.
Top comments (1)
tr.ee/dev-to