Summary
The matcher uses token-F1 + bigram recall only. Extracted triggers ('when git push fails with non-fast-forward') don't match reference trajectories that express the same failure differently ('Updates were rejected because the remote contains work'). Result: recall 0.02-0.04 (target ≥0.10) and golden pass rate 1/10 (target ≥70%) on ALL models. The #489 reference expansion added 288 paraphrases, but token-F1 can't exploit them.
Evidence
- reference-expansion recall: 0.020-0.039 across all 4 models (target ≥0.10)
- golden 1/10 on all 4 models (cloud == local — not model capability)
- 60-70% of extracted candidates are inconclusive because trigger ≠ reference at threshold, even though the trigger is CORRECT
Root cause
Lexical-only matching. The diverse phrasings added to the reference corpus (#489) require semantic/paraphrase awareness the current scorer lacks.
Proposed design
Add an embedding-based similarity layer to match_score():
- Load a lightweight local sentence embedder (sentence-transformers/all-MiniLM-L6-v2, ~22MB)
- score = 0.4 * token_f1 + 0.3 * bigram_recall + 0.3 * semantic_cosine(trigger, haystack)
- Cache embeddings per trajectory (compute once per corpus load)
- Offline fallback to current token-F1 when embedder unavailable
Impact
Highest — directly addresses the sole v0.3.0 release blocker (quality gate + recall). This is the v0.4.0 lever identified in FIELD_TEST_REPORT.md Conclusions.
References
Summary
The matcher uses token-F1 + bigram recall only. Extracted triggers ('when git push fails with non-fast-forward') don't match reference trajectories that express the same failure differently ('Updates were rejected because the remote contains work'). Result: recall 0.02-0.04 (target ≥0.10) and golden pass rate 1/10 (target ≥70%) on ALL models. The #489 reference expansion added 288 paraphrases, but token-F1 can't exploit them.
Evidence
Root cause
Lexical-only matching. The diverse phrasings added to the reference corpus (#489) require semantic/paraphrase awareness the current scorer lacks.
Proposed design
Add an embedding-based similarity layer to match_score():
Impact
Highest — directly addresses the sole v0.3.0 release blocker (quality gate + recall). This is the v0.4.0 lever identified in FIELD_TEST_REPORT.md Conclusions.
References