Skip to content

[v0.3.0·M7→v0.4.0] Semantic matching — embedding-based trigger↔reference similarity to bridge paraphrase gap #680

Description

@deghosal-2026

Summary

The matcher uses token-F1 + bigram recall only. Extracted triggers ('when git push fails with non-fast-forward') don't match reference trajectories that express the same failure differently ('Updates were rejected because the remote contains work'). Result: recall 0.02-0.04 (target ≥0.10) and golden pass rate 1/10 (target ≥70%) on ALL models. The #489 reference expansion added 288 paraphrases, but token-F1 can't exploit them.

Evidence

  • reference-expansion recall: 0.020-0.039 across all 4 models (target ≥0.10)
  • golden 1/10 on all 4 models (cloud == local — not model capability)
  • 60-70% of extracted candidates are inconclusive because trigger ≠ reference at threshold, even though the trigger is CORRECT

Root cause

Lexical-only matching. The diverse phrasings added to the reference corpus (#489) require semantic/paraphrase awareness the current scorer lacks.

Proposed design

Add an embedding-based similarity layer to match_score():

  • Load a lightweight local sentence embedder (sentence-transformers/all-MiniLM-L6-v2, ~22MB)
  • score = 0.4 * token_f1 + 0.3 * bigram_recall + 0.3 * semantic_cosine(trigger, haystack)
  • Cache embeddings per trajectory (compute once per corpus load)
  • Offline fallback to current token-F1 when embedder unavailable

Impact

Highest — directly addresses the sole v0.3.0 release blocker (quality gate + recall). This is the v0.4.0 lever identified in FIELD_TEST_REPORT.md Conclusions.

References

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions