What is a GTM engineer, and what tools and data providers do they need in 2026?
A GTM engineer turns a company’s sales and marketing playbook into systems that run every week: the ICP written as filters, lists built and enriched, signals that say who to contact now, research for every first line, and routing for inbound, with a person approving what goes out. GTM engineering is the practice; the GTM engineer is the person who owns it. The job sits between RevOps, sales and engineering. What changed in 2026 is where it runs: more of it now runs from an agent such as Claude Code or Codex, where you describe the job in a prompt and the agent makes the calls.
What a GTM engineer does in a week
- Write the ICP down from the deals the team actually won and kept, as fields a provider can filter on (chapter 1).
- Build and enrich lists cheaply: qualify on the fields you already have, then pay for people, emails and verification only on the rows that pass (chapters 4 to 8).
- Watch for timing: hiring, funding, job changes and posts, each with a link and a date, scored against fit (chapters 9 and 10).
- Hand reps the reason to write, with its source, and keep a person on the send button (chapter 11).
- Keep the data honest: verify before sending, re-check people who just moved, and test a signal before it gets a weight. Our job-change study shows why.
- Measure and roll out: tag every row so pipeline traces back to a list, a signal and a provider, and automate one step at a time (chapters 14 to 16).
GTM engineer, RevOps or sales ops?
Titles vary and the work overlaps, and in many teams GTM engineers sit inside RevOps. For the pipeline work in this playbook, this is a useful way to divide it:
| Role | Typical focus | Example measures |
|---|---|---|
| GTM engineer | Automated workflows: lists, enrichment, scoring, signals, routing, research for outreach | Qualified pipeline per hour and per dollar of data |
| RevOps | Revenue processes, systems, planning, data governance and reporting across teams | Data quality, forecast reliability, revenue efficiency |
| Sales or marketing ops | Tool admin, sequences, campaigns and lead handoff | Tools that run and leads that reach the right rep |
| SDR | Conversations: first touches, follow-ups, booked meetings | Meetings held |
In small teams one person holds several of these. Roughly: RevOps decides how the revenue system should work, and a GTM engineer builds the automated parts of it.
When GTM engineering helps, and what the work takes
The workflows in this playbook pay off once the ICP can be written down and someone is doing the same research by hand every week. A small workflow can be owned by someone already on the team; a dedicated role makes sense when the volume and value justify it. Either way, the skills are less about code than about precision:
- Writing rules a machine can follow: what qualifies, what disqualifies, what the agent may do alone.
- Data judgement: cost per correct result rather than price per row, and knowing which fields go stale (found is not deliverable).
- Testing before trusting: a timing signal earns a weight only after it separates your own wins from losses (our test).
- Enough sales sense to know a weak reason to write when the agent produces one.
- Comfort with an agent, a CLI and a spreadsheet. Little code is needed when the agent makes the calls.
The stack has five layers. Only one of them is treg.
set up treg — https://treg.to/llms.txt
Define before you build
Three decisions every later chapter inherits. Skip them and the agent will be fast and wrong.
Write your ICP as fields a provider can filter on
An agent cannot search for “values data”. It can search for country, headcount band, category, funding and technology. Anything you cannot turn into one of those becomes a check the agent runs on the rows afterwards (chapter 5), not a filter.
- Split the ICP into filters (fields providers index) and checks (judgement on each row).
- Write the buyer as a function (“marketing or growth”), not a title, so chapter 6 can search for it.
- Keep the exclusions explicit: agencies, consumer apps, companies that were acquired.
ICP: the one we ran this playbook on
filters country = US · employees = 51-200 · category = B2B and SaaS
checks sells to sales, marketing, revenue or growth teams, or runs outbound itself
would buy lead data, enrichment or prospecting APIs
buyer function = marketing or growth (Head of Marketing, Head of Growth)
exclude agencies, services firms, consumer apps, hardware, acquired companiesUsing treg, turn the ICP in icp.md into filters for two company-search providers. Use their free count endpoints only. For each provider show the exact filter object and the count it returns. Do not return any rows yet.
Count the market for free before you pay for a single row
Several providers will count matches for free. Use that before anything else, and use two of them: the same ICP, expressed in each provider’s vocabulary, returns very different markets.
series_a or an industry label can quietly shrink a market by a factor of 400.- Count on two providers. A gap of 2× is normal; 100× means one filter is not doing what you think.
- Pull 10 sample rows from each (cheap) and read them before you trust either number.
- Size the adjacent bands too. Your real market is often the band next door.
Using treg's free company-count endpoints, count my ICP on two providers and in the headcount bands either side of it. Show every filter you used. If two counts differ by more than 3x, tell me which filter is likely wrong before we pay for anything.
Write down the rules the agent follows, and who can stop it
What makes a lead good enough? What evidence moves it forward? When can Claude act? LinkedIn
An agent acts at volume on whatever its instructions leave open. A written rule is still an instruction, not a guarantee, so it needs a number, a log and an owner.
icp_check pass if fit >= 0.5, judged on the fields the row already has (one cheap enrichment first if it has too few) people, email and news steps only on rows that passed icp_check drops log every dropped row with its reason email send only to verifier = deliverable; unknown and catch-all go to a separate list first_line a person approves every first line before a sequence starts spend_cap stop and ask if a run will cost more than $5 owner <name> can pause every sequence with one command
Go deeperThe run the thresholds came from
Build the list
How GTM engineers use Claude Code to build and enrich lead lists: candidates first, a check on the fields you already have, then expensive lookups only where they can pay back.
Lookalikes of your best accounts are candidates, not leads
We took the three best-fit accounts from the 23 Sep run as seeds, asked for 25 lookalikes each in the same country, enriched every one and ran the same ICP check as chapter 5.
treg.companies.enrich ($0.20 for 71, served by four providers cheapest-first). Lookalike rows arrive with too few fields to judge, so this one cheap enrichment came before the check. The check was jev on the team’s own key, judging fit at 50%; it does not enforce size, and 3 of the 8 passes were outside 51–200 staff.- Seed with 3 to 5 accounts that actually closed, not the logos you wish you had.
- Constrain the lookalike by country and size if the provider allows it; we only constrained country, and most results were smaller companies.
- Enrich once, cheaply, so there are fields to judge; then enforce your size filter and run the same check as any other list.
Using treg, find 25 lookalikes for each of these customer domains, same country and same size band. Enrich each lookalike, then run my ICP check from rules.md on it. Return only the ones that pass, with the reason each failed one was dropped.
Go deeperCompany enrichment comparedFindymail
Qualify on the fields you already have before any expensive lookup
I rarely enrich everything. r/gtmengineering
- Judge every row on the fields the list already carries: description, industry, size, funding.
- Drop the rows that fail and keep the reason next to them.
- Run finders, verifiers and news only on what passed; use your own provider keys where you already pay.
Using treg, list 50 US software companies with 51 to 200 staff that raised a Series A. Before you find anyone, judge each company against my ICP from the list fields alone, drop anything below 50%, and show me the price of every paid step before you run it.
Which people search APIs work inside an AI agent? Search by role, not “decision makers”
On 8 accounts that passed chapter 5, we asked a “decision makers” endpoint for the committee, then asked a routed people search for the buyer’s function directly.
treg.people.search with title: marketing, capped at 3 rows and $0.20 a call; it was served by QuickEnrich on 7 accounts and Dropleads on 1. Contact details are a separate, paid step (chapter 8).- Map the committee yourself: who pays, who champions, who uses it. Write each as a function.
- Search each account for those functions, cheapest rows first, then pick the most senior per function.
- Reveal contact details only for the people you will actually write to.
Using treg, for each passed account search people by company domain with title "marketing" and then "growth", at most 3 rows each and $0.20 per call. Return name, title and a seniority guess. Do not reveal emails yet.
Go deeperPeople search APIs comparedClaude for people searchPeople Search BenchQuickEnrich
The cheapest way to run waterfall email enrichment from an AI agent
Half the emails bounce now. r/gtmengineering
Coverage depends on the segment, so the only honest test is your own rows. And the number that matters is cost per correct result, because a cheap provider that misses is expensive.
- Take 20 to 50 rows of your real list and run them through a routed finder, which tries providers cheapest first; misses on per-success providers are not billed.
- Record which provider found each email, and compare cost per correct, not price per row.
- Widen adjacent titles before switching vendors; thin results are often a title filter.
Using treg, find work emails for these 30 people with the routed email finder, then verify each one. Report found, verified and cost per verified email, and which provider found each address.
Find, then verify, and keep the unknowns apart
Half our “valid” contacts were catch-alls. r/gtmengineering
Those companies had passed a fit check first. In a separate sample built from title searches alone, the gap between found and deliverable was much wider. We asked the routed finder for each person’s email and verified every address it returned.
- Verify as its own step and store the verdict next to the address. “Found” only means a finder returned an address; in the title-search sample above, 53% of found addresses verified deliverable.
- Send only to deliverable. Unknown and catch-all go to a smaller, slower list, or nowhere. A blank catch-all field means unknown, not safe.
- Watch bounces and complaints per sending domain.
Go deeperEmail verifiers comparedKittTomba
Know when to reach out
How to set up signal-based outbound with an AI agent: signals you can check, and a score that combines them with fit.
How to set up signal-based outbound: use signals you can open and date
Another alert feed everyone ignores? r/sales
An intent score cannot be checked, so reps discount it. A job post, a hire, a funding announcement or a post about the problem can be opened, dated and quoted in the first line.
treg.companies.jobs.search (served by PredictLeads, $0.04 a call, up to 25 postings each); funding from Aviato ($0.01 a call; 26 of 27 had rounds on record, the latest returned from May 2025). $1.35 for both checks. Open hiring was the live signal; funding data was stale for this list.Job changes: most contact records do not show the new job yet
A champion who moves to a new company is one of the best reasons to write. The trap is the data: in the first weeks after a move, most of the stored records we checked did not show the new company, so the email can go to an inbox they no longer read and the first line names the wrong company.
- Keep only signals with a source link and a date, and drop closed job postings: more than half of those returned were closed.
- Write one line of “why now” per signal. If you cannot, it is noise.
- For a job change, confirm the new company with a live profile read before anything else. A second database helps little: when two both had the person, both missed the new job 64% of the time.
- Give the feed an owner who triages it on a fixed day.
npx skills add superdesigndev/treg --skill lead-signals
Using treg, for each person in job-changes.csv read their live LinkedIn profile and return current company, title and start date. Compare with the company in our CRM and flag every mismatch. For mismatches only, find and verify an email at the new company. Show the cost before you start.
Go deeperClaude for lead signalsThe skillPredictLeadsLive profile reads
Score fit and timing together, then work the top tier first
Timing signals deserve suspicion before they get a weight. We tested a popular one: can you see a startup’s next round coming in its hiring or its news? Hiring could not tell them apart, and news ran the wrong way.
| Signal | Raised next | No round found |
|---|---|---|
| Opened a GTM role | 30% | 24% |
| Opened two or more GTM roles | 19% | 15% |
| Opened any job | 47% | 48% |
| Opened a senior role | 18% | 19% |
| Hiring sped up | 25% | 22% |
| Was in the news | 26% | 54% |
- Keep fit and timing as two numbers. Multiply them only to sort, never to decide.
- Define tiers with numbers, and write them in rules.md.
- Re-score weekly; timing decays, fit barely moves.
- Before a timing signal gets a weight, check it on your own past deals: the share of wins that showed it against the share of losses. If it does not separate them, use it as a reason to write, not as a score. For funding, act on the announcement itself.
Using treg, for each passed account pull open jobs and funding rounds. Count sales, marketing and growth roles that are still open and were posted or first seen in the last 60 days. Tier A if fit >= 0.6 and 2+ such roles, B if 1+, else C. Show the tier, the roles and their links.
Using treg, take wins.csv and losses.csv (domain and close date). For each account pull job postings and news from the 90 days before its close date. For each signal, report the share of wins and the share of losses that showed it, and tell me which signals separate them by more than 15 points. Show the cost before you start.
Go deeperLead signalsAviato
Reach out
The agent finds the reason to write. A person decides whether it is good enough to send.
Let the agent research. Let a person write, or at least approve.
Cold outreaches always come off super robotic. r/sales
- The agent gathers the reason to write, with its source.
- Score the reason before anyone writes. Drop weak reasons; do not polish them.
- A person writes or approves the first line. Read ten out loud before a sequence goes live.
Sending infrastructure and replies: what practitioners recommend
I was spending more time on infrastructure than on selling. r/b2bmarketing
Sending
Inboxes, domains and warm-up belong to your sending tool. The advice that held up across threads: conservative volume, and watching bounces and complaints per domain. Absolute rules (“never use links”) did not.
Replies
Classify every reply as interested, objection, referral or out-of-office; route it to the rep with the context the agent gathered; let a person approve the answer; and write the outcome back to the row (chapter 14).
Inbound
The leads that come to you deserve the fastest answer, and most of the routing needs one lookup.
Enrich an inbound lead from its email domain and route it in seconds
A work email gives you a domain, and a domain gives you most of what routing needs. We enriched 20 company domains the way a form handler would.
treg.companies.enrich on 20 domains not looked up before; served by TheCompaniesAPI 13, Hunter 5, Dropleads 2; slowest answer 4.8 s. A second set of 20 domains that had already been enriched came back in a median 1.15 s for $0.011 in total. This measures the lookup, not a full routing flow.- On submit, enrich the company from the email domain, and store the result so a repeat never pays full price.
- Route on two or three fields you trust: headcount band, country, industry. Send gaps to a person, not to nurture.
- Answer tier-A inbound within the hour; it already told you the timing.
treg call treg.companies.enrich --method POST --data '{"domain":"acme.com"}'Go deeperCompany enrichment comparedPerson enrichmentTheCompaniesAPI
Operate and improve
A playbook that is not measured decays. These three chapters keep it honest.
Tag every row, so you can tell which list, signal and provider paid off
Which positive replies became opportunities? Which meetings went nowhere, and why? LinkedIn
- Store per row: source list, signal, the provider that found the email, the verifier’s verdict.
- Monthly, join rows to replies, meetings and closed-won in the CRM.
- Cut the sources and providers that never appear on the winning side of that join.
Go deeperDownload the tagged CSV
The metric stack: system health, performance, efficiency
| Layer | Metric | Ours, from the runs |
|---|---|---|
| System health | Two market counts agree; enrichment fill rate on routing fields; share of verified-deliverable | One filter changed a count by over 400× (ch. 2); 16 to 20 of 20 routing fields filled (ch. 13); 20 of 21 deliverable (ch. 8) |
| Performance | Reply rate, meetings, opportunities, by source, signal and tier | Yours to measure: needs the chapter 14 tags |
| Efficiency | Cost per usable result; share of rows dropped before paying; time to first touch on inbound | $0.12 per deliverable lead; 21 of 48 dropped before paying (ch. 5); 2.2 s median to enrich a new domain (ch. 13) |
Go deeperAll workflows with receipts
Roll automation out like software: shadow, small segment, human gate, then autonomy
We usually find out days later. r/RevOps
- Shadow. Run the play next to the manual process for a week and compare outputs row by row.
- Small segment. Turn it on for one segment or one rep.
- Human gate. A person approves each send until the error rate is known.
- Autonomy only for the steps that passed; keep the gate on first lines.
What broke in our own runs
- Alert on expected counts per stage per run, not on errors.
- Store raw provider responses for a week, so a bad parse can be fixed without paying twice.
- Set a spend cap per run and have the agent stop and ask above it.
Go deeperAll workflows with receiptsProvider success rates in the catalog
The best Claude Code skills for GTM engineering, and the data step each leaves to you
A skill is the method. Most either ask for a vendor API key or leave live data to the agent. The last column is the catalog job that covers that step. We have not yet run these with treg end to end.
| Skill | Stars | GTM job | The data step | Covered by |
|---|---|---|---|---|
| coreyhaines31/marketingskills | 51.8k | Cold email, competitor profiling | Signals you supply; keyword and backlink data via DataForSEO | Buying signals; ranked keywords, backlinks |
| mvanhorn/last30days-skill | 63.1k | Research what people say about a topic | Social posts; optional ScrapeCreators or Apify keys | X, Reddit, TikTok, YouTube search and comments |
| AgriciDaniel/claude-seo | 17.9k | SEO audits and research | DataForSEO; Google OAuth for Search Console | SERP, keyword volume, backlinks; your own Search Console |
| zubair-trabzada/geo-seo-claude | 10.9k | AI visibility and brand mentions | Page fetches and heuristic scoring | ChatGPT, Perplexity and AI Mode answers with citations |
| phuryn/pm-skills | 26.6k | Competitive battlecards | Web search plus your win/loss notes | Company enrichment, news, pricing pages, ads |
| swan-gtm/gtm-skills | 161 | Account research briefs | Tool-agnostic company, people and search evidence | Firmographics, decision makers, hiring |
| superdesigndev/treg | ours | Buyer signals, UGC videos | Runs on treg directly | All of the above |
Stars read from GitHub on 29 Sep 2026. The six third-party repositories are MIT-licensed.
Clay alternatives for GTM engineers who work in Claude Code or Codex
If you already think in prompts, the question is not which table to use but where the data comes from and what it costs per correct result. Measured on the same 292 people:
| Option | How you work | Exact match | Per correct email |
|---|---|---|---|
| Claude Code or Codex + treg.to | Prompts; per-call data layer; your own keys first | 90.4% | $0.0056 |
| Clay | Visual table, shared workspace, credits | 89.7% | $0.0395 |
| Freckle | Table-based enrichment | 90.1% | $0.0427 |
| Deepline | Batch enrichment | 86.6% | $0.0924 |
Keep Clay if your team needs a shared visual table and already pays for it. Move the work into your agent when you want the whole playbook in prompts, paid per call. The full bench.
Every recorded run behind this playbook
| Run | Date | What came back | Metered |
|---|---|---|---|
| Playbook runs: counts, lookalikes, committee, timing, inbound (ch. 2, 4, 6, 9, 10, 13) | 30 Sep | 294 calls on one ICP; figures in each chapter | $2.70 |
| Study: job changes against five databases (ch. 9) | 7 Oct | 148 people; 68% of records did not show the new job | $11.10 |
| Study: hiring and news before a raise (ch. 10) | 6 Oct | 57 raised against 54 with no round found; hiring did not separate them | $13.65 |
| Study: found against deliverable (ch. 8) | 7 Oct | 60 found, 32 deliverable | $0.65 |
| Build a verified lead list (ch. 3, 5, 8, 11, 14) | 23 Sep | 20 deliverable leads from 50 companies | $2.33 |
| Work email bench (ch. 7) | 16 Sep | 292 people across five aggregators | $1.49 on treg |
| Find creators in a niche | 14 Sep | 25 creators from 153 matches | $0.20 |
| Screen creators before outreach | 14 Sep | 18 of 20 profiles, median engagement 5.3% | $0.036 |
| Price keyword demand | 14 Sep | 50 keywords, 742,970 monthly searches | $0.11 |
| Mine a competitor’s ads | 14 Sep | 20 Meta ads and 17 Google creatives | $0.015* |
The three studies are separate tests, each on its own sample, and are not part of the one-ICP runs; the hiring study’s figure includes building its two samples. The $5.03 in the header is the 30 Sep and 23 Sep runs, which follow one ICP; it includes $1.08 we spent re-running a step after our own parsing bug (ch. 16). * Two of that run’s calls used the team’s own Apify key, which treg never meters.
Glossary and questions
| ICP | Ideal customer profile, written as filters plus checks (ch. 1). |
| TAM | The count of companies that match your filters (ch. 2). |
| Check | A judgement on each row that no filter can express, run before paid steps (ch. 5). |
| Waterfall | Trying providers in order until one answers. treg’s routed endpoints try them cheapest first. |
| Catch-all | A domain that accepts any address, so a verifier cannot confirm a specific inbox (ch. 8). |
| Signal | A dated, linkable event that makes a message timely (ch. 9). |
| Live read | Fetching a person’s public profile at the moment you need it, instead of a stored database record (ch. 9). |
| Tier | A priority bucket from fit and timing (ch. 10). |
| Shadow mode | Running a play alongside the manual process to compare before switching (ch. 16). |
What is the difference between a GTM engineer and RevOps?
RevOps coordinates revenue processes, systems, planning and reporting across teams. A GTM engineer builds and runs specific workflows, such as enrichment, scoring, signals and routing, and often sits inside RevOps. Titles vary and the work overlaps; in practice RevOps decides how the revenue system should work, and a GTM engineer builds the automated parts of it.
How accurate is contact data after someone changes jobs?
Not very, in the first month. In our 7 Oct 2026 study of 148 people who had announced a new job 1 to 29 days earlier, 68% of contact database records on average did not show the new employer, and when two databases both had a person, both missed it 64% of the time. A live profile read was current for 47 of the 49 people it found, but it found only a third of them. For job-change plays, confirm the new company live before you write.
Do hiring or news signals predict that a startup is about to raise?
Hiring did not, in our test. Across 57 US startups that announced a seed to Series B round in August to October 2026 and 54 similar startups with no round found, the share that opened GTM roles, any job or senior roles in the weeks before was within a few points, and recorded news was more common in the comparison group with no round found (54% against 26%). Act on the funding announcement itself, and test any timing signal on your own wins and losses before you weight it.
Do I need to write code?
No. You describe the job in a prompt, the agent chooses the calls, and it shows the price before it spends. Writing the rules in chapter 3 is the part that needs you.
Does treg.to choose the data provider for me?
The catalog lists every provider of a job with measured success rates and prices, and your agent picks. Some jobs also have a routed endpoint (treg.<capability>) that tries providers cheapest first when you call it; the result names the provider that answered.
What does a miss cost?
It depends on the provider: some bill per successful result, some per call. Each catalog entry says which, routed endpoints do not bill misses on per-success providers, and every receipt shows what was actually charged.
Can I use the provider keys my team already pays for?
Yes. Register the key once and calls to that provider go through it. A team’s own key always wins over treg.to’s, and those calls are never metered.
Which agents does this work with?
The setup line works in Claude Code, Codex and Cursor; anything that can run a shell command can use the treg CLI. The prompts in each chapter are plain language and do not depend on one agent.