Data steps tested on recorded runs · 3 new studies, 8 Oct 2026

The GTM engineering playbook
with Claude Code

Sixteen chapters, from defining your ICP to rolling automation out safely. Each one starts from a problem GTM engineers actually post about, then gives the play, a prompt to run in your agent, and the rule to keep, and for every data step, what happened when we ran it. We ran the data steps on one ICP so the numbers connect; the process chapters are marked Method. New on 8 Oct: three studies of our own, on contact data after a job change, whether hiring predicts a raise and found against deliverable emails.

16chapters + 4 appendices
6recorded runs on one ICP
$5.03total metered, all runs
502logged calls behind it
Find where to start Read the chapters
›
agent idle$0.00
An abbreviated replay of the recorded 23 Sep 2026 lead-list run (chapters 5, 8 and 11). Rows are illustrative; the counts are the run’s. The $2.33 total also covers two steps not shown: a news lookup ($0.84) and opener scoring.

Start here:
which of these sounds like your week?

We read 460 Reddit threads and 1,456 LinkedIn posts about GTM engineering. The problems cluster into the symptoms below. Tick the ones you recognise and the playbook tells you where to start.

ChaptersContents
Chapter 0 · Start Method

What is a GTM engineer, and what tools and data providers do they need in 2026?

A GTM engineer turns a company’s sales and marketing playbook into systems that run every week: the ICP written as filters, lists built and enriched, signals that say who to contact now, research for every first line, and routing for inbound, with a person approving what goes out. GTM engineering is the practice; the GTM engineer is the person who owns it. The job sits between RevOps, sales and engineering. What changed in 2026 is where it runs: more of it now runs from an agent such as Claude Code or Codex, where you describe the job in a prompt and the agent makes the calls.

What a GTM engineer does in a week

  1. Write the ICP down from the deals the team actually won and kept, as fields a provider can filter on (chapter 1).
  2. Build and enrich lists cheaply: qualify on the fields you already have, then pay for people, emails and verification only on the rows that pass (chapters 4 to 8).
  3. Watch for timing: hiring, funding, job changes and posts, each with a link and a date, scored against fit (chapters 9 and 10).
  4. Hand reps the reason to write, with its source, and keep a person on the send button (chapter 11).
  5. Keep the data honest: verify before sending, re-check people who just moved, and test a signal before it gets a weight. Our job-change study shows why.
  6. Measure and roll out: tag every row so pipeline traces back to a list, a signal and a provider, and automate one step at a time (chapters 14 to 16).

GTM engineer, RevOps or sales ops?

Titles vary and the work overlaps, and in many teams GTM engineers sit inside RevOps. For the pipeline work in this playbook, this is a useful way to divide it:

RoleTypical focusExample measures
GTM engineerAutomated workflows: lists, enrichment, scoring, signals, routing, research for outreachQualified pipeline per hour and per dollar of data
RevOpsRevenue processes, systems, planning, data governance and reporting across teamsData quality, forecast reliability, revenue efficiency
Sales or marketing opsTool admin, sequences, campaigns and lead handoffTools that run and leads that reach the right rep
SDRConversations: first touches, follow-ups, booked meetingsMeetings held

In small teams one person holds several of these. Roughly: RevOps decides how the revenue system should work, and a GTM engineer builds the automated parts of it.

When GTM engineering helps, and what the work takes

The workflows in this playbook pay off once the ICP can be written down and someone is doing the same research by hand every week. A small workflow can be owned by someone already on the team; a dedicated role makes sense when the volume and value justify it. Either way, the skills are less about code than about precision:

  • Writing rules a machine can follow: what qualifies, what disqualifies, what the agent may do alone.
  • Data judgement: cost per correct result rather than price per row, and knowing which fields go stale (found is not deliverable).
  • Testing before trusting: a timing signal earns a weight only after it separates your own wins from losses (our test).
  • Enough sales sense to know a weak reason to write when the agent produces one.
  • Comfort with an agent, a CLI and a spreadsheet. Little code is needed when the agent makes the calls.

The stack has five layers. Only one of them is treg.

AgentClaude Code, Codex, Cursor, Hermes or another agent that runs tools. It plans the steps and makes the calls.
Data layer · tregCompany and people search, enrichment, email finding and verification, hiring, funding and news, social posts, SEO and ad data. One key, priced per call, your own provider keys first.
SkillsThe method, written down: how to qualify, how to write the first line, how to score. Open source ones are in appendix A.
SendingYour inboxes and sequencer. Not treg (chapter 12).
CRMWhere outcomes live, so you can learn from them (chapter 14).
set up once, in your agent
set up treg — https://treg.to/llms.txt

Go deeperThe catalogSet up in Claude CodeTutorial

Part I

Define before you build

Three decisions every later chapter inherits. Skip them and the agent will be fast and wrong.

Chapter 1 · ICP Run 30 Sep

Write your ICP as fields a provider can filter on

SymptomYour ICP is a sentence in a deck (“mid-market B2B companies that value data”), so every list is built from a different reading of it.

An agent cannot search for “values data”. It can search for country, headcount band, category, funding and technology. Anything you cannot turn into one of those becomes a check the agent runs on the rows afterwards (chapter 5), not a filter.

  1. Split the ICP into filters (fields providers index) and checks (judgement on each row).
  2. Write the buyer as a function (“marketing or growth”), not a title, so chapter 6 can search for it.
  3. Keep the exclusions explicit: agencies, consumer apps, companies that were acquired.
worksheet · icp.md
ICP: the one we ran this playbook on
filters   country = US · employees = 51-200 · category = B2B and SaaS
checks    sells to sales, marketing, revenue or growth teams, or runs outbound itself
          would buy lead data, enrichment or prospecting APIs
buyer     function = marketing or growth (Head of Marketing, Head of Growth)
exclude   agencies, services firms, consumer apps, hardware, acquired companies
prompt
Using treg, turn the ICP in icp.md into filters for two company-search providers.
Use their free count endpoints only. For each provider show the exact filter
object and the count it returns. Do not return any rows yet.
The ruleIf a part of the ICP cannot be a filter, it is a check, and it runs before any expensive step.

Go deeperCompany research for agentsCompany data providers

Chapter 2 · Market size Run 30 Sep · free counts

Count the market for free before you pay for a single row

SymptomYou buy a list, and only later learn the filter meant something different to the provider.

Several providers will count matches for free. Use that before anything else, and use two of them: the same ICP, expressed in each provider’s vocabulary, returns very different markets.

One ICP, eight counts · US, 51–200 staff
CompanyEnrich: B2B + SaaS10,403
CompanyEnrich: B2B + SaaS, funded 2024+1,357
CompanyEnrich: B2B + SaaS, round = series_a25
Dropleads: “IT & Services”15,873
Dropleads: keyword “saas”5,395
Dropleads: “Computer Software”21
Apollo: Series A filter (paid page, 23 Sep)958
Recorded 30 Sep 2026. All eight provider counts were free (six here, two neighbouring bands below); Apollo’s number came with its paid first list page ($0.026) on 23 Sep. Each CompanyEnrich row adds one filter to the first row, not to each other. Neighbouring size bands on CompanyEnrich: 42,981 at 11–50 staff, 2,621 at 201–500. A tag like series_a or an industry label can quietly shrink a market by a factor of 400.
  1. Count on two providers. A gap of 2× is normal; 100× means one filter is not doing what you think.
  2. Pull 10 sample rows from each (cheap) and read them before you trust either number.
  3. Size the adjacent bands too. Your real market is often the band next door.
prompt
Using treg's free company-count endpoints, count my ICP on two providers and in the
headcount bands either side of it. Show every filter you used. If two counts differ by
more than 3x, tell me which filter is likely wrong before we pay for anything.
The ruleNo list is bought until two free counts roughly agree and ten sample rows look right.
treg doesFree count endpoints on several providers behind one key.
You decideWhich vocabulary matches your market, after reading the sample.

Go deeperCompany data providersCompanyEnrichDropleads

Chapter 3 · Rules From the 23 Sep run

Write down the rules the agent follows, and who can stop it

SymptomThe agent did something at volume that nobody would have approved one by one.What makes a lead good enough? What evidence moves it forward? When can Claude act? LinkedIn

An agent acts at volume on whatever its instructions leave open. A written rule is still an instruction, not a guarantee, so it needs a number, a log and an owner.

The threshold is a business decision
Pass at 50%27
Would pass at 60%14
Out of 48 companies in the 23 Sep run. Ten points on the threshold roughly halves the list.
worksheet · rules.md
icp_check      pass if fit >= 0.5, judged on the fields the row already has (one cheap enrichment first if it has too few)
people, email and news steps   only on rows that passed icp_check
drops          log every dropped row with its reason
email          send only to verifier = deliverable; unknown and catch-all go to a separate list
first_line     a person approves every first line before a sequence starts
spend_cap      stop and ask if a run will cost more than $5
owner          <name> can pause every sequence with one command
The ruleEvery threshold is a number in the prompt, every drop has a reason, and one person owns the stop.

Go deeperThe run the thresholds came from

Part II

Build the list

How GTM engineers use Claude Code to build and enrich lead lists: candidates first, a check on the fields you already have, then expensive lookups only where they can pay back.

Chapter 4 · Lookalikes Run 30 Sep · $0.20

Lookalikes of your best accounts are candidates, not leads

SymptomA lookalike list feels on-target because the seeds were your best customers, so it skips the qualification step.

We took the three best-fit accounts from the 23 Sep run as seeds, asked for 25 lookalikes each in the same country, enriched every one and ran the same ICP check as chapter 5.

71unique lookalikes from 3 seeds, $0.006
13in the 51–200 band; 53 had 50 staff or fewer
8passed the model check (11%); only 5 of them in the 51–200 band
Recorded 30 Sep 2026. Lookalikes from Findymail; enrichment through treg.companies.enrich ($0.20 for 71, served by four providers cheapest-first). Lookalike rows arrive with too few fields to judge, so this one cheap enrichment came before the check. The check was jev on the team’s own key, judging fit at 50%; it does not enforce size, and 3 of the 8 passes were outside 51–200 staff.
  1. Seed with 3 to 5 accounts that actually closed, not the logos you wish you had.
  2. Constrain the lookalike by country and size if the provider allows it; we only constrained country, and most results were smaller companies.
  3. Enrich once, cheaply, so there are fields to judge; then enforce your size filter and run the same check as any other list.
prompt
Using treg, find 25 lookalikes for each of these customer domains, same country and
same size band. Enrich each lookalike, then run my ICP check from rules.md on it.
Return only the ones that pass, with the reason each failed one was dropped.
The ruleA lookalike goes through the same size filter and ICP check as any other row.

Go deeperCompany enrichment comparedFindymail

Chapter 5 · Cost Run 23 Sep · $2.33

Qualify on the fields you already have before any expensive lookup

SymptomCredits burn faster than planned, because every finder, verifier and news call runs on every row.I rarely enrich everything. r/gtmengineering
Companies listed50
Usable domain48
Passed the ICP check27
Dropped before paying21
Verified deliverable20
≈ $4.12estimate if all 48 were enriched
$2.33metered · $0.12 per deliverable lead
The 23 Sep 2026 run. The check was jev on the free list fields at a 50% threshold, on the team’s own key. The $4.12 adds the run’s own estimate for the 21 dropped rows ($1.79); those rows were never enriched. Prompt, steps and CSV.
  1. Judge every row on the fields the list already carries: description, industry, size, funding.
  2. Drop the rows that fail and keep the reason next to them.
  3. Run finders, verifiers and news only on what passed; use your own provider keys where you already pay.
prompt
Using treg, list 50 US software companies with 51 to 200 staff that raised a Series A.
Before you find anyone, judge each company against my ICP from the list fields alone,
drop anything below 50%, and show me the price of every paid step before you run it.
The ruleNo expensive lookup (people, emails, news) runs on a row that has not passed the check. If a row lacks the fields to judge, one cheap enrichment comes first.
treg doesPrices every step before it runs, meters per call, and puts your own keys first, unmetered.
You decideWhat the check is. The agent applies it; it cannot invent your ICP.

Go deeperThe full lead-list runLead enrichmentApollo

Chapter 6 · Buying committee Run 30 Sep · $0.00

Which people search APIs work inside an AI agent? Search by role, not “decision makers”

SymptomYou email one senior person per account, and they are rarely the one who buys what you sell.

On 8 accounts that passed chapter 5, we asked a “decision makers” endpoint for the committee, then asked a routed people search for the buyer’s function directly.

“Decision makers” returned 45 senior people at 5 accounts
Sales, revenue, partnerships10
Ops, strategy, chief of staff7
Product, engineering, design6
Customer success, support5
People and HR5
Finance4
Founders and assistants4
Growth (no “marketing”)1
3 of 8accounts rate-limited (429) on the decision-maker endpoint
8 of 8accounts returned marketing people when searched by role
20people, names and titles, $0.00 metered
Recorded 30 Sep 2026, titles mapped to functions by hand; 3 were too vague to place, and no title contained “marketing”. A buying-group endpoint (Lusha) was out of capacity on treg’s key that day and charged nothing. The role search was treg.people.search with title: marketing, capped at 3 rows and $0.20 a call; it was served by QuickEnrich on 7 accounts and Dropleads on 1. Contact details are a separate, paid step (chapter 8).
  1. Map the committee yourself: who pays, who champions, who uses it. Write each as a function.
  2. Search each account for those functions, cheapest rows first, then pick the most senior per function.
  3. Reveal contact details only for the people you will actually write to.
prompt
Using treg, for each passed account search people by company domain with title
"marketing" and then "growth", at most 3 rows each and $0.20 per call. Return name,
title and a seniority guess. Do not reveal emails yet.
The ruleSearch for the buyer’s function. A “decision makers” list is a map of the company, not your committee.
treg doesRouted people search that tries providers cheapest first and names the one that answered.
You decideWho is on your buying committee. No endpoint knows what you sell.

Go deeperPeople search APIs comparedClaude for people searchPeople Search BenchQuickEnrich

Chapter 7 · Coverage and cost Bench 16 Sep

The cheapest way to run waterfall email enrichment from an AI agent

SymptomA vendor that works for someone else misses half of your market, and its per-row price hides it.Half the emails bounce now. r/gtmengineering

Coverage depends on the segment, so the only honest test is your own rows. And the number that matters is cost per correct result, because a cheap provider that misses is expensive.

Cost per correct work email · same 292 people
treg.to$0.0056
Clay$0.0395
Freckle$0.0427
Deepline$0.0924
16 Sep 2026, 88 companies, name and domain only. Exact-match rates were close: 90.4% treg.to, 90.1% Freckle, 89.7% Clay and 86.6% Deepline. Everyone had a public team-page email, so find rates run higher than on a cold list. Method and table.
  1. Take 20 to 50 rows of your real list and run them through a routed finder, which tries providers cheapest first; misses on per-success providers are not billed.
  2. Record which provider found each email, and compare cost per correct, not price per row.
  3. Widen adjacent titles before switching vendors; thin results are often a title filter.
prompt
Using treg, find work emails for these 30 people with the routed email finder, then
verify each one. Report found, verified and cost per verified email, and which
provider found each address.
The rulePick providers per segment from a test on your own rows, and re-test when the segment changes.
treg doesEvery provider of a job behind one key, and a routed endpoint that tries them cheapest first and names the one that answered.
You decideWhich provider wins for your segment; the catalog compares, your agent chooses.

Go deeperEmail finders comparedThe 292-person benchHunter

Chapter 8 · Deliverability Run 23 Sep Study 7 Oct

Find, then verify, and keep the unknowns apart

Symptom“Valid” from a finder turns out to mean the domain accepts everything, and the bounces land on your sending domain.Half our “valid” contacts were catch-alls. r/gtmengineering
20 deliverable1 unknown, kept separate0 invalid6 not found by either finder
27 named people: 21 emails found (18 by Hunter, 3 more by Kitt on Hunter’s misses), then verified. A third finder, Tomba, was out of capacity that day. The verifier returned no catch-all flag on any row. The receipt.

Those companies had passed a fit check first. In a separate sample built from title searches alone, the gap between found and deliverable was much wider. We asked the routed finder for each person’s email and verified every address it returned.

Study · 60 people from a title search: every email “found”, about half deliverable
32 deliverable16 risky4 unknown8 invalid
60 of 60people came back with an address from the finder
53%of those addresses verified deliverable
$0.65metered for the whole study, live checks included
Study run 7 Oct 2026. Method: 60 people from six US title searches covering sales, marketing, growth, revenue operations and GTM engineering, up to ten per search, with no industry filter. The routed email finder ran on every row, then the routed verifier on every address. “Deliverable” is the verifier’s verdict, not a measured inbox delivery. Of the 24 people with a usable live profile, 22 were at the listed company; we could not check the other 36. Results vary by segment, so run it on your own rows.
  1. Verify as its own step and store the verdict next to the address. “Found” only means a finder returned an address; in the title-search sample above, 53% of found addresses verified deliverable.
  2. Send only to deliverable. Unknown and catch-all go to a smaller, slower list, or nowhere. A blank catch-all field means unknown, not safe.
  3. Watch bounces and complaints per sending domain.
The ruleOnly a verified-deliverable address goes into the main sequence.
treg doesFinding and verification as separate calls, with the verifier’s verdict and a catch-all flag where the verifier returns one.
Not tregInboxes, warm-up and sending limits.

Go deeperEmail verifiers comparedKittTomba

Part III

Know when to reach out

How to set up signal-based outbound with an AI agent: signals you can check, and a score that combines them with fit.

Chapter 9 · Signals Run 30 Sep Study 7 Oct

How to set up signal-based outbound: use signals you can open and date

SymptomThe team has a signal feed and nobody acts on it.Another alert feed everyone ignores? r/sales

An intent score cannot be checked, so reps discount it. A job post, a hire, a funding announcement or a post about the problem can be opened, dated and quoted in the first line.

16 of 27accounts had open sales, marketing or growth roles posted or first seen in the last 60 days
355 of 675job postings returned were already closed
0 of 27funding rounds from the last 12 months returned by the funding provider
Recorded 30 Sep 2026 on the 27 accounts that passed chapter 5. Jobs through treg.companies.jobs.search (served by PredictLeads, $0.04 a call, up to 25 postings each); funding from Aviato ($0.01 a call; 26 of 27 had rounds on record, the latest returned from May 2025). $1.35 for both checks. Open hiring was the live signal; funding data was stale for this list.
# the shape of one row from the lead-signals skill person Head of Growth, <company> signal posted 3 SDR roles in the last 30 days why_now building outbound now; the list will need data source https://…/careers/sdr email found and verified, or blank

Job changes: most contact records do not show the new job yet

A champion who moves to a new company is one of the best reasons to write. The trap is the data: in the first weeks after a move, most of the stored records we checked did not show the new company, so the email can go to an inbox they no longer read and the first line names the wrong company.

Study · 148 people who announced a new job 1 to 29 days earlier
One database68%
Both of two64%
Live read4%
Share of records found that did not show the new employer. “Both of two” counts people two databases both had, where neither did.
72%still not showing it three to four weeks after the move (61% in the first week)
84 of 139people a database had: none of the databases that had them showed the new job
1 in 3people the live profile read could find at all
Study run 7 Oct 2026. Method: we took people who had posted that they were starting a new role 1 to 29 days earlier (median about 13), and used each post as the answer. We looked each person up once by LinkedIn URL in five contact databases and with one live profile read, and compared the company each returned with the announced one by name. 68% is the average of the five databases’ rates; 64% compares two databases on the people both had, where one alone missed 69%. A hand check of 40 non-matching records found every one named a different organisation, not a spelling of the new one; a few may be side roles held alongside the new job. The week groups are different people seen once, not records followed over time. The live read comes from the same profile the person updates, so its agreement is partly expected, and it found 49 of the 148. Everyone here announced their move publicly, so results may differ for people who do not. Databases are not named, and results will differ by database and segment.
  1. Keep only signals with a source link and a date, and drop closed job postings: more than half of those returned were closed.
  2. Write one line of “why now” per signal. If you cannot, it is noise.
  3. For a job change, confirm the new company with a live profile read before anything else. A second database helps little: when two both had the person, both missed the new job 64% of the time.
  4. Give the feed an owner who triages it on a fixed day.
install the signals skill
npx skills add superdesigndev/treg --skill lead-signals
prompt: check job changes live
Using treg, for each person in job-changes.csv read their live LinkedIn profile and
return current company, title and start date. Compare with the company in our CRM and
flag every mismatch. For mismatches only, find and verify an email at the new company.
Show the cost before you start.
The ruleNo source link, no signal. No “why now”, no outreach.

Go deeperClaude for lead signalsThe skillPredictLeadsLive profile reads

Chapter 10 · Priority Run 30 Sep Study 6 Oct

Score fit and timing together, then work the top tier first

SymptomEvery account on the list gets the same sequence on the same day, whether or not anything is happening there.
27 accounts, one tier each
A · fit ≥ 60% and 2+ open GTM roles8
B · at least 1 open GTM role8
C · nothing open now11
Recorded 30 Sep 2026. Fit is the chapter 5 ICP score (50% to 67% on these accounts); timing is open GTM roles posted or first seen in the last 60 days. Tier A gets a person and a first line this week; tier C waits for a signal.

Timing signals deserve suspicion before they get a weight. We tested a popular one: can you see a startup’s next round coming in its hiring or its news? Hiring could not tell them apart, and news ran the wrong way.

Study · what showed up before the reference date, for startups that raised and startups with no round found
SignalRaised nextNo round found
Opened a GTM role30%24%
Opened two or more GTM roles19%15%
Opened any job47%48%
Opened a senior role18%19%
Hiring sped up25%22%
Was in the news26%54%
Study run 6 Oct 2026. Method: 57 US startups that announced a seed to Series B round between mid-August and early October 2026 (reference date: the announcement), against similar startups whose last round was in 2025 and for which our source showed no 2026 round (reference date: the median announcement date). 54 of those 61 had any job or news data. Job postings by first-seen date (posted date if missing) and news by found date, from 90 to 7 days before the reference date; the final week is left out so the round itself does not count. GTM roles cover sales, marketing, growth, revenue operations, business development, partnerships and customer success. “Hiring sped up” means two or more postings and more than in the 90 days before that. One source, with incomplete coverage, and an unmatched comparison: gaps of a few points are noise, and differences in company age and coverage could explain the news gap; we did not establish the cause. This is no evidence of a hiring signal, not proof that none exists.
  1. Keep fit and timing as two numbers. Multiply them only to sort, never to decide.
  2. Define tiers with numbers, and write them in rules.md.
  3. Re-score weekly; timing decays, fit barely moves.
  4. Before a timing signal gets a weight, check it on your own past deals: the share of wins that showed it against the share of losses. If it does not separate them, use it as a reason to write, not as a score. For funding, act on the announcement itself.
prompt
Using treg, for each passed account pull open jobs and funding rounds. Count sales,
marketing and growth roles that are still open and were posted or first seen in the last
60 days. Tier A if fit >= 0.6 and 2+
such roles, B if 1+, else C. Show the tier, the roles and their links.
prompt: test a signal on your own deals
Using treg, take wins.csv and losses.csv (domain and close date). For each account pull
job postings and news from the 90 days before its close date. For each signal, report
the share of wins and the share of losses that showed it, and tell me which signals
separate them by more than 15 points. Show the cost before you start.
The ruleWork tier A this week. Tier C gets nothing until a signal moves it.

Go deeperLead signalsAviato

Part IV

Reach out

The agent finds the reason to write. A person decides whether it is good enough to send.

Chapter 11 · First line Run 23 Sep

Let the agent research. Let a person write, or at least approve.

SymptomThousands of emails go out and the replies do not come.Cold outreaches always come off super robotic. r/sales
In the 23 Sep run, 19 companies had a news event from the last year. jev scored each as a first line on a 0 to 3 scale: 4 of 19 were decent or better, mean 1.79. Most available “personalisation” was not worth sending.
  1. The agent gathers the reason to write, with its source.
  2. Score the reason before anyone writes. Drop weak reasons; do not polish them.
  3. A person writes or approves the first line. Read ten out loud before a sequence goes live.
The ruleNothing is sent that a person has not read.
treg doesThe research: news, hiring, posts, company data, each with a source.
Not tregWriting and sending. For the method, see the cold-email skill in appendix A.

Go deeperHow the opener scoring worksCompany research

Chapter 12 · Sending and replies Method · not treg

Sending infrastructure and replies: what practitioners recommend

SymptomMore time goes into inboxes and reply triage than into selling.I was spending more time on infrastructure than on selling. r/b2bmarketing

Sending

Inboxes, domains and warm-up belong to your sending tool. The advice that held up across threads: conservative volume, and watching bounces and complaints per domain. Absolute rules (“never use links”) did not.

Replies

Classify every reply as interested, objection, referral or out-of-office; route it to the rep with the context the agent gathered; let a person approve the answer; and write the outcome back to the row (chapter 14).

The ruleReplies are routed within a working day, and every outcome is written back to the row.
Part V

Inbound

The leads that come to you deserve the fastest answer, and most of the routing needs one lookup.

Chapter 13 · Inbound routing Run 30 Sep · $0.053

Enrich an inbound lead from its email domain and route it in seconds

SymptomA demo request sits in a queue while someone looks the company up by hand.

A work email gives you a domain, and a domain gives you most of what routing needs. We enriched 20 company domains the way a form handler would.

2.2 smedian to enrich a new domain (p90 4.5 s)
$0.053for 20 new domains, $0.0026 each
~⅕of the price, and about half the time, for a repeat lookup of the same domain
Employee count20/20
Location20/20
Industry16/20
Description13/20
Recorded 30 Sep 2026 with treg.companies.enrich on 20 domains not looked up before; served by TheCompaniesAPI 13, Hunter 5, Dropleads 2; slowest answer 4.8 s. A second set of 20 domains that had already been enriched came back in a median 1.15 s for $0.011 in total. This measures the lookup, not a full routing flow.
  1. On submit, enrich the company from the email domain, and store the result so a repeat never pays full price.
  2. Route on two or three fields you trust: headcount band, country, industry. Send gaps to a person, not to nurture.
  3. Answer tier-A inbound within the hour; it already told you the timing.
from a form handler or your agent
treg call treg.companies.enrich --method POST --data '{"domain":"acme.com"}'
The ruleEvery inbound lead is enriched and routed before a person opens it.

Go deeperCompany enrichment comparedPerson enrichmentTheCompaniesAPI

Part VI

Operate and improve

A playbook that is not measured decays. These three chapters keep it honest.

Chapter 14 · Attribution Run 23 Sep

Tag every row, so you can tell which list, signal and provider paid off

SymptomNobody can say which source produced last quarter’s pipeline.Which positive replies became opportunities? Which meetings went nowhere, and why? LinkedIn
companydomainpersontitleemailwhich provider found itverifier’s verdictcatch-all flagthe event to lead withopener score
The columns of the 23 Sep run’s CSV. The green ones make a later join to outcomes possible. Download it.
  1. Store per row: source list, signal, the provider that found the email, the verifier’s verdict.
  2. Monthly, join rows to replies, meetings and closed-won in the CRM.
  3. Cut the sources and providers that never appear on the winning side of that join.
The ruleA row without its source and provider does not go into a sequence.
treg doesNames the provider that answered each call and what it cost, so your agent can write both onto the row.
Not tregThe durable record and the CRM join. Keep provider and outcome on your own rows.

Go deeperDownload the tagged CSV

Chapter 15 · Metrics Method, with our numbers

The metric stack: system health, performance, efficiency

SymptomThe only number anyone tracks is “emails sent”.
LayerMetricOurs, from the runs
System healthTwo market counts agree; enrichment fill rate on routing fields; share of verified-deliverableOne filter changed a count by over 400× (ch. 2); 16 to 20 of 20 routing fields filled (ch. 13); 20 of 21 deliverable (ch. 8)
PerformanceReply rate, meetings, opportunities, by source, signal and tierYours to measure: needs the chapter 14 tags
EfficiencyCost per usable result; share of rows dropped before paying; time to first touch on inbound$0.12 per deliverable lead; 21 of 48 dropped before paying (ch. 5); 2.2 s median to enrich a new domain (ch. 13)
The ruleReview the three layers weekly, and change one rule at a time.

Go deeperAll workflows with receipts

Chapter 16 · Rollout Lessons from our runs

Roll automation out like software: shadow, small segment, human gate, then autonomy

SymptomAn automation quietly stops, or quietly does the wrong thing, and nobody notices for days.We usually find out days later. r/RevOps
  1. Shadow. Run the play next to the manual process for a week and compare outputs row by row.
  2. Small segment. Turn it on for one segment or one rep.
  3. Human gate. A person approves each send until the error rate is known.
  4. Autonomy only for the steps that passed; keep the gate on first lines.

What broke in our own runs

Tomba out of capacity (23 Sep)Lusha out of capacity (30 Sep)3 accounts rate-limited (429)our parser read the wrong field: $1.08 re-run
None of these threw an error that stopped the run. Each would have produced a quietly thinner list. The routed endpoints fell back to the next provider; the $1.08 was our own bug, caught because we stored the raw responses and compared counts.
  1. Alert on expected counts per stage per run, not on errors.
  2. Store raw provider responses for a week, so a bad parse can be fixed without paying twice.
  3. Set a spend cap per run and have the agent stop and ask above it.
The ruleNothing goes autonomous until it has run in shadow, and every stage reports its row count.

Go deeperAll workflows with receiptsProvider success rates in the catalog

Appendix A

The best Claude Code skills for GTM engineering, and the data step each leaves to you

A skill is the method. Most either ask for a vendor API key or leave live data to the agent. The last column is the catalog job that covers that step. We have not yet run these with treg end to end.

SkillStarsGTM jobThe data stepCovered by
coreyhaines31/marketingskills51.8kCold email, competitor profilingSignals you supply; keyword and backlink data via DataForSEOBuying signals; ranked keywords, backlinks
mvanhorn/last30days-skill63.1kResearch what people say about a topicSocial posts; optional ScrapeCreators or Apify keysX, Reddit, TikTok, YouTube search and comments
AgriciDaniel/claude-seo17.9kSEO audits and researchDataForSEO; Google OAuth for Search ConsoleSERP, keyword volume, backlinks; your own Search Console
zubair-trabzada/geo-seo-claude10.9kAI visibility and brand mentionsPage fetches and heuristic scoringChatGPT, Perplexity and AI Mode answers with citations
phuryn/pm-skills26.6kCompetitive battlecardsWeb search plus your win/loss notesCompany enrichment, news, pricing pages, ads
swan-gtm/gtm-skills161Account research briefsTool-agnostic company, people and search evidenceFirmographics, decision makers, hiring
superdesigndev/tregoursBuyer signals, UGC videosRuns on treg directlyAll of the above

Stars read from GitHub on 29 Sep 2026. The six third-party repositories are MIT-licensed.

Appendix B

Clay alternatives for GTM engineers who work in Claude Code or Codex

If you already think in prompts, the question is not which table to use but where the data comes from and what it costs per correct result. Measured on the same 292 people:

OptionHow you workExact matchPer correct email
Claude Code or Codex + treg.toPrompts; per-call data layer; your own keys first90.4%$0.0056
ClayVisual table, shared workspace, credits89.7%$0.0395
FreckleTable-based enrichment90.1%$0.0427
DeeplineBatch enrichment86.6%$0.0924

Keep Clay if your team needs a shared visual table and already pays for it. Move the work into your agent when you want the whole playbook in prompts, paid per call. The full bench.

Appendix C

Every recorded run behind this playbook

RunDateWhat came backMetered
Playbook runs: counts, lookalikes, committee, timing, inbound (ch. 2, 4, 6, 9, 10, 13)30 Sep294 calls on one ICP; figures in each chapter$2.70
Study: job changes against five databases (ch. 9)7 Oct148 people; 68% of records did not show the new job$11.10
Study: hiring and news before a raise (ch. 10)6 Oct57 raised against 54 with no round found; hiring did not separate them$13.65
Study: found against deliverable (ch. 8)7 Oct60 found, 32 deliverable$0.65
Build a verified lead list (ch. 3, 5, 8, 11, 14)23 Sep20 deliverable leads from 50 companies$2.33
Work email bench (ch. 7)16 Sep292 people across five aggregators$1.49 on treg
Find creators in a niche14 Sep25 creators from 153 matches$0.20
Screen creators before outreach14 Sep18 of 20 profiles, median engagement 5.3%$0.036
Price keyword demand14 Sep50 keywords, 742,970 monthly searches$0.11
Mine a competitor’s ads14 Sep20 Meta ads and 17 Google creatives$0.015*

The three studies are separate tests, each on its own sample, and are not part of the one-ICP runs; the hiring study’s figure includes building its two samples. The $5.03 in the header is the 30 Sep and 23 Sep runs, which follow one ICP; it includes $1.08 we spent re-running a step after our own parsing bug (ch. 16). * Two of that run’s calls used the team’s own Apify key, which treg never meters.

Appendix D

Glossary and questions

ICPIdeal customer profile, written as filters plus checks (ch. 1).
TAMThe count of companies that match your filters (ch. 2).
CheckA judgement on each row that no filter can express, run before paid steps (ch. 5).
WaterfallTrying providers in order until one answers. treg’s routed endpoints try them cheapest first.
Catch-allA domain that accepts any address, so a verifier cannot confirm a specific inbox (ch. 8).
SignalA dated, linkable event that makes a message timely (ch. 9).
Live readFetching a person’s public profile at the moment you need it, instead of a stored database record (ch. 9).
TierA priority bucket from fit and timing (ch. 10).
Shadow modeRunning a play alongside the manual process to compare before switching (ch. 16).
What is the difference between a GTM engineer and RevOps?

RevOps coordinates revenue processes, systems, planning and reporting across teams. A GTM engineer builds and runs specific workflows, such as enrichment, scoring, signals and routing, and often sits inside RevOps. Titles vary and the work overlaps; in practice RevOps decides how the revenue system should work, and a GTM engineer builds the automated parts of it.

How accurate is contact data after someone changes jobs?

Not very, in the first month. In our 7 Oct 2026 study of 148 people who had announced a new job 1 to 29 days earlier, 68% of contact database records on average did not show the new employer, and when two databases both had a person, both missed it 64% of the time. A live profile read was current for 47 of the 49 people it found, but it found only a third of them. For job-change plays, confirm the new company live before you write.

Do hiring or news signals predict that a startup is about to raise?

Hiring did not, in our test. Across 57 US startups that announced a seed to Series B round in August to October 2026 and 54 similar startups with no round found, the share that opened GTM roles, any job or senior roles in the weeks before was within a few points, and recorded news was more common in the comparison group with no round found (54% against 26%). Act on the funding announcement itself, and test any timing signal on your own wins and losses before you weight it.

Do I need to write code?

No. You describe the job in a prompt, the agent chooses the calls, and it shows the price before it spends. Writing the rules in chapter 3 is the part that needs you.

Does treg.to choose the data provider for me?

The catalog lists every provider of a job with measured success rates and prices, and your agent picks. Some jobs also have a routed endpoint (treg.<capability>) that tries providers cheapest first when you call it; the result names the provider that answered.

What does a miss cost?

It depends on the provider: some bill per successful result, some per call. Each catalog entry says which, routed endpoints do not bill misses on per-success providers, and every receipt shows what was actually charged.

Can I use the provider keys my team already pays for?

Yes. Register the key once and calls to that provider go through it. A team’s own key always wins over treg.to’s, and those calls are never metered.

Which agents does this work with?

The setup line works in Claude Code, Codex and Cursor; anything that can run a shell command can use the treg CLI. The prompts in each chapter are plain language and do not depend on one agent.

Run chapter 1
in your agent today.

One key for 3,900+ tools across 111 providers. $1.00 of free credit once per new verified account, and the counts in chapter 2 are free.

›set up treg — https://treg.to/llms.txt
$npx skills add superdesigndev/treg
Set up tregRead as Markdown