Template · 8 min read
The GTM experiment brief
A lightweight template for running revenue experiments that produce a decision—not just a busy spreadsheet.
An experiment is complete when the team can confidently decide what to scale, change, or stop.
Use this before building
Most GTM experiments fail in the planning stage: the hypothesis is vague, the audience changes mid-test, or success is defined after the results arrive. Use one brief per meaningful test.
| Prompt | Write this down |
|---|---|
| Decision | What will we decide when this test ends? |
| Audience | Which accounts, people, or routeable records are eligible? |
| Hypothesis | If we change [input], then [outcome] will improve because [reason]. |
| Control | What stays the same so the result is interpretable? |
| Owner | Who can make the decision and who runs the work? |
| End date | When do we stop collecting data and review the result? |
The experiment canvas
- Name the current problem
Use a measurable observation: “qualified demo follow-up takes more than two hours” rather than “inbound is slow.”
- Choose one change
Test a revised routing rule, enrichment step, message angle, or handoff—not several at once.
- Define a primary measure
Pick the metric closest to the decision: meetings held, accepted opportunities, time to first response, or reviewer acceptance.
- Add guardrails
Set thresholds for downside: low-quality meetings, opt-outs, complaints, incorrect routing, or excess manual work.
- Write the decision rule
State the outcome that means scale, iterate, or stop before the team sees the results.
A worked example
| Part of the brief | Example |
|---|---|
| Hypothesis | Routing high-fit demo requests by territory and product interest will reduce first-response time without lowering meeting quality. |
| Control | Keep form, SLA, territory definitions, and sales capacity constant. |
| Primary measure | Median minutes to meaningful first response. |
| Guardrails | Meeting-held rate and lead reassignment rate must not worsen. |
| Decision | Scale if response time falls by 30% or more and both guardrails remain within their baseline range. |
The review note
- What changed, and what did not?
- What did the primary metric do?
- Did any guardrail trigger?
- What did users or customers experience that the numbers do not show?
- What is the next decision: scale, iterate, stop, or gather more evidence?