Template · 8 min read

The GTM experiment brief

A lightweight template for running revenue experiments that produce a decision—not just a busy spreadsheet.

Key idea

An experiment is complete when the team can confidently decide what to scale, change, or stop.

Use this before building

Most GTM experiments fail in the planning stage: the hypothesis is vague, the audience changes mid-test, or success is defined after the results arrive. Use one brief per meaningful test.

PromptWrite this down
DecisionWhat will we decide when this test ends?
AudienceWhich accounts, people, or routeable records are eligible?
HypothesisIf we change [input], then [outcome] will improve because [reason].
ControlWhat stays the same so the result is interpretable?
OwnerWho can make the decision and who runs the work?
End dateWhen do we stop collecting data and review the result?

The experiment canvas

  1. Name the current problem

    Use a measurable observation: “qualified demo follow-up takes more than two hours” rather than “inbound is slow.”

  2. Choose one change

    Test a revised routing rule, enrichment step, message angle, or handoff—not several at once.

  3. Define a primary measure

    Pick the metric closest to the decision: meetings held, accepted opportunities, time to first response, or reviewer acceptance.

  4. Add guardrails

    Set thresholds for downside: low-quality meetings, opt-outs, complaints, incorrect routing, or excess manual work.

  5. Write the decision rule

    State the outcome that means scale, iterate, or stop before the team sees the results.

A worked example

Part of the briefExample
HypothesisRouting high-fit demo requests by territory and product interest will reduce first-response time without lowering meeting quality.
ControlKeep form, SLA, territory definitions, and sales capacity constant.
Primary measureMedian minutes to meaningful first response.
GuardrailsMeeting-held rate and lead reassignment rate must not worsen.
DecisionScale if response time falls by 30% or more and both guardrails remain within their baseline range.

The review note

  • What changed, and what did not?
  • What did the primary metric do?
  • Did any guardrail trigger?
  • What did users or customers experience that the numbers do not show?
  • What is the next decision: scale, iterate, stop, or gather more evidence?