Experiment brief
The experiment
This is the page I would bring to a steering review. It says what the prototype is testing, how a pilot would measure it, and what result would advance it, redirect it or stop it.
Status: proposal. No pilot has been run. Nothing on this page is a result.
01
The problem
New agents pay for their own leads and learn on live ones. Every early mistake costs money, and many leave before they get good.
This is the premise being tested. It is not a measured fact about any particular company, and the first job of a pilot is to check it against that company's own numbers.
02
Hypothesis
New agents who complete a set number of scored practice calls before their first live leads will book appointments at a higher rate in their first 30 days than agents who do not.
The number of practice calls is agreed before the pilot starts and does not change during it.
03
Pilot design
- Cohorts
- Two matched groups of new agents from the same onboarding class.
- Difference
- One group practises before taking live leads. The other does not.
- Held constant
- Same lead type for both groups.
- Length
- 30 days from each agent's first live lead.
04
Metrics
Primary
Appointments set per 100 leads worked in the first 30 days.
Secondary
- Days to first placed policy
- 60-day retention
- Lead spend per appointment
05
Decision rule
Agreed before the pilot starts, so the result decides what happens and not the mood in the room. Lift is measured on the primary metric, practising cohort against comparison cohort.
Advance
15 percent relative lift or better
The practising cohort books at least 15 percent more appointments per 100 leads than the comparison cohort. Hand the prototype to the product team with the pilot data.
Redirect
Between 5 and 15 percent
There is a signal but not a case. Change one thing, such as the number of practice calls or the rubric, and run one more cohort before deciding.
Stop
Below 5 percent
Practice calls did not move the metric. Write up what was learned, archive the prototype, and spend the time on the next idea.
These thresholds are proposals, not findings. The real ones are set with the business before the pilot starts.
06
What this prototype does not prove yet
- The grader has not been calibrated against a human sales coach.
- Scores have not been tied to real outcomes. A high score here has not been shown to predict a booked appointment.
- The leads are invented. Real leads may object in ways these three do not.
What I would do first
Have an experienced coach grade a sample of calls blind, then compare with the AI grader item by item. Nobody should trust a score before that comparison exists.
07
What data it would need
- Outcomes per lead
- From the CRM or dialer: contacted, appointment set, policy placed.
- Onboarding cohort
- Which class each agent started in, so the two groups can be matched.
- Lead cost
- What each agent paid per lead, to work out spend per appointment.
The pilot as designed needs no call audio. Listening to real calls would need recording consent handled state by state before any real audio is used.