Build notes
How it was built
The request flow, the decisions that keep the score honest, and how the work was done. This is a prototype built to illustrate an approach, not a product in use.
The flow
One call, start to scorecard
- Browser
The call screen
The transcript is held in the page. Nothing is stored.
- Server
/api/lead
AI plays the lead. The lead's hidden motivation never leaves the server.
- Server
/api/score
AI grades each rubric item and must quote the trainee.
- Code
Checks and arithmetic
Every quote is checked against the transcript. The total is computed from fixed weights.
- Browser
Scorecard
Score, quoted evidence, and one coaching point.
Optional side path
Coach mode: help during the call
Switching on “Coach me” adds one more request to each turn. When the lead replies, the browser also calls /api/coach, a third AI role that runs beside the lead and the grader and shares nothing with either.
It sees what the agent sees.
The coach is given the lead card and the transcript so far. It is never given the lead's hidden reasons, so it cannot hand over the answer.
It suggests. The agent decides.
After each thing the lead says, it returns a one-line read and two lines the agent could say next. A click puts one in the text box to edit. Nothing is sent for them.
A coached call is labelled.
It is off by default, and a scorecard from a call where it was on says so. A coached score is never mistaken for an unaided one.
Practising with suggestions showing measures something different from practising without them, which makes coach on against coach off a natural second arm for the pilot.
Decisions that matter
Why the score can be checked
01
The model grades. Code does the arithmetic.
The model scores each rubric item from 0 to 4. The total is computed in code from fixed weights, so it cannot drift or be talked up.
02
Every quote is checked.
Each evidence quote is verified against the transcript. A quote that cannot be found is dropped and the score for that item is capped.
03
The answer stays on the server.
The lead's hidden reasons are never sent to the browser during a call, so the answer cannot be read from the network tab.
04
The transcript is untrusted input.
The grader treats what the trainee typed as material to grade, not as instructions. Asking for a perfect score in the call does not produce one.
05
Nothing is stored.
The transcript lives in the page for the length of the call. There is no database and no account.
06
Inputs are capped.
Each turn is limited to 600 characters and each call to 40 turns. Requests are rate limited.
How I worked
Contract first, then agents
01
Set the contract by hand
The types, the rubric and the personas were written first. They fix what a call, a score and a lead are.
02
Ran coding agents in parallel
Separate agents built the API, the call screen and these pages against that contract, at the same time.
03
Reviewed and tested the output
Everything the agents produced was read, run and corrected before it went in.
Stack
- Next.js
- TypeScript
- Tailwind
- Claude via OpenRouter
- Deployed on Vercel