1

Build notes

How it was built

The request flow, the decisions that keep the score honest, and how the work was done. This is a prototype built to illustrate an approach, not a product in use.

The flow

One call, start to scorecard

  1. Browser

    The call screen

    The transcript is held in the page. Nothing is stored.

  2. Server

    /api/lead

    AI plays the lead. The lead's hidden motivation never leaves the server.

  3. Server

    /api/score

    AI grades each rubric item and must quote the trainee.

  4. Code

    Checks and arithmetic

    Every quote is checked against the transcript. The total is computed from fixed weights.

  5. Browser

    Scorecard

    Score, quoted evidence, and one coaching point.

Optional side path

Coach mode: help during the call

Switching on “Coach me” adds one more request to each turn. When the lead replies, the browser also calls /api/coach, a third AI role that runs beside the lead and the grader and shares nothing with either.

  • It sees what the agent sees.

    The coach is given the lead card and the transcript so far. It is never given the lead's hidden reasons, so it cannot hand over the answer.

  • It suggests. The agent decides.

    After each thing the lead says, it returns a one-line read and two lines the agent could say next. A click puts one in the text box to edit. Nothing is sent for them.

  • A coached call is labelled.

    It is off by default, and a scorecard from a call where it was on says so. A coached score is never mistaken for an unaided one.

Practising with suggestions showing measures something different from practising without them, which makes coach on against coach off a natural second arm for the pilot.

Decisions that matter

Why the score can be checked

  • 01

    The model grades. Code does the arithmetic.

    The model scores each rubric item from 0 to 4. The total is computed in code from fixed weights, so it cannot drift or be talked up.

  • 02

    Every quote is checked.

    Each evidence quote is verified against the transcript. A quote that cannot be found is dropped and the score for that item is capped.

  • 03

    The answer stays on the server.

    The lead's hidden reasons are never sent to the browser during a call, so the answer cannot be read from the network tab.

  • 04

    The transcript is untrusted input.

    The grader treats what the trainee typed as material to grade, not as instructions. Asking for a perfect score in the call does not produce one.

  • 05

    Nothing is stored.

    The transcript lives in the page for the length of the call. There is no database and no account.

  • 06

    Inputs are capped.

    Each turn is limited to 600 characters and each call to 40 turns. Requests are rate limited.

How I worked

Contract first, then agents

  1. 01

    Set the contract by hand

    The types, the rubric and the personas were written first. They fix what a call, a score and a lead are.

  2. 02

    Ran coding agents in parallel

    Separate agents built the API, the call screen and these pages against that contract, at the same time.

  3. 03

    Reviewed and tested the output

    Everything the agents produced was read, run and corrected before it went in.

Stack

  • Next.js
  • TypeScript
  • Tailwind
  • Claude via OpenRouter
  • Deployed on Vercel
Next, step 4The rolloutWhat it would take to get this to a whole agent network.