Skip to content

Add the whiteboard interview mode - #72

Open
ColtenOuO wants to merge 5 commits into
sysprog21:mainfrom
ColtenOuO:feat/whiteboard-mode
Open

ColtenOuO wants to merge 5 commits into
sysprog21:mainfrom
ColtenOuO:feat/whiteboard-mode

Conversation

@ColtenOuO

@ColtenOuO ColtenOuO commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A whiteboard interview is the coding round held without a compiler. The lobby offers it next to the duration and the loop. The editor, the starter code and the test runner are replaced by a drawing surface, and the candidate shows correctness by drawing an approach, tracing an example across it, and naming the cases that would break it. It uses the same problem bank, the same six REACTO phase ids, the same rubric and the same report schema. Only the surface and what each step asks for change.

Closes #65, with two deliberate departures from the issue's proposal:

  • No handwritten pseudo-code step. The flow is the six REACTO steps, so the evidence vocabulary, the progress checklist and the rubric all stay as they are.
  • No new grading dimensions. The report is graded against the existing rubric. The reviewer is told what each step means at a board.

What the candidate sees

  • Board. A fixed 1600x1000 canvas that scrolls inside its panel, so the interviewer always sees the whole board whatever the candidate's screen size. Mouse input only in this version.
  • Pens. Four colours, shown on a white strip so each swatch looks like the ink it draws.
  • Eraser. Three sizes: 12, 24 and 56 board px. The cursor becomes a circle the size of the erasure.
  • Undo, redo and Clear board. Clear is a single undoable step.
  • Opening. The greeting does not ask for a language, since there are no tabs and nothing compiles. It goes straight to restating the problem.
  • Checklist. The step list beside the timer uses the whiteboard names below.
Phase id At the board What is asked
repeat Repeat restate inputs, outputs, constraints and ambiguities
example Example draw one ordinary case and one boundary case
algorithm Approach draw the approach and its cost before tracing it
coding Trace walk one of their examples through the drawing
test Edge cases name what would break it, and what it does on each
optimizations Complexity confirm time and space, name one optimization

How a board reaches the interviewer

The browser keeps the drawing as vector strokes (web/whiteboard.js). That is what makes undo and replay possible. Rendering a board to pixels happens only when something needs an image.

flowchart LR
    Stroke["Stroke ends"] --> Settle["1 s with no new stroke"]
    Settle --> JPEG["Canvas exported as JPEG<br/>quality 0.72"]
    JPEG -->|"LiveKit byte stream<br/>topic board_image"| Drain["Agent drain task<br/>bounded at 512 KiB"]
    Drain -->|"queue of 2"| Loop["Room loop: pump_board"]
    Loop --> Latest["Latest board in memory"]
    Loop -->|"at most one per second<br/>realtimeInput.video"| Live["Gemini Live"]
    Tool["read_board tool call"] --> Resend["Resend latest board"]
    Latest --> Resend
    Resend --> Live
Loading
  • Why a byte stream. A board is tens of kilobytes, more than one data packet carries. The stream header carries a strokes count and, for a phase checkpoint, a checkpoint id.
  • What the agent refuses. A stream is refused unless it is on the board topic, from the candidate, in a whiteboard interview, and declares no more than 512 KiB. The same bound is enforced again while reading, for a sender that declared nothing.
  • Keeping the room responsive. The stream is read on its own task, so audio and room events never wait behind a board. The queue holds two boards and the newest is the one that matters, because each board is the whole drawing rather than a change to it.
  • The one-second floor. The browser already debounces, but the agent still sends at most one board per second. A board that arrives too soon is kept as the latest, so read_board answers with it.
  • read_board. Returns JSON with the stroke count, the board's age and the timer, then resends the latest image, because a tool response cannot carry a picture.
  • Pause and silence. A paused interview drops boards the way it drops audio. Drawing counts as activity, so the silence nudge does not interrupt a candidate who is drawing without talking.

What the interviewer is told

There is one live prompt. The sentences that name the surface sit in a Surface struct with a coding and a whiteboard value, so a rule cannot end up in one prompt and not the other. Rebased onto #203, #192 and #202: the coding prompt is byte-identical to main's, and only the whiteboard entries are new in tests/golden/prompts.json. The thinking-time rule from #192 says "board changes" rather than "editor changes" at a board.

Phase checkpoints and grading

The report reviewer sees the board as it stood when each step closed, not just the final one. A candidate who clears the board between steps does not lose that evidence.

sequenceDiagram
    participant B as Browser
    participant A as Interview agent
    participant G as Gemini report model

    A->>B: framework_state (phases with evidence)
    B->>B: New phase ids, in REACTO order
    B->>A: JPEG + checkpoint id (after pointerup if mid-stroke)
    A->>A: Validate the id, keep one image per phase
    B->>A: Final JPEG when End is pressed
    B->>A: end_interview, once the uploads finish
    A->>G: generateContent(labelled inline JPEGs + prompt)
    Note over A,G: Retries and schema repairs resend every image
    G-->>A: Structured JSON report
    A-->>B: Report, stamped with interviewMode
Loading
  • Capture. Each framework_state update is reduced to known phase ids that have not been captured yet. A phase that completes while the pointer is down waits for pointerup, so the image and the replay describe the same stroke.
  • Labels. The agent validates the id again and assigns the label itself (Repeat, Example, Approach, Trace, Edge cases, Complexity). A string from the browser therefore never sits next to an image as an instruction to the reviewer.
  • Images sent to the reviewer. At most one image per phase, newest wins. The final board is added after the checkpoints when it differs from the last one.
  • Ordering. Uploads run one at a time, and end_interview is sent only after they finish. That way the report is never frozen on an older board.
  • Report prompt. It tells the reviewer what each phase means at a board, and says so when no board arrived.

Replay and recording

Replay stores no JPEGs. A board as a JPEG is over the per-event size limit and would use up the per-interview replay budget within a few frames. Replay instead stores the operations that drew the board: stroke, undo, redo and clear, in batches under 24 KiB, each optionally tagged with a checkpoint id. The replay page and the recording template rebuild any moment of the board with the same module the candidate drew on (createBoard + applyOp). Every operation is validated on the way back in: colour, width, point count and coordinates.

The report card

A whiteboard report opens on the board, below the committee summary:

  • "Your board, step by step." Each of the six steps shows the board as that step closed, the step's score, and what the interviewer recorded about it. A step with nothing recorded says so, rather than disappearing.
  • Final board. Shown after the six steps.
  • Names. Steps, Practice next and Framework evidence use the whiteboard names. The coding score is labelled Board work.
  • Schema. Unchanged. The card maps the REACTO ids it already carries to display names.
  • Images are not saved. The step images are the page's own copies of the checkpoints it sent. Six data URLs would be most of what a saved account report may weigh, so a report reopened from history says the pictures are in the recording.
  • Export. The .md file lists the steps and their evidence without base64.

Testing

  • Run locally on the rebased branch. cargo test (1204 passed), clippy with -D warnings, the browser suite (839 passed, 1 skipped for no Java), ESLint, the drift checks and the hook suite, and scripts/indent.sh --check with commentflow. All passed.
  • Not run locally. actionlint could not run from a git worktree. ruff, shellcheck and cargo-audit are not installed here.
  • Not exercised. No live interview against Gemini was run after the rebase. The board path is covered by the unit tests in tests/unit/livekit/board.rs and tests/agent/whiteboard.rs, and the report card by tests/browser/render.test.js.

Whiteboard

Whiteboard interview

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Need some time for testing and self-review; will mark this PR as ready for review once it's set.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Perhaps we will need to record the state of the whiteboard at each stage so that we can provide more information to the user at the end of the interview.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Additionally, pseudocode should be graded separately rather than being combined with optimization like it is currently.

@alanhc

alanhc commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

On the concern from #65 about whether Gemini can read a drawn board reliably: one option worth considering is Excalidraw as the board layer.

The drawing tools are not the point; the data model is. An Excalidraw scene is a list of typed elements, and boxes, arrows (bound to the boxes they connect) and text keep their actual strings. So next to the JPEG, the agent can get a plain-text summary of the board, e.g. box "left" -> box "mid", and typed pseudo-code arrives as exact text rather than something the model has to read off pixels. The report prompt can quote it too.

I checked whether it fits the vendoring rules, following the three-vrm.js precedent (one esbuild bundle, committed, with a reproduce recipe):

  • @excalidraw/excalidraw@0.18.1 + React 19 bundle into a single ES module: 4.9 MB (1.6 MB gzip), plus 145 KB of CSS. That is with @excalidraw/mermaid-to-excalidraw aliased to a stub; without it, 8.5 MB.
  • Fonts load from window.EXCALIDRAW_ASSET_PATH. The Latin ones are ~0.5 MB. The CJK font (Xiaolai) is 13 MB; I left it out, and Excalidraw then falls back to esm.sh for it. The current CSP refuses those loads (230 console violations in my run, drawing and export unaffected), but it needs either vendoring or silencing.
  • Served locally with every non-local request aborted, under the policy from src/web/policy.rs (script-src 'self', style-src 'self', nothing inline): it mounts in ~200 ms, freehand, shapes and text all work, and exportToBlob gives a ~20 KB JPEG. The only inline bit was setting EXCALIDRAW_ASSET_PATH, which moves to a same-origin script.
  • A 41-point freehand stroke is ~1.5 KB of JSON, so replay would still journal per-element changes from onChange rather than whole scenes.

The costs:

  • Far bigger than the 280-line whiteboard.js, and it brings React into the page.
  • The structured benefit exists only when the candidate uses shapes and the text tool. Freehand pseudo-code is still just points. Whether typed text belongs in a whiteboard interview is a product call: less realistic, much easier to grade.
  • Text input opens the integrity side: paste and library/file import would have to be disabled.

Only the board layer changes. The byte stream, read_board, the report attachment and the replay plumbing stay as they are. Happy to put together the vendored bundle and the scene-to-text summary, either on top of this branch or as a follow-up once it lands, whichever you prefer.

@ColtenOuO

Copy link
Copy Markdown
Collaborator Author

Hi, @alanhc

Thanks for the discussion!

Excalidraw was also brought up in the discussions under the Facebook post back then. After looking into it, I think it seems like a solid option to consider, but I also have a few thoughts.

If our goal is to simulate an actual interview environment as closely as possible, would Excalidraw make things too convenient for the user? After all, in a real interview, you only have a marker and a physical whiteboard. That raw experience is precisely what we want to deliver in this mode—allowing users to practice explaining their problem-solving approach and thought process purely through drawing on a blank board.

Currently, a tester, @Eason0729 , has tested this feature (the non-Excalidraw version) and provided a lot of feedback. Here is a brief summary of the key concerns raised from the testing:

  1. Should we support tablet touch input? Drawing with a mouse can significantly degrade the user's drawing experience and heavily impact their performance. Therefore, I would strongly prefer having this supported.

  2. Unreliable AI recognition will severely impact the quality of the questions asked.

For now, I'll focus on addressing the points raised in the feedback first. As for Excalidraw, we can discuss it further as a potential future improvement.

Overall, I actually think this is quite a promising direction. However, my primary concern is that users should have an experience that feels as close to reality as possible; we probably shouldn't compromise on that just to make recognition easier for the model.

WDYH

@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch 4 times, most recently from 897c017 to 2228cf7 Compare October 1, 2026 22:52
A whiteboard interview takes the same problem bank and the same six
steps and swaps the editor and the test runner for a board. Nothing
runs, so correctness is what the candidate can defend by tracing an
example across their own drawing.

The board travels as a LiveKit byte stream rather than on a data topic,
because one board is tens of kilobytes and a data packet carries
fifteen, and it reaches Gemini the way a camera frame does, as a
realtime image. A tool response is JSON and cannot carry a picture, so
read_board asks for the board to be sent again and answers with what an
image cannot say: how much is on it and how long ago it was drawn.

Bundle 27, live prompt 19, follows the surface. A whiteboard
session is told it has no editor and no test runner, is offered
read_board in place of read_editor, and records board_snapshot where
the other records an editor snapshot or a test event; neither may
record the other's source. The phases about written work are gated on
strokes instead of on characters, and Test on the cases named against
the drawing instead of on a run, or an interview with no editor would
be refused the second half of its own flow. The report still reads the
transcript alone and the replay still carries no board.
The whiteboard was live but nowhere else: the report was written from a
transcript and an editor nobody opened, the recording kept no drawing,
and the checklist beside the timer called step four Coding while the
interviewer was asking the candidate to trace.

The board now rides the report request as an inline image, ahead of the
brief and on every repair, and the brief says what it is looking at:
nothing ran, so correctness is the trace the candidate walked, and the
phases keep their names with Coding meaning that trace, Test the cases
named against the drawing, and Optimizations the complexity they
confirmed. The system instruction is left as it is, so every report
call still shares its prefix, and the brief tells a whiteboard reviewer
how to read the rules that speak of code. That is report prompt 16,
inside the same bundle 27, and the reviewer is told when no board
arrived rather than sent looking for an attachment that is not there.

The recording keeps the drawing as the operations that made it. One
board as a JPEG is past the per-event ceiling on its own and would
spend the whole per-interview budget on a handful of frames; the same
board as strokes is a few kilobytes, so the replay page and the
recording template redraw any moment of the interview with the module
the candidate drew on, rather than the few moments a photograph could
afford.
A final board cannot represent work cleared between phases. Keep one
JPEG checkpoint per phase for grading and a vector marker for replay,
so review survives a clear without putting images into the replay
budget. Clear itself becomes one undoable step: with the earlier phases
kept, an accidental clear is the costliest mistake left at the board.
One fixed width was either too wide to take out a character or too
narrow to clear a region. Three offered sizes cover both, and a fixed
list rather than a slider keeps every width on the replay one a button
can make. The cursor becomes a circle the size of the erasure.
A whiteboard report read as an editor one: a Coding score, REACTO step
names the candidate was never shown, and one board folded at the foot
of the card. The page already captured the board as each step closed,
for the grader, and kept none of it.

The report now opens on the board, step by step: the image as that step
closed, its score and what the interviewer wrote down, then the final
board. Steps, plans and evidence take the whiteboard names, and the
score is Board work. Only the display changes; the report keeps the
REACTO ids it is graded against. The images are never saved, so a
report reopened from history says where they are instead.
@ColtenOuO
ColtenOuO force-pushed the feat/whiteboard-mode branch from 2228cf7 to a124a50 Compare October 2, 2026 04:55
@ColtenOuO
ColtenOuO marked this pull request as ready for review October 2, 2026 05:16

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

48 issues found across 57 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="docs/interview-contract-versions.md">

<violation number="1" location="docs/interview-contract-versions.md:15">
P2: `session_timing` is also valid for skipped STAR evidence, so this overstates the whiteboard source restriction. Qualify the claim as applying to observed evidence or mention the skipped-evidence exception.</violation>
</file>

<file name="src/recording/replay.rs">

<violation number="1" location="src/recording/replay.rs:60">
P1: `Board` is accepted by replay parsing, but the `replay_events` schema still rejects `kind = "board"`. Every whiteboard batch therefore fails at insertion and returns 500 instead of being replayed; add a schema migration that permits `board` before enabling this kind.</violation>
</file>

<file name="tests/recording.rs">

<violation number="1" location="tests/recording.rs:3600">
P3: Board is now in the schema list but never asserted in either classification loop, so its is_snapshot() result (false — it Accumulates in replay.rs:109) has no test pinning it. Add ReplayKind::Board to the accumulates list so the new kind's late-join semantics are locked down like the other six.</violation>
</file>

<file name="tests/agent/report.rs">

<violation number="1" location="tests/agent/report.rs:285">
P3: This contract test now covers only the Coding variant of the mode-parameterized `greeting`. The Whiteboard branch this PR adds carries the same obligations (hint offer, no published material, introduce THE EXERCISE), but it is not exercised here, and its text would actually fail the existing assertions: it capitalizes "Do not volunteer a constraint, edge case, or hint" sentence-initially while the assertions match the Coding greeting's lowercase mid-sentence form. Iterate over both `InterviewMode` variants in this test so the new mode's greeting is held to the same contract (using a case-insensitive match or aligning the variant text).</violation>
</file>

<file name="web/replay.js">

<violation number="1" location="web/replay.js:328">
P2: Empty whiteboard sessions are classified as coding because this flag only looks for board moments. Show the empty whiteboard for whiteboard recordings too, using a persisted recording-mode signal rather than requiring a drawing operation.</violation>

<violation number="2" location="web/replay.js:488">
P2: The report can show an earlier selected moment as “Your final board”: `loadReport()` waits for `/api/reports` while timeline buttons remain clickable, and `showMoment()` redraws this canvas before this line reads it. Capture the final board image before the await or restore the final moment before rendering the report.</violation>
</file>

<file name="tests/browser/account.test.js">

<violation number="1" location="tests/browser/account.test.js:141">
P2: The replaced guards removed the only checks on the lobby→room mode link without adding positive ones. `app.js` must do `destination.searchParams.set("mode", mode)` and `interview.js` must read `params.get("mode")` for the selected surface to survive into the token payload; if either line is dropped (mode hardcoded to "coding", or the URL value never set), a whiteboard interview silently runs as a coding interview while this test stays green because it only asserts the data-mode vocabulary and the `interviewMode: mode` string in the payload. Add `assert.match(app, /destination\.searchParams\.set\("mode", mode\)/)` and `assert.match(interview, /params\.get\("mode"\)/)`.</violation>

<violation number="2" location="tests/browser/account.test.js:191">
P3: The test name now contradicts what it asserts: the lobby offers two surfaces ("coding", "whiteboard") and `interviewMode: mode` is carried to the room, yet the title still reads "offers one interview and carries no mode to the room". The body comment was updated for the new contract; the title was not. Rename it (e.g. "the lobby offers coding or whiteboard and carries the mode into the room").</violation>
</file>

<file name="src/gemini.rs">

<violation number="1" location="src/gemini.rs:1458">
P2: Whiteboard interviews still expose a `log_hint` description that tells Jim to give the clue with the candidate's editor, which does not exist in this mode. Make the hint description mode-specific and direct him to fit the clue to the board.</violation>

<violation number="2" location="src/gemini.rs:1473">
P2: This declaration is sent in whiteboard mode, but its description still tells Gemini that evidence can come from an editor or test run. Make the description mode-specific so the interviewer is not guided toward nonexistent evidence sources.</violation>
</file>

<file name="src/livekit/report.rs">

<violation number="1" location="src/livekit/report.rs:71">
P2: Whiteboard images now reach Gemini but never enter `record_model_input`, so the reported model-input byte total omits every attached board. Include board payload bytes in the report metrics so the total remains accurate.</violation>
</file>

<file name="web/render.js">

<violation number="1" location="web/render.js:398">
P2: The board report hides explicitly recorded `skipped` evidence and presents that phase as missing. Keep the skipped outcome visible, or render an explicit skipped status, instead of treating a recorded outcome as absent.

(Based on your team's feedback about skipped framework evidence.) [5de29b75-9991-48cb-96b6-10a6ce846cb4].</violation>

<violation number="2" location="web/render.js:431">
P2: Saved whiteboard exports can contain no images, and recording is optional, so this promises access to a board that may not exist anywhere. Say the images are not included, and qualify replay availability on having a recording; update the matching history-card fallback too.</violation>
</file>

<file name="web/replay.html">

<violation number="1" location="web/replay.html:108">
P2: `<canvas id="replay-board">` has no accessible name: the `<h3>` heading above it does not label a following canvas in the accessibility tree, so a screen reader announces an unnamed canvas and a user has no indication that a board drawing is shown. The constraint cited in the comment does not actually block an `aria-label` — `history.test.js` already allowlists `"Whiteboard"` (it is used for the heading), and this page already carries attributes such as `role="status"`. Give the canvas `aria-label="Whiteboard"`, matching the live board in `web/interview.html`.</violation>
</file>

<file name="scripts/browser-check.cjs">

<violation number="1" location="scripts/browser-check.cjs:448">
P2: `checkWhiteboardInterview` never clicks Undo; it only observes that the model enabled the button, so a broken DOM handler still passes. Click Undo and assert the stroke returns, then exercise Redo as well.</violation>
</file>

<file name="tests/unit/livekit/board.rs">

<violation number="1" location="tests/unit/livekit/board.rs:335">
P3: The "too soon to go out" assertion assumes less than one wall-clock second elapses between the first `pump_board` and the second, because `too_soon` compares `BOARD_SEND_INTERVAL` against `Instant::now()` inside `pump_board`. The gap is normally a few milliseconds, but a scheduling stall on a loaded CI host past one second makes the second board get sent and this assertion fails spuriously. Make the timing deterministic (e.g., a pauseable clock or an injectable interval) or assert the throttle through `too_soon` directly rather than through a wall-clock race.</violation>
</file>

<file name="src/agent/prompts.rs">

<violation number="1" location="src/agent/prompts.rs:192">
P2: The whiteboard mode selects this surface, but the shared interview flow still tells the interviewer to “let them code” after a solid answer despite there being no editor. Replace that shared transition with the whiteboard equivalent: continue drawing and tracing.</violation>

<violation number="2" location="src/agent/prompts.rs:256">
P2: This note says every board change reaches Gemini, but rate-limited snapshots can leave Gemini viewing an older image after the candidate stops drawing. State that snapshots may lag and require `read_board` for confirmation, or schedule the deferred latest send.</violation>

<violation number="3" location="src/agent/prompts.rs:1984">
P2: Whiteboard mode reaches the final report, but idle reviews still receive an editor-only prompt with an empty editor. Adapt or skip interim reviews for whiteboards so their rolling assessment does not omit board work or report a misleading surface.</violation>

<violation number="4" location="src/agent/prompts.rs:2054">
P1: Whiteboard reports still receive the shared rubric that treats the absent editor as a failed solution and requires working code for HIRE. Make the report system instruction mode-aware; the brief's board reinterpretation cannot override these higher-priority instructions.</violation>
</file>

<file name="src/agent.rs">

<violation number="1" location="src/agent.rs:1548">
P2: `board_age_seconds` measures snapshot receipt, not when Gemini actually received the image. During the throttle, `read_board` can report a fresh board while the model still sees the previous frame; update the timestamp only after a successful send, or track receipt and send ages separately.</violation>

<violation number="2" location="src/agent.rs:1561">
P2: `board_strokes` comes from the candidate-controlled stream attribute and is not checked against the image or drawing data. A client can claim three strokes with an empty board, passing this guard and allowing Coding, Test, or Optimizations evidence; validate the count before using it as a gate.</violation>
</file>

<file name="src/livekit.rs">

<violation number="1" location="src/livekit.rs:2106">
P2: Snapshots received while paused are dropped, and resume does not send the current board again. Since the board remains editable, Jim misses work drawn during the pause until another edit produces a snapshot; retain the latest snapshot and send it on resume.</violation>

<violation number="2" location="src/livekit.rs:2918">
P1: The report snapshots `Board` without first consuming completed board streams already queued from the reader tasks. An end packet can therefore freeze a stale board (or no board), causing whiteboard reports to miss the candidate’s final work; drain or otherwise await the board queue before freezing report attachments.</violation>
</file>

<file name="tests/golden/prompts.json">

<violation number="1" location="tests/golden/prompts.json:7">
P2: This no-board prompt says no board was available, then asserts all six phases ran and that the candidate traced, tested, and confirmed complexity on it. That can make report generation invent board evidence; make this instruction conditional on attached evidence and otherwise direct scoring from the transcript and rolling assessment only.</violation>

<violation number="2" location="tests/golden/prompts.json:8">
P2: boardSilenceDrawn freezes code-mode test claims into a whiteboard prompt: "tests: browser-reported claims (unverified): 1 of 3 passing, 1 edit-and-run cycles" contradicts board mode, where nothing runs and there is no edit-and-run. The evidence string passed to `board_silence_nudge` in `prompt_samples` (tests/agent.rs:308-309) reuses the coding sample's `working` string. Give the board sample a board-appropriate evidence string (board strokes/phases, tests: not run) and regenerate the golden, so goldens stop asserting editor/test semantics for the whiteboard mode.</violation>
</file>

<file name="tests/browser/recording-template.test.js">

<violation number="1" location="tests/browser/recording-template.test.js:288">
P2: The file now claims "seven kinds, seven call sites" and adds `board` to the rendered kinds, but unlike every other kind (envelope, transcript, editor, tests, stage, avatar, lifecycle) `board` has no `replay-producer` test. Nothing pins the producer payload or drain: `recordReplay("board", { ops, checkpoint })` in `web/interview.js` `recordBoardOps` would drift (e.g. renaming `ops`, dropping the empty-ops checkpoint emit) without this tripwire failing, and a payload-shape change would break the template's `payload.ops` in production only. Add a producer test in the file's existing style, asserting `recordBoardOps` sends `recordReplay("board", ...)`, batches via `board.model.takeOps()`, and attaches the checkpoint to the last batch.</violation>
</file>

<file name="scripts/gen-wire-fixtures.mjs">

<violation number="1" location="scripts/gen-wire-fixtures.mjs:380">
P2: The new fixture carries a `checkpoint: "algorithm"` attribute, but `reads_the_stroke_count_the_browser_sent` in tests/unit/livekit/board.rs only asserts the topic and stroke counts. If the producer renamed the `checkpoint` attribute (e.g. to `phase`), `gen-wire-fixtures.mjs` would regenerate the fixture, `--check` and `cargo test` would pass, and the agent would silently stop labeling checkpoint images in whiteboard report generation. The comment says the fixture pins "the attributes the agent reads off it"; assert the per-case `checkpoint` value in the Rust test so that key is actually pinned.</violation>
</file>

<file name="web/recording/index.html">

<violation number="1" location="web/recording/index.html:35">
P2: The `hidden` attribute on this canvas has no effect: `recording.css` sets `.board { display: block; }`, an author rule that overrides the user-agent `[hidden] { display: none }`. Recording switches the panels only via the `hidden` property (`recording.js:201-203`), so in editor interviews the recording shows an empty whiteboard panel for the entire video — the layout the comment above this element says it prevents. `web/styles.css:23` works around this exact case with `[hidden] { display: none !important; }`, but the recording page loads only `recording.css`. Add the same reset to `web/recording/recording.css`.</violation>
</file>

<file name="tests/browser/whiteboard.test.js">

<violation number="1" location="tests/browser/whiteboard.test.js:70">
P2: This only proves the returned array is detached; its stroke objects and `points` are still shared, so mutating a returned point silently changes the board without journaling. Also mutate a returned point and assert the board remains unchanged to cover the stated copy contract.</violation>
</file>

<file name="web/index.html">

<violation number="1" location="web/index.html:952">
P2: The default Editor selection is only a CSS class until a mode button is clicked, so screen readers cannot tell which mode Start will use. Set initial `aria-pressed` values on both mode buttons; the existing click handler keeps them updated.</violation>

<violation number="2" location="web/index.html:962">
P3: The new `mode-note` paragraph mirrors `duration-note` directly above it, but omits its `aria-live="polite"` annotation. When the candidate toggles a mode, the note text appears or disappears without any announcement, while every sibling status element on the page (`duration-note`, grounding statuses, `progress-summary`) is announced. Add `aria-live="polite"` so the whiteboard explanation is announced to screen-reader users.</violation>
</file>

<file name="tests/interview_behavior.rs">

<violation number="1" location="tests/interview_behavior.rs:996">
P2: The behavior suite's mode plumbing is adapted to pass InterviewMode, but every new call site hard-codes InterviewMode::Coding, so none of the new whiteboard surfaces (whiteboard greeting, read_board tool, mode-specific instruction text) are exercised by the only tests that assert prompt behavior. Add a whiteboard variant to one of the live tests (e.g. loop interview_mode in `live_interviewer_poses_the_variant_and_serves_hints_in_order` alongside interview_loop) so a whiteboard prompt regression cannot land uncaught.</violation>
</file>

<file name="src/livekit/turn.rs">

<violation number="1" location="src/livekit/turn.rs:1155">
P2: `board_strokes` can describe a throttled snapshot that Gemini never received, but this prompt tells Gemini it has the latest image. Track the last-sent board state or resend the latest image before making that claim.</violation>
</file>

<file name="src/livekit/board.rs">

<violation number="1" location="src/livekit/board.rs:242">
P1: Each accepted board stream gets its own 512 KiB buffer, but the queue does not cap concurrent drain tasks and checkpoint tasks wait rather than being rejected. A candidate can open many streams and exhaust agent memory; cap active readers (for example with a semaphore) before spawning, and release the permit when draining finishes.</violation>

<violation number="2" location="src/livekit/board.rs:300">
P1: Repeated valid checkpoint uploads can accumulate one blocked task and up to 512 KiB per snapshot after the channel fills. Bound or coalesce checkpoint backlog before retaining snapshots, or a candidate can exhaust the agent's memory.</violation>

<violation number="3" location="src/livekit/board.rs:301">
P1: This drops the newest board whenever the queue is full, leaving Gemini with older snapshots; if the candidate stops drawing, nothing publishes the final board afterward. Replace an older queued snapshot before enqueueing the newest one.</violation>
</file>

<file name="web/interview.js">

<violation number="1" location="web/interview.js:2989">
P2: The editor is disabled during a pause, but this board input path and its toolbar remain active, and completed snapshots are still sent to Jim. Disable the board controls and reject pointer edits until the interview resumes.</violation>

<violation number="2" location="web/interview.js:3220">
P1: A board settled during a LiveKit reconnect is discarded and never resent after recovery. Retain the latest board or enqueue a fresh board snapshot from the `Reconnected` handler so the interviewer catches up even when the candidate makes no further edits.</violation>

<violation number="3" location="web/interview.js:3235">
P2: Whiteboard recordings still emit an empty editor snapshot at connection time. Skip editor replay events for whiteboard sessions, otherwise replay shows an irrelevant empty Code moment alongside the board timeline.</violation>
</file>

<file name="web/whiteboard.js">

<violation number="1" location="web/whiteboard.js:68">
P3: This production module exposes `strokeColor`, `strokeWidth`, and `drawStroke` without any runtime consumer; the first two are used directly only by tests, and `drawStroke` is only called internally. Keep these helpers private and exercise them through `createBoard()` and `drawBoard()` so the production API stays focused.

(Based on your team's feedback about keeping test-only helpers out of the production module API.) [ada3c0d3-a44b-4ef2-9575-7c3e2ed7190c]</violation>

<violation number="2" location="web/whiteboard.js:128">
P2: `undo()` loses the clear redo marker when it undoes a prior stroke after restoring a clear. After Clear → Undo → Undo → Redo, the next Redo is unavailable, so the candidate cannot restore the cleared board.</violation>
</file>

<file name="web/recording/recording.js">

<violation number="1" location="web/recording/recording.js:201">
P2: The recorder leaves the blank code surface visible until the first `board` event. An untouched or early whiteboard recording therefore shows a code panel instead of the whiteboard; pass the interview mode to the recording template or emit an initial board marker before recording starts.</violation>
</file>

<file name="src/livekit/session.rs">

<violation number="1" location="src/livekit/session.rs:985">
P2: `read_board` drops the pending candidate context when Gemini context compression is enabled. Add `TOOL_READ_BOARD` to the continuity-response set so the model can continue the current exchange after reading the refreshed board.</violation>

<violation number="2" location="src/livekit/session.rs:1034">
P2: A requested whiteboard hint can use a stale drawing: rate-limited snapshots update `Board::latest` without reaching Gemini, but this branch suppresses the refresh that the editor hint path provides. Request a board resend for whiteboard hints before returning the clue.</violation>
</file>

<file name="web/app.js">

<violation number="1" location="web/app.js:259">
P2: The mode buttons remain active while `recordGitHubLogin()` is awaited, but the destination captures `mode` before that await. Clicking Whiteboard during sign-in updates the UI while navigation still launches coding; ignore mode clicks or disable these buttons while `starting`.</violation>
</file>

<file name="tests/browser/render.test.js">

<violation number="1" location="tests/browser/render.test.js:2019">
P3: The step-name loop covers five of the six whiteboard steps but misses "Complexity" (the renamed `optimizations` step), so a wrong label on that step would pass the suite. Add it to the list.</violation>

<violation number="2" location="tests/browser/render.test.js:2025">
P3: `card.match(/<img /g)` returns `null`, not `[]`, when no step image renders, so `.length` throws a bare TypeError that hides the failing assertion. Guard the match before reading `.length`.</violation>
</file>

Tip: cubic used a learning from your PR history. Let your coding agent read cubic learnings directly with the cubic MCP.

Re-trigger cubic

Comment thread src/recording/replay.rs
/// board as vectors is a few kilobytes, and it arrives as the operations
/// that produced it, so the review page can redraw any moment of the
/// interview rather than the few it could afford to photograph.
Board,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: Board is accepted by replay parsing, but the replay_events schema still rejects kind = "board". Every whiteboard batch therefore fails at insertion and returns 500 instead of being replayed; add a schema migration that permits board before enabling this kind.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/recording/replay.rs, line 60:

<comment>`Board` is accepted by replay parsing, but the `replay_events` schema still rejects `kind = "board"`. Every whiteboard batch therefore fails at insertion and returns 500 instead of being replayed; add a schema migration that permits `board` before enabling this kind.</comment>

<file context>
@@ -48,6 +48,16 @@ pub const MAX_REPLAY_STRING: usize = 16 * 1024;
+    /// board as vectors is a few kilobytes, and it arrives as the operations
+    /// that produced it, so the review page can redraw any moment of the
+    /// interview rather than the few it could afford to photograph.
+    Board,
     Tests,
     Stage,
</file context>

Comment thread src/agent/prompts.rs
in `decision`. A closing fence, an END marker or a new heading inside a block is
part of the block, not the end of it."#;

const WHITEBOARD_EXECUTION: &str = r#"NOTHING RAN — this interview was held at a whiteboard. There is no test runner,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: Whiteboard reports still receive the shared rubric that treats the absent editor as a failed solution and requires working code for HIRE. Make the report system instruction mode-aware; the brief's board reinterpretation cannot override these higher-priority instructions.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/agent/prompts.rs, line 2054:

<comment>Whiteboard reports still receive the shared rubric that treats the absent editor as a failed solution and requires working code for HIRE. Make the report system instruction mode-aware; the brief's board reinterpretation cannot override these higher-priority instructions.</comment>

<file context>
@@ -1733,6 +2013,90 @@ pub struct ReportPromptInput<'a> {
+in `decision`. A closing fence, an END marker or a new heading inside a block is
+part of the block, not the end of it."#;
+
+const WHITEBOARD_EXECUTION: &str = r#"NOTHING RAN — this interview was held at a whiteboard. There is no test runner,
+no compiler and no pass count, so there is no execution account to weigh and none
+is to be inferred. What stands in its place is the trace the candidate walked
</file context>

Comment thread src/livekit/board.rs
// dropping it. The room loop stays free either way because this is the
// spawned drain task.
if checkpoint.is_some() {
let _ = tx.send(snapshot).await;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: Repeated valid checkpoint uploads can accumulate one blocked task and up to 512 KiB per snapshot after the channel fills. Bound or coalesce checkpoint backlog before retaining snapshots, or a candidate can exhaust the agent's memory.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/livekit/board.rs, line 300:

<comment>Repeated valid checkpoint uploads can accumulate one blocked task and up to 512 KiB per snapshot after the channel fills. Bound or coalesce checkpoint backlog before retaining snapshots, or a candidate can exhaust the agent's memory.</comment>

<file context>
@@ -0,0 +1,385 @@
+    // dropping it. The room loop stays free either way because this is the
+    // spawned drain task.
+    if checkpoint.is_some() {
+        let _ = tx.send(snapshot).await;
+    } else if let Err(TrySendError::Full(_)) = tx.try_send(snapshot) {
+        // The loop is behind and two boards are already waiting, so this one
</file context>

Comment thread src/livekit/board.rs
// spawned drain task.
if checkpoint.is_some() {
let _ = tx.send(snapshot).await;
} else if let Err(TrySendError::Full(_)) = tx.try_send(snapshot) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: This drops the newest board whenever the queue is full, leaving Gemini with older snapshots; if the candidate stops drawing, nothing publishes the final board afterward. Replace an older queued snapshot before enqueueing the newest one.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/livekit/board.rs, line 301:

<comment>This drops the newest board whenever the queue is full, leaving Gemini with older snapshots; if the candidate stops drawing, nothing publishes the final board afterward. Replace an older queued snapshot before enqueueing the newest one.</comment>

<file context>
@@ -0,0 +1,385 @@
+    // spawned drain task.
+    if checkpoint.is_some() {
+        let _ = tx.send(snapshot).await;
+    } else if let Err(TrySendError::Full(_)) = tx.try_send(snapshot) {
+        // The loop is behind and two boards are already waiting, so this one
+        // would be shown after both of them were already stale. See
</file context>

Comment thread web/interview.js
/// everything this one would have, where the publish queue would deliver a
/// stale board first.
async function publishBoard(blob, strokes, checkpoint) {
if (!state.room || !state.connected || !blob) return;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1: A board settled during a LiveKit reconnect is discarded and never resent after recovery. Retain the latest board or enqueue a fresh board snapshot from the Reconnected handler so the interviewer catches up even when the candidate makes no further edits.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/interview.js, line 3220:

<comment>A board settled during a LiveKit reconnect is discarded and never resent after recovery. Retain the latest board or enqueue a fresh board snapshot from the `Reconnected` handler so the interviewer catches up even when the candidate makes no further edits.</comment>

<file context>
@@ -2797,12 +2905,345 @@ function paintEditor() {
+/// everything this one would have, where the publish queue would deliver a
+/// stale board first.
+async function publishBoard(blob, strokes, checkpoint) {
+  if (!state.room || !state.connected || !blob) return;
+  const bytes = new Uint8Array(await blob.arrayBuffer());
+  board.sequence += 1;
</file context>

strokes: 6,
checkpoint: None,
};
pump_board(&mut board, &mut gemini, &mut state, second)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The "too soon to go out" assertion assumes less than one wall-clock second elapses between the first pump_board and the second, because too_soon compares BOARD_SEND_INTERVAL against Instant::now() inside pump_board. The gap is normally a few milliseconds, but a scheduling stall on a loaded CI host past one second makes the second board get sent and this assertion fails spuriously. Make the timing deterministic (e.g., a pauseable clock or an injectable interval) or assert the throttle through too_soon directly rather than through a wall-clock race.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/unit/livekit/board.rs, line 335:

<comment>The "too soon to go out" assertion assumes less than one wall-clock second elapses between the first `pump_board` and the second, because `too_soon` compares `BOARD_SEND_INTERVAL` against `Instant::now()` inside `pump_board`. The gap is normally a few milliseconds, but a scheduling stall on a loaded CI host past one second makes the second board get sent and this assertion fails spuriously. Make the timing deterministic (e.g., a pauseable clock or an injectable interval) or assert the throttle through `too_soon` directly rather than through a wall-clock race.</comment>

<file context>
@@ -0,0 +1,404 @@
+        strokes: 6,
+        checkpoint: None,
+    };
+    pump_board(&mut board, &mut gemini, &mut state, second)
+        .await
+        .unwrap();
</file context>

Comment thread web/index.html
Whiteboard
</button>
</section>
<p class="duration-note" id="mode-note" hidden></p>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The new mode-note paragraph mirrors duration-note directly above it, but omits its aria-live="polite" annotation. When the candidate toggles a mode, the note text appears or disappears without any announcement, while every sibling status element on the page (duration-note, grounding statuses, progress-summary) is announced. Add aria-live="polite" so the whiteboard explanation is announced to screen-reader users.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/index.html, line 962:

<comment>The new `mode-note` paragraph mirrors `duration-note` directly above it, but omits its `aria-live="polite"` annotation. When the candidate toggles a mode, the note text appears or disappears without any announcement, while every sibling status element on the page (`duration-note`, grounding statuses, `progress-summary`) is announced. Add `aria-live="polite"` so the whiteboard explanation is announced to screen-reader users.</comment>

<file context>
@@ -939,7 +947,19 @@ <h1>Practice a live technical interview</h1>
+          Whiteboard
+        </button>
       </section>
+      <p class="duration-note" id="mode-note" hidden></p>
 
       <section class="start-row">
</file context>

Comment thread web/whiteboard.js
/// other, and on an opaque board it is indistinguishable from the real thing.
/// The alternative, compositing with `destination-out`, punches holes that are
/// transparent and so come out black in the JPEG.
export function strokeColor(tool, color) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: This production module exposes strokeColor, strokeWidth, and drawStroke without any runtime consumer; the first two are used directly only by tests, and drawStroke is only called internally. Keep these helpers private and exercise them through createBoard() and drawBoard() so the production API stays focused.

(Based on your team's feedback about keeping test-only helpers out of the production module API.) [ada3c0d3-a44b-4ef2-9575-7c3e2ed7190c]

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/whiteboard.js, line 68:

<comment>This production module exposes `strokeColor`, `strokeWidth`, and `drawStroke` without any runtime consumer; the first two are used directly only by tests, and `drawStroke` is only called internally. Keep these helpers private and exercise them through `createBoard()` and `drawBoard()` so the production API stays focused.

(Based on your team's feedback about keeping test-only helpers out of the production module API.) [ada3c0d3-a44b-4ef2-9575-7c3e2ed7190c]</comment>

<file context>
@@ -0,0 +1,341 @@
+/// other, and on an opaque board it is indistinguishable from the real thing.
+/// The alternative, compositing with `destination-out`, punches holes that are
+/// transparent and so come out black in the JPEG.
+export function strokeColor(tool, color) {
+  return tool === "eraser" ? BOARD_BACKGROUND : color;
+}
</file context>

assert.match(card, /Your board, step by step/);
assert.match(card, /<p>Board work<\/p>/);
assert.doesNotMatch(card, /<p>Coding<\/p>/);
for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases"]) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The step-name loop covers five of the six whiteboard steps but misses "Complexity" (the renamed optimizations step), so a wrong label on that step would pass the suite. Add it to the list.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/browser/render.test.js, line 2019:

<comment>The step-name loop covers five of the six whiteboard steps but misses "Complexity" (the renamed `optimizations` step), so a wrong label on that step would pass the suite. Add it to the list.</comment>

<file context>
@@ -1902,3 +1902,155 @@ test("each transcript row is classed by who is speaking", () => {
+  assert.match(card, /Your board, step by step/);
+  assert.match(card, /<p>Board work<\/p>/);
+  assert.doesNotMatch(card, /<p>Coding<\/p>/);
+  for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases"]) {
+    assert.match(card, new RegExp(`\\d\\. ${name}</strong>`));
+  }
</file context>
Suggested change
for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases"]) {
for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases", "Complexity"]) {

assert.match(card, /4\. Trace<\/strong><span>55 \/ 100<\/span>/);
assert.match(card, /Walked the second example across the table\./);
// Two step images and the final board, and nothing from the refused one.
assert.equal(card.match(/<img /g).length, 3);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: card.match(/<img /g) returns null, not [], when no step image renders, so .length throws a bare TypeError that hides the failing assertion. Guard the match before reading .length.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/browser/render.test.js, line 2025:

<comment>`card.match(/<img /g)` returns `null`, not `[]`, when no step image renders, so `.length` throws a bare TypeError that hides the failing assertion. Guard the match before reading `.length`.</comment>

<file context>
@@ -1902,3 +1902,155 @@ test("each transcript row is classed by who is speaking", () => {
+  assert.match(card, /4\. Trace<\/strong><span>55 \/ 100<\/span>/);
+  assert.match(card, /Walked the second example across the table\./);
+  // Two step images and the final board, and nothing from the refused one.
+  assert.equal(card.match(/<img /g).length, 3);
+  assert.doesNotMatch(card, /example\.test/);
+  assert.match(card, /Your final board/);
</file context>
Suggested change
assert.equal(card.match(/<img /g).length, 3);
assert.equal((card.match(/<img /g) || []).length, 3);

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add a whiteboard interview mode

2 participants