Conversation
|
Need some time for testing and self-review; will mark this PR as ready for review once it's set. |
|
Perhaps we will need to record the state of the whiteboard at each stage so that we can provide more information to the user at the end of the interview. |
|
Additionally, pseudocode should be graded separately rather than being combined with optimization like it is currently. |
|
On the concern from #65 about whether Gemini can read a drawn board reliably: one option worth considering is Excalidraw as the board layer. The drawing tools are not the point; the data model is. An Excalidraw scene is a list of typed elements, and boxes, arrows (bound to the boxes they connect) and text keep their actual strings. So next to the JPEG, the agent can get a plain-text summary of the board, e.g. I checked whether it fits the vendoring rules, following the
The costs:
Only the board layer changes. The byte stream, |
|
Hi, @alanhc Thanks for the discussion! Excalidraw was also brought up in the discussions under the Facebook post back then. After looking into it, I think it seems like a solid option to consider, but I also have a few thoughts. If our goal is to simulate an actual interview environment as closely as possible, would Excalidraw make things too convenient for the user? After all, in a real interview, you only have a marker and a physical whiteboard. That raw experience is precisely what we want to deliver in this mode—allowing users to practice explaining their problem-solving approach and thought process purely through drawing on a blank board. Currently, a tester, @Eason0729 , has tested this feature (the non-Excalidraw version) and provided a lot of feedback. Here is a brief summary of the key concerns raised from the testing:
For now, I'll focus on addressing the points raised in the feedback first. As for Excalidraw, we can discuss it further as a potential future improvement. Overall, I actually think this is quite a promising direction. However, my primary concern is that users should have an experience that feels as close to reality as possible; we probably shouldn't compromise on that just to make recognition easier for the model. WDYH |
897c017 to
2228cf7
Compare
A whiteboard interview takes the same problem bank and the same six steps and swaps the editor and the test runner for a board. Nothing runs, so correctness is what the candidate can defend by tracing an example across their own drawing. The board travels as a LiveKit byte stream rather than on a data topic, because one board is tens of kilobytes and a data packet carries fifteen, and it reaches Gemini the way a camera frame does, as a realtime image. A tool response is JSON and cannot carry a picture, so read_board asks for the board to be sent again and answers with what an image cannot say: how much is on it and how long ago it was drawn. Bundle 27, live prompt 19, follows the surface. A whiteboard session is told it has no editor and no test runner, is offered read_board in place of read_editor, and records board_snapshot where the other records an editor snapshot or a test event; neither may record the other's source. The phases about written work are gated on strokes instead of on characters, and Test on the cases named against the drawing instead of on a run, or an interview with no editor would be refused the second half of its own flow. The report still reads the transcript alone and the replay still carries no board.
The whiteboard was live but nowhere else: the report was written from a transcript and an editor nobody opened, the recording kept no drawing, and the checklist beside the timer called step four Coding while the interviewer was asking the candidate to trace. The board now rides the report request as an inline image, ahead of the brief and on every repair, and the brief says what it is looking at: nothing ran, so correctness is the trace the candidate walked, and the phases keep their names with Coding meaning that trace, Test the cases named against the drawing, and Optimizations the complexity they confirmed. The system instruction is left as it is, so every report call still shares its prefix, and the brief tells a whiteboard reviewer how to read the rules that speak of code. That is report prompt 16, inside the same bundle 27, and the reviewer is told when no board arrived rather than sent looking for an attachment that is not there. The recording keeps the drawing as the operations that made it. One board as a JPEG is past the per-event ceiling on its own and would spend the whole per-interview budget on a handful of frames; the same board as strokes is a few kilobytes, so the replay page and the recording template redraw any moment of the interview with the module the candidate drew on, rather than the few moments a photograph could afford.
A final board cannot represent work cleared between phases. Keep one JPEG checkpoint per phase for grading and a vector marker for replay, so review survives a clear without putting images into the replay budget. Clear itself becomes one undoable step: with the earlier phases kept, an accidental clear is the costliest mistake left at the board.
One fixed width was either too wide to take out a character or too narrow to clear a region. Three offered sizes cover both, and a fixed list rather than a slider keeps every width on the replay one a button can make. The cursor becomes a circle the size of the erasure.
A whiteboard report read as an editor one: a Coding score, REACTO step names the candidate was never shown, and one board folded at the foot of the card. The page already captured the board as each step closed, for the grader, and kept none of it. The report now opens on the board, step by step: the image as that step closed, its score and what the interviewer wrote down, then the final board. Steps, plans and evidence take the whiteboard names, and the score is Board work. Only the display changes; the report keeps the REACTO ids it is graded against. The images are never saved, so a report reopened from history says where they are instead.
2228cf7 to
a124a50
Compare
There was a problem hiding this comment.
48 issues found across 57 files
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="docs/interview-contract-versions.md">
<violation number="1" location="docs/interview-contract-versions.md:15">
P2: `session_timing` is also valid for skipped STAR evidence, so this overstates the whiteboard source restriction. Qualify the claim as applying to observed evidence or mention the skipped-evidence exception.</violation>
</file>
<file name="src/recording/replay.rs">
<violation number="1" location="src/recording/replay.rs:60">
P1: `Board` is accepted by replay parsing, but the `replay_events` schema still rejects `kind = "board"`. Every whiteboard batch therefore fails at insertion and returns 500 instead of being replayed; add a schema migration that permits `board` before enabling this kind.</violation>
</file>
<file name="tests/recording.rs">
<violation number="1" location="tests/recording.rs:3600">
P3: Board is now in the schema list but never asserted in either classification loop, so its is_snapshot() result (false — it Accumulates in replay.rs:109) has no test pinning it. Add ReplayKind::Board to the accumulates list so the new kind's late-join semantics are locked down like the other six.</violation>
</file>
<file name="tests/agent/report.rs">
<violation number="1" location="tests/agent/report.rs:285">
P3: This contract test now covers only the Coding variant of the mode-parameterized `greeting`. The Whiteboard branch this PR adds carries the same obligations (hint offer, no published material, introduce THE EXERCISE), but it is not exercised here, and its text would actually fail the existing assertions: it capitalizes "Do not volunteer a constraint, edge case, or hint" sentence-initially while the assertions match the Coding greeting's lowercase mid-sentence form. Iterate over both `InterviewMode` variants in this test so the new mode's greeting is held to the same contract (using a case-insensitive match or aligning the variant text).</violation>
</file>
<file name="web/replay.js">
<violation number="1" location="web/replay.js:328">
P2: Empty whiteboard sessions are classified as coding because this flag only looks for board moments. Show the empty whiteboard for whiteboard recordings too, using a persisted recording-mode signal rather than requiring a drawing operation.</violation>
<violation number="2" location="web/replay.js:488">
P2: The report can show an earlier selected moment as “Your final board”: `loadReport()` waits for `/api/reports` while timeline buttons remain clickable, and `showMoment()` redraws this canvas before this line reads it. Capture the final board image before the await or restore the final moment before rendering the report.</violation>
</file>
<file name="tests/browser/account.test.js">
<violation number="1" location="tests/browser/account.test.js:141">
P2: The replaced guards removed the only checks on the lobby→room mode link without adding positive ones. `app.js` must do `destination.searchParams.set("mode", mode)` and `interview.js` must read `params.get("mode")` for the selected surface to survive into the token payload; if either line is dropped (mode hardcoded to "coding", or the URL value never set), a whiteboard interview silently runs as a coding interview while this test stays green because it only asserts the data-mode vocabulary and the `interviewMode: mode` string in the payload. Add `assert.match(app, /destination\.searchParams\.set\("mode", mode\)/)` and `assert.match(interview, /params\.get\("mode"\)/)`.</violation>
<violation number="2" location="tests/browser/account.test.js:191">
P3: The test name now contradicts what it asserts: the lobby offers two surfaces ("coding", "whiteboard") and `interviewMode: mode` is carried to the room, yet the title still reads "offers one interview and carries no mode to the room". The body comment was updated for the new contract; the title was not. Rename it (e.g. "the lobby offers coding or whiteboard and carries the mode into the room").</violation>
</file>
<file name="src/gemini.rs">
<violation number="1" location="src/gemini.rs:1458">
P2: Whiteboard interviews still expose a `log_hint` description that tells Jim to give the clue with the candidate's editor, which does not exist in this mode. Make the hint description mode-specific and direct him to fit the clue to the board.</violation>
<violation number="2" location="src/gemini.rs:1473">
P2: This declaration is sent in whiteboard mode, but its description still tells Gemini that evidence can come from an editor or test run. Make the description mode-specific so the interviewer is not guided toward nonexistent evidence sources.</violation>
</file>
<file name="src/livekit/report.rs">
<violation number="1" location="src/livekit/report.rs:71">
P2: Whiteboard images now reach Gemini but never enter `record_model_input`, so the reported model-input byte total omits every attached board. Include board payload bytes in the report metrics so the total remains accurate.</violation>
</file>
<file name="web/render.js">
<violation number="1" location="web/render.js:398">
P2: The board report hides explicitly recorded `skipped` evidence and presents that phase as missing. Keep the skipped outcome visible, or render an explicit skipped status, instead of treating a recorded outcome as absent.
(Based on your team's feedback about skipped framework evidence.) [5de29b75-9991-48cb-96b6-10a6ce846cb4].</violation>
<violation number="2" location="web/render.js:431">
P2: Saved whiteboard exports can contain no images, and recording is optional, so this promises access to a board that may not exist anywhere. Say the images are not included, and qualify replay availability on having a recording; update the matching history-card fallback too.</violation>
</file>
<file name="web/replay.html">
<violation number="1" location="web/replay.html:108">
P2: `<canvas id="replay-board">` has no accessible name: the `<h3>` heading above it does not label a following canvas in the accessibility tree, so a screen reader announces an unnamed canvas and a user has no indication that a board drawing is shown. The constraint cited in the comment does not actually block an `aria-label` — `history.test.js` already allowlists `"Whiteboard"` (it is used for the heading), and this page already carries attributes such as `role="status"`. Give the canvas `aria-label="Whiteboard"`, matching the live board in `web/interview.html`.</violation>
</file>
<file name="scripts/browser-check.cjs">
<violation number="1" location="scripts/browser-check.cjs:448">
P2: `checkWhiteboardInterview` never clicks Undo; it only observes that the model enabled the button, so a broken DOM handler still passes. Click Undo and assert the stroke returns, then exercise Redo as well.</violation>
</file>
<file name="tests/unit/livekit/board.rs">
<violation number="1" location="tests/unit/livekit/board.rs:335">
P3: The "too soon to go out" assertion assumes less than one wall-clock second elapses between the first `pump_board` and the second, because `too_soon` compares `BOARD_SEND_INTERVAL` against `Instant::now()` inside `pump_board`. The gap is normally a few milliseconds, but a scheduling stall on a loaded CI host past one second makes the second board get sent and this assertion fails spuriously. Make the timing deterministic (e.g., a pauseable clock or an injectable interval) or assert the throttle through `too_soon` directly rather than through a wall-clock race.</violation>
</file>
<file name="src/agent/prompts.rs">
<violation number="1" location="src/agent/prompts.rs:192">
P2: The whiteboard mode selects this surface, but the shared interview flow still tells the interviewer to “let them code” after a solid answer despite there being no editor. Replace that shared transition with the whiteboard equivalent: continue drawing and tracing.</violation>
<violation number="2" location="src/agent/prompts.rs:256">
P2: This note says every board change reaches Gemini, but rate-limited snapshots can leave Gemini viewing an older image after the candidate stops drawing. State that snapshots may lag and require `read_board` for confirmation, or schedule the deferred latest send.</violation>
<violation number="3" location="src/agent/prompts.rs:1984">
P2: Whiteboard mode reaches the final report, but idle reviews still receive an editor-only prompt with an empty editor. Adapt or skip interim reviews for whiteboards so their rolling assessment does not omit board work or report a misleading surface.</violation>
<violation number="4" location="src/agent/prompts.rs:2054">
P1: Whiteboard reports still receive the shared rubric that treats the absent editor as a failed solution and requires working code for HIRE. Make the report system instruction mode-aware; the brief's board reinterpretation cannot override these higher-priority instructions.</violation>
</file>
<file name="src/agent.rs">
<violation number="1" location="src/agent.rs:1548">
P2: `board_age_seconds` measures snapshot receipt, not when Gemini actually received the image. During the throttle, `read_board` can report a fresh board while the model still sees the previous frame; update the timestamp only after a successful send, or track receipt and send ages separately.</violation>
<violation number="2" location="src/agent.rs:1561">
P2: `board_strokes` comes from the candidate-controlled stream attribute and is not checked against the image or drawing data. A client can claim three strokes with an empty board, passing this guard and allowing Coding, Test, or Optimizations evidence; validate the count before using it as a gate.</violation>
</file>
<file name="src/livekit.rs">
<violation number="1" location="src/livekit.rs:2106">
P2: Snapshots received while paused are dropped, and resume does not send the current board again. Since the board remains editable, Jim misses work drawn during the pause until another edit produces a snapshot; retain the latest snapshot and send it on resume.</violation>
<violation number="2" location="src/livekit.rs:2918">
P1: The report snapshots `Board` without first consuming completed board streams already queued from the reader tasks. An end packet can therefore freeze a stale board (or no board), causing whiteboard reports to miss the candidate’s final work; drain or otherwise await the board queue before freezing report attachments.</violation>
</file>
<file name="tests/golden/prompts.json">
<violation number="1" location="tests/golden/prompts.json:7">
P2: This no-board prompt says no board was available, then asserts all six phases ran and that the candidate traced, tested, and confirmed complexity on it. That can make report generation invent board evidence; make this instruction conditional on attached evidence and otherwise direct scoring from the transcript and rolling assessment only.</violation>
<violation number="2" location="tests/golden/prompts.json:8">
P2: boardSilenceDrawn freezes code-mode test claims into a whiteboard prompt: "tests: browser-reported claims (unverified): 1 of 3 passing, 1 edit-and-run cycles" contradicts board mode, where nothing runs and there is no edit-and-run. The evidence string passed to `board_silence_nudge` in `prompt_samples` (tests/agent.rs:308-309) reuses the coding sample's `working` string. Give the board sample a board-appropriate evidence string (board strokes/phases, tests: not run) and regenerate the golden, so goldens stop asserting editor/test semantics for the whiteboard mode.</violation>
</file>
<file name="tests/browser/recording-template.test.js">
<violation number="1" location="tests/browser/recording-template.test.js:288">
P2: The file now claims "seven kinds, seven call sites" and adds `board` to the rendered kinds, but unlike every other kind (envelope, transcript, editor, tests, stage, avatar, lifecycle) `board` has no `replay-producer` test. Nothing pins the producer payload or drain: `recordReplay("board", { ops, checkpoint })` in `web/interview.js` `recordBoardOps` would drift (e.g. renaming `ops`, dropping the empty-ops checkpoint emit) without this tripwire failing, and a payload-shape change would break the template's `payload.ops` in production only. Add a producer test in the file's existing style, asserting `recordBoardOps` sends `recordReplay("board", ...)`, batches via `board.model.takeOps()`, and attaches the checkpoint to the last batch.</violation>
</file>
<file name="scripts/gen-wire-fixtures.mjs">
<violation number="1" location="scripts/gen-wire-fixtures.mjs:380">
P2: The new fixture carries a `checkpoint: "algorithm"` attribute, but `reads_the_stroke_count_the_browser_sent` in tests/unit/livekit/board.rs only asserts the topic and stroke counts. If the producer renamed the `checkpoint` attribute (e.g. to `phase`), `gen-wire-fixtures.mjs` would regenerate the fixture, `--check` and `cargo test` would pass, and the agent would silently stop labeling checkpoint images in whiteboard report generation. The comment says the fixture pins "the attributes the agent reads off it"; assert the per-case `checkpoint` value in the Rust test so that key is actually pinned.</violation>
</file>
<file name="web/recording/index.html">
<violation number="1" location="web/recording/index.html:35">
P2: The `hidden` attribute on this canvas has no effect: `recording.css` sets `.board { display: block; }`, an author rule that overrides the user-agent `[hidden] { display: none }`. Recording switches the panels only via the `hidden` property (`recording.js:201-203`), so in editor interviews the recording shows an empty whiteboard panel for the entire video — the layout the comment above this element says it prevents. `web/styles.css:23` works around this exact case with `[hidden] { display: none !important; }`, but the recording page loads only `recording.css`. Add the same reset to `web/recording/recording.css`.</violation>
</file>
<file name="tests/browser/whiteboard.test.js">
<violation number="1" location="tests/browser/whiteboard.test.js:70">
P2: This only proves the returned array is detached; its stroke objects and `points` are still shared, so mutating a returned point silently changes the board without journaling. Also mutate a returned point and assert the board remains unchanged to cover the stated copy contract.</violation>
</file>
<file name="web/index.html">
<violation number="1" location="web/index.html:952">
P2: The default Editor selection is only a CSS class until a mode button is clicked, so screen readers cannot tell which mode Start will use. Set initial `aria-pressed` values on both mode buttons; the existing click handler keeps them updated.</violation>
<violation number="2" location="web/index.html:962">
P3: The new `mode-note` paragraph mirrors `duration-note` directly above it, but omits its `aria-live="polite"` annotation. When the candidate toggles a mode, the note text appears or disappears without any announcement, while every sibling status element on the page (`duration-note`, grounding statuses, `progress-summary`) is announced. Add `aria-live="polite"` so the whiteboard explanation is announced to screen-reader users.</violation>
</file>
<file name="tests/interview_behavior.rs">
<violation number="1" location="tests/interview_behavior.rs:996">
P2: The behavior suite's mode plumbing is adapted to pass InterviewMode, but every new call site hard-codes InterviewMode::Coding, so none of the new whiteboard surfaces (whiteboard greeting, read_board tool, mode-specific instruction text) are exercised by the only tests that assert prompt behavior. Add a whiteboard variant to one of the live tests (e.g. loop interview_mode in `live_interviewer_poses_the_variant_and_serves_hints_in_order` alongside interview_loop) so a whiteboard prompt regression cannot land uncaught.</violation>
</file>
<file name="src/livekit/turn.rs">
<violation number="1" location="src/livekit/turn.rs:1155">
P2: `board_strokes` can describe a throttled snapshot that Gemini never received, but this prompt tells Gemini it has the latest image. Track the last-sent board state or resend the latest image before making that claim.</violation>
</file>
<file name="src/livekit/board.rs">
<violation number="1" location="src/livekit/board.rs:242">
P1: Each accepted board stream gets its own 512 KiB buffer, but the queue does not cap concurrent drain tasks and checkpoint tasks wait rather than being rejected. A candidate can open many streams and exhaust agent memory; cap active readers (for example with a semaphore) before spawning, and release the permit when draining finishes.</violation>
<violation number="2" location="src/livekit/board.rs:300">
P1: Repeated valid checkpoint uploads can accumulate one blocked task and up to 512 KiB per snapshot after the channel fills. Bound or coalesce checkpoint backlog before retaining snapshots, or a candidate can exhaust the agent's memory.</violation>
<violation number="3" location="src/livekit/board.rs:301">
P1: This drops the newest board whenever the queue is full, leaving Gemini with older snapshots; if the candidate stops drawing, nothing publishes the final board afterward. Replace an older queued snapshot before enqueueing the newest one.</violation>
</file>
<file name="web/interview.js">
<violation number="1" location="web/interview.js:2989">
P2: The editor is disabled during a pause, but this board input path and its toolbar remain active, and completed snapshots are still sent to Jim. Disable the board controls and reject pointer edits until the interview resumes.</violation>
<violation number="2" location="web/interview.js:3220">
P1: A board settled during a LiveKit reconnect is discarded and never resent after recovery. Retain the latest board or enqueue a fresh board snapshot from the `Reconnected` handler so the interviewer catches up even when the candidate makes no further edits.</violation>
<violation number="3" location="web/interview.js:3235">
P2: Whiteboard recordings still emit an empty editor snapshot at connection time. Skip editor replay events for whiteboard sessions, otherwise replay shows an irrelevant empty Code moment alongside the board timeline.</violation>
</file>
<file name="web/whiteboard.js">
<violation number="1" location="web/whiteboard.js:68">
P3: This production module exposes `strokeColor`, `strokeWidth`, and `drawStroke` without any runtime consumer; the first two are used directly only by tests, and `drawStroke` is only called internally. Keep these helpers private and exercise them through `createBoard()` and `drawBoard()` so the production API stays focused.
(Based on your team's feedback about keeping test-only helpers out of the production module API.) [ada3c0d3-a44b-4ef2-9575-7c3e2ed7190c]</violation>
<violation number="2" location="web/whiteboard.js:128">
P2: `undo()` loses the clear redo marker when it undoes a prior stroke after restoring a clear. After Clear → Undo → Undo → Redo, the next Redo is unavailable, so the candidate cannot restore the cleared board.</violation>
</file>
<file name="web/recording/recording.js">
<violation number="1" location="web/recording/recording.js:201">
P2: The recorder leaves the blank code surface visible until the first `board` event. An untouched or early whiteboard recording therefore shows a code panel instead of the whiteboard; pass the interview mode to the recording template or emit an initial board marker before recording starts.</violation>
</file>
<file name="src/livekit/session.rs">
<violation number="1" location="src/livekit/session.rs:985">
P2: `read_board` drops the pending candidate context when Gemini context compression is enabled. Add `TOOL_READ_BOARD` to the continuity-response set so the model can continue the current exchange after reading the refreshed board.</violation>
<violation number="2" location="src/livekit/session.rs:1034">
P2: A requested whiteboard hint can use a stale drawing: rate-limited snapshots update `Board::latest` without reaching Gemini, but this branch suppresses the refresh that the editor hint path provides. Request a board resend for whiteboard hints before returning the clue.</violation>
</file>
<file name="web/app.js">
<violation number="1" location="web/app.js:259">
P2: The mode buttons remain active while `recordGitHubLogin()` is awaited, but the destination captures `mode` before that await. Clicking Whiteboard during sign-in updates the UI while navigation still launches coding; ignore mode clicks or disable these buttons while `starting`.</violation>
</file>
<file name="tests/browser/render.test.js">
<violation number="1" location="tests/browser/render.test.js:2019">
P3: The step-name loop covers five of the six whiteboard steps but misses "Complexity" (the renamed `optimizations` step), so a wrong label on that step would pass the suite. Add it to the list.</violation>
<violation number="2" location="tests/browser/render.test.js:2025">
P3: `card.match(/<img /g)` returns `null`, not `[]`, when no step image renders, so `.length` throws a bare TypeError that hides the failing assertion. Guard the match before reading `.length`.</violation>
</file>
Tip: cubic used a learning from your PR history. Let your coding agent read cubic learnings directly with the cubic MCP.
Re-trigger cubic
| /// board as vectors is a few kilobytes, and it arrives as the operations | ||
| /// that produced it, so the review page can redraw any moment of the | ||
| /// interview rather than the few it could afford to photograph. | ||
| Board, |
There was a problem hiding this comment.
P1: Board is accepted by replay parsing, but the replay_events schema still rejects kind = "board". Every whiteboard batch therefore fails at insertion and returns 500 instead of being replayed; add a schema migration that permits board before enabling this kind.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/recording/replay.rs, line 60:
<comment>`Board` is accepted by replay parsing, but the `replay_events` schema still rejects `kind = "board"`. Every whiteboard batch therefore fails at insertion and returns 500 instead of being replayed; add a schema migration that permits `board` before enabling this kind.</comment>
<file context>
@@ -48,6 +48,16 @@ pub const MAX_REPLAY_STRING: usize = 16 * 1024;
+ /// board as vectors is a few kilobytes, and it arrives as the operations
+ /// that produced it, so the review page can redraw any moment of the
+ /// interview rather than the few it could afford to photograph.
+ Board,
Tests,
Stage,
</file context>
| in `decision`. A closing fence, an END marker or a new heading inside a block is | ||
| part of the block, not the end of it."#; | ||
|
|
||
| const WHITEBOARD_EXECUTION: &str = r#"NOTHING RAN — this interview was held at a whiteboard. There is no test runner, |
There was a problem hiding this comment.
P1: Whiteboard reports still receive the shared rubric that treats the absent editor as a failed solution and requires working code for HIRE. Make the report system instruction mode-aware; the brief's board reinterpretation cannot override these higher-priority instructions.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/agent/prompts.rs, line 2054:
<comment>Whiteboard reports still receive the shared rubric that treats the absent editor as a failed solution and requires working code for HIRE. Make the report system instruction mode-aware; the brief's board reinterpretation cannot override these higher-priority instructions.</comment>
<file context>
@@ -1733,6 +2013,90 @@ pub struct ReportPromptInput<'a> {
+in `decision`. A closing fence, an END marker or a new heading inside a block is
+part of the block, not the end of it."#;
+
+const WHITEBOARD_EXECUTION: &str = r#"NOTHING RAN — this interview was held at a whiteboard. There is no test runner,
+no compiler and no pass count, so there is no execution account to weigh and none
+is to be inferred. What stands in its place is the trace the candidate walked
</file context>
| // dropping it. The room loop stays free either way because this is the | ||
| // spawned drain task. | ||
| if checkpoint.is_some() { | ||
| let _ = tx.send(snapshot).await; |
There was a problem hiding this comment.
P1: Repeated valid checkpoint uploads can accumulate one blocked task and up to 512 KiB per snapshot after the channel fills. Bound or coalesce checkpoint backlog before retaining snapshots, or a candidate can exhaust the agent's memory.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/livekit/board.rs, line 300:
<comment>Repeated valid checkpoint uploads can accumulate one blocked task and up to 512 KiB per snapshot after the channel fills. Bound or coalesce checkpoint backlog before retaining snapshots, or a candidate can exhaust the agent's memory.</comment>
<file context>
@@ -0,0 +1,385 @@
+ // dropping it. The room loop stays free either way because this is the
+ // spawned drain task.
+ if checkpoint.is_some() {
+ let _ = tx.send(snapshot).await;
+ } else if let Err(TrySendError::Full(_)) = tx.try_send(snapshot) {
+ // The loop is behind and two boards are already waiting, so this one
</file context>
| // spawned drain task. | ||
| if checkpoint.is_some() { | ||
| let _ = tx.send(snapshot).await; | ||
| } else if let Err(TrySendError::Full(_)) = tx.try_send(snapshot) { |
There was a problem hiding this comment.
P1: This drops the newest board whenever the queue is full, leaving Gemini with older snapshots; if the candidate stops drawing, nothing publishes the final board afterward. Replace an older queued snapshot before enqueueing the newest one.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/livekit/board.rs, line 301:
<comment>This drops the newest board whenever the queue is full, leaving Gemini with older snapshots; if the candidate stops drawing, nothing publishes the final board afterward. Replace an older queued snapshot before enqueueing the newest one.</comment>
<file context>
@@ -0,0 +1,385 @@
+ // spawned drain task.
+ if checkpoint.is_some() {
+ let _ = tx.send(snapshot).await;
+ } else if let Err(TrySendError::Full(_)) = tx.try_send(snapshot) {
+ // The loop is behind and two boards are already waiting, so this one
+ // would be shown after both of them were already stale. See
</file context>
| /// everything this one would have, where the publish queue would deliver a | ||
| /// stale board first. | ||
| async function publishBoard(blob, strokes, checkpoint) { | ||
| if (!state.room || !state.connected || !blob) return; |
There was a problem hiding this comment.
P1: A board settled during a LiveKit reconnect is discarded and never resent after recovery. Retain the latest board or enqueue a fresh board snapshot from the Reconnected handler so the interviewer catches up even when the candidate makes no further edits.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/interview.js, line 3220:
<comment>A board settled during a LiveKit reconnect is discarded and never resent after recovery. Retain the latest board or enqueue a fresh board snapshot from the `Reconnected` handler so the interviewer catches up even when the candidate makes no further edits.</comment>
<file context>
@@ -2797,12 +2905,345 @@ function paintEditor() {
+/// everything this one would have, where the publish queue would deliver a
+/// stale board first.
+async function publishBoard(blob, strokes, checkpoint) {
+ if (!state.room || !state.connected || !blob) return;
+ const bytes = new Uint8Array(await blob.arrayBuffer());
+ board.sequence += 1;
</file context>
| strokes: 6, | ||
| checkpoint: None, | ||
| }; | ||
| pump_board(&mut board, &mut gemini, &mut state, second) |
There was a problem hiding this comment.
P3: The "too soon to go out" assertion assumes less than one wall-clock second elapses between the first pump_board and the second, because too_soon compares BOARD_SEND_INTERVAL against Instant::now() inside pump_board. The gap is normally a few milliseconds, but a scheduling stall on a loaded CI host past one second makes the second board get sent and this assertion fails spuriously. Make the timing deterministic (e.g., a pauseable clock or an injectable interval) or assert the throttle through too_soon directly rather than through a wall-clock race.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/unit/livekit/board.rs, line 335:
<comment>The "too soon to go out" assertion assumes less than one wall-clock second elapses between the first `pump_board` and the second, because `too_soon` compares `BOARD_SEND_INTERVAL` against `Instant::now()` inside `pump_board`. The gap is normally a few milliseconds, but a scheduling stall on a loaded CI host past one second makes the second board get sent and this assertion fails spuriously. Make the timing deterministic (e.g., a pauseable clock or an injectable interval) or assert the throttle through `too_soon` directly rather than through a wall-clock race.</comment>
<file context>
@@ -0,0 +1,404 @@
+ strokes: 6,
+ checkpoint: None,
+ };
+ pump_board(&mut board, &mut gemini, &mut state, second)
+ .await
+ .unwrap();
</file context>
| Whiteboard | ||
| </button> | ||
| </section> | ||
| <p class="duration-note" id="mode-note" hidden></p> |
There was a problem hiding this comment.
P3: The new mode-note paragraph mirrors duration-note directly above it, but omits its aria-live="polite" annotation. When the candidate toggles a mode, the note text appears or disappears without any announcement, while every sibling status element on the page (duration-note, grounding statuses, progress-summary) is announced. Add aria-live="polite" so the whiteboard explanation is announced to screen-reader users.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/index.html, line 962:
<comment>The new `mode-note` paragraph mirrors `duration-note` directly above it, but omits its `aria-live="polite"` annotation. When the candidate toggles a mode, the note text appears or disappears without any announcement, while every sibling status element on the page (`duration-note`, grounding statuses, `progress-summary`) is announced. Add `aria-live="polite"` so the whiteboard explanation is announced to screen-reader users.</comment>
<file context>
@@ -939,7 +947,19 @@ <h1>Practice a live technical interview</h1>
+ Whiteboard
+ </button>
</section>
+ <p class="duration-note" id="mode-note" hidden></p>
<section class="start-row">
</file context>
| /// other, and on an opaque board it is indistinguishable from the real thing. | ||
| /// The alternative, compositing with `destination-out`, punches holes that are | ||
| /// transparent and so come out black in the JPEG. | ||
| export function strokeColor(tool, color) { |
There was a problem hiding this comment.
P3: This production module exposes strokeColor, strokeWidth, and drawStroke without any runtime consumer; the first two are used directly only by tests, and drawStroke is only called internally. Keep these helpers private and exercise them through createBoard() and drawBoard() so the production API stays focused.
(Based on your team's feedback about keeping test-only helpers out of the production module API.) [ada3c0d3-a44b-4ef2-9575-7c3e2ed7190c]
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At web/whiteboard.js, line 68:
<comment>This production module exposes `strokeColor`, `strokeWidth`, and `drawStroke` without any runtime consumer; the first two are used directly only by tests, and `drawStroke` is only called internally. Keep these helpers private and exercise them through `createBoard()` and `drawBoard()` so the production API stays focused.
(Based on your team's feedback about keeping test-only helpers out of the production module API.) [ada3c0d3-a44b-4ef2-9575-7c3e2ed7190c]</comment>
<file context>
@@ -0,0 +1,341 @@
+/// other, and on an opaque board it is indistinguishable from the real thing.
+/// The alternative, compositing with `destination-out`, punches holes that are
+/// transparent and so come out black in the JPEG.
+export function strokeColor(tool, color) {
+ return tool === "eraser" ? BOARD_BACKGROUND : color;
+}
</file context>
| assert.match(card, /Your board, step by step/); | ||
| assert.match(card, /<p>Board work<\/p>/); | ||
| assert.doesNotMatch(card, /<p>Coding<\/p>/); | ||
| for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases"]) { |
There was a problem hiding this comment.
P3: The step-name loop covers five of the six whiteboard steps but misses "Complexity" (the renamed optimizations step), so a wrong label on that step would pass the suite. Add it to the list.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/browser/render.test.js, line 2019:
<comment>The step-name loop covers five of the six whiteboard steps but misses "Complexity" (the renamed `optimizations` step), so a wrong label on that step would pass the suite. Add it to the list.</comment>
<file context>
@@ -1902,3 +1902,155 @@ test("each transcript row is classed by who is speaking", () => {
+ assert.match(card, /Your board, step by step/);
+ assert.match(card, /<p>Board work<\/p>/);
+ assert.doesNotMatch(card, /<p>Coding<\/p>/);
+ for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases"]) {
+ assert.match(card, new RegExp(`\\d\\. ${name}</strong>`));
+ }
</file context>
| for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases"]) { | |
| for (const name of ["Repeat", "Example", "Approach", "Trace", "Edge cases", "Complexity"]) { |
| assert.match(card, /4\. Trace<\/strong><span>55 \/ 100<\/span>/); | ||
| assert.match(card, /Walked the second example across the table\./); | ||
| // Two step images and the final board, and nothing from the refused one. | ||
| assert.equal(card.match(/<img /g).length, 3); |
There was a problem hiding this comment.
P3: card.match(/<img /g) returns null, not [], when no step image renders, so .length throws a bare TypeError that hides the failing assertion. Guard the match before reading .length.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At tests/browser/render.test.js, line 2025:
<comment>`card.match(/<img /g)` returns `null`, not `[]`, when no step image renders, so `.length` throws a bare TypeError that hides the failing assertion. Guard the match before reading `.length`.</comment>
<file context>
@@ -1902,3 +1902,155 @@ test("each transcript row is classed by who is speaking", () => {
+ assert.match(card, /4\. Trace<\/strong><span>55 \/ 100<\/span>/);
+ assert.match(card, /Walked the second example across the table\./);
+ // Two step images and the final board, and nothing from the refused one.
+ assert.equal(card.match(/<img /g).length, 3);
+ assert.doesNotMatch(card, /example\.test/);
+ assert.match(card, /Your final board/);
</file context>
| assert.equal(card.match(/<img /g).length, 3); | |
| assert.equal((card.match(/<img /g) || []).length, 3); |
Summary
A whiteboard interview is the coding round held without a compiler. The lobby offers it next to the duration and the loop. The editor, the starter code and the test runner are replaced by a drawing surface, and the candidate shows correctness by drawing an approach, tracing an example across it, and naming the cases that would break it. It uses the same problem bank, the same six REACTO phase ids, the same rubric and the same report schema. Only the surface and what each step asks for change.
Closes #65, with two deliberate departures from the issue's proposal:
What the candidate sees
repeatexamplealgorithmcodingtestoptimizationsHow a board reaches the interviewer
The browser keeps the drawing as vector strokes (
web/whiteboard.js). That is what makes undo and replay possible. Rendering a board to pixels happens only when something needs an image.flowchart LR Stroke["Stroke ends"] --> Settle["1 s with no new stroke"] Settle --> JPEG["Canvas exported as JPEG<br/>quality 0.72"] JPEG -->|"LiveKit byte stream<br/>topic board_image"| Drain["Agent drain task<br/>bounded at 512 KiB"] Drain -->|"queue of 2"| Loop["Room loop: pump_board"] Loop --> Latest["Latest board in memory"] Loop -->|"at most one per second<br/>realtimeInput.video"| Live["Gemini Live"] Tool["read_board tool call"] --> Resend["Resend latest board"] Latest --> Resend Resend --> Livestrokescount and, for a phase checkpoint, acheckpointid.read_boardanswers with it.read_board. Returns JSON with the stroke count, the board's age and the timer, then resends the latest image, because a tool response cannot carry a picture.What the interviewer is told
There is one live prompt. The sentences that name the surface sit in a
Surfacestruct with a coding and a whiteboard value, so a rule cannot end up in one prompt and not the other. Rebased onto #203, #192 and #202: the coding prompt is byte-identical to main's, and only the whiteboard entries are new intests/golden/prompts.json. The thinking-time rule from #192 says "board changes" rather than "editor changes" at a board.read_boardin place ofread_editor. It is told nothing will run.board_snapshotis a valid evidence source only at a board, and Test is recorded from the cases named against the drawing.Phase checkpoints and grading
The report reviewer sees the board as it stood when each step closed, not just the final one. A candidate who clears the board between steps does not lose that evidence.
sequenceDiagram participant B as Browser participant A as Interview agent participant G as Gemini report model A->>B: framework_state (phases with evidence) B->>B: New phase ids, in REACTO order B->>A: JPEG + checkpoint id (after pointerup if mid-stroke) A->>A: Validate the id, keep one image per phase B->>A: Final JPEG when End is pressed B->>A: end_interview, once the uploads finish A->>G: generateContent(labelled inline JPEGs + prompt) Note over A,G: Retries and schema repairs resend every image G-->>A: Structured JSON report A-->>B: Report, stamped with interviewModeframework_stateupdate is reduced to known phase ids that have not been captured yet. A phase that completes while the pointer is down waits for pointerup, so the image and the replay describe the same stroke.end_interviewis sent only after they finish. That way the report is never frozen on an older board.Replay and recording
Replay stores no JPEGs. A board as a JPEG is over the per-event size limit and would use up the per-interview replay budget within a few frames. Replay instead stores the operations that drew the board:
stroke,undo,redoandclear, in batches under 24 KiB, each optionally tagged with a checkpoint id. The replay page and the recording template rebuild any moment of the board with the same module the candidate drew on (createBoard+applyOp). Every operation is validated on the way back in: colour, width, point count and coordinates.The report card
A whiteboard report opens on the board, below the committee summary:
.mdfile lists the steps and their evidence without base64.Testing
cargo test(1204 passed), clippy with-D warnings, the browser suite (839 passed, 1 skipped for no Java), ESLint, the drift checks and the hook suite, andscripts/indent.sh --checkwith commentflow. All passed.tests/unit/livekit/board.rsandtests/agent/whiteboard.rs, and the report card bytests/browser/render.test.js.Whiteboard