Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,26 @@ big-file mode badge. Files: extension.js + fileOps/lineOps/columnOps/encodingEol
- `aiEdit.js` — **edit-with-diff**: select code → `Cmd+Alt+E` → instruction → side-by-side diff → ✓ Keep / ✗ Discard
buttons on the diff toolbar (gated on `levelcode.ai.diffActive`).
- `inlineReview.js` — **dead code** (an inline per-hunk Keep/Undo attempt that was reverted; nothing imports it).
- `diagram/` — **rich diagrams, phase 1** (`docs/RICH-DIAGRAMS.md`). The agent calls a `render_diagram` tool with
STRUCTURE only (nodes, edges, groups, one accent — never coordinates or colours); the editor validates it, lays it
out in one house style and paints it in the chat, themed, with nodes that link to code. Graph JSON only — Mermaid,
Vega-Lite and raw SVG are later phases and are **not built**. Not yet run in the packaged editor or against a live
model; the eval (`scripts/diagram-eval.js`) exists and has not been run.
- Shared UMD modules (`theme` `schema` `validate` `repair` `layout` `scene` `text` `ascii`) run in Node AND are
inlined into `chat.html` by `diagram/bundle.js` under the page's existing nonce — the CSP is unchanged. Host-only:
`tool` (tool + prompt block + result text), `service` (ids, the one repair pass, records, stubs), `links`,
`exportCheck`, `stats`.
- Repair ladder: lossless auto-fix → every error back to the model ONCE → degrade with a banner, or source + Retry.
Never a blank card, never a second automatic repair. Records are stored in the session log and re-validated (not
re-repaired) on reopen, by the host and by the page.
- Layout is in-house (layered + orthogonal routing), **not ELK** (EPL-2.0, ~1.5 MB, no build step here);
`layout.layout()` is the one swap point. The validator is a small JSON-Schema-subset interpreter, not Ajv.
- Gated per run by `client.render` (`rich`|`ascii`): off via `levelcode.ai.diagrams.enabled`, or per model with
`diagrams: false` in `providers/catalog.js` `CAPS`. Costs ~970 tokens of tool + prompt per request while on.
- Checks: `test/diagram*.test.js` (in the gate), `scripts/diagram-browser-check.js` (real page in headless Chrome,
not in the gate), `scripts/diagram-editor-check.js` (the REAL editor: a throwaway instance of the dev build with
this checkout's extension and a stand-in provider), `scripts/diagram-eval.js --dry-run`. Local counters: command
`AI: Diagram Statistics`.

## Deferred / known limits (don't waste time re-hitting these)

Expand All @@ -189,4 +209,13 @@ big-file mode badge. Files: extension.js + fileOps/lineOps/columnOps/encodingEol
- `// @ts-check` + JSDoc at top of JS files.
- Test JS logic with `node --check` and small unit snippets before wiring into the editor.
- After any change, `./scripts/run-dev.sh` to verify; package with `./scripts/build-macos.sh`.
- `run-dev.sh` runs the extensions of the checkout that HAS `vscode/`. A git worktree has none, so work in a worktree
is not in the editor until you load it: `./scripts/run-dev.sh --extensionDevelopmentPath=<worktree>/extensions/levelcode-ai`
(from the main checkout; the dev extension replaces the built-in one). Uncommitted work is not "on the branch" —
checking the branch out somewhere else gets none of it. Run it in the editor before telling anyone to try it.
- Commit `extensions/`, `patches/`, `branding/`, `scripts/`, `docs/`, `PLAN.md`, `CLAUDE.md`. Never commit `vscode/`.
- `extensions/levelcode-ai/diagram/` modules listed in `bundle.FILES` are pasted INTO a script block in `chat.html`.
They must never contain the text of a script tag or an HTML comment opener — not even in a comment — or the block
ends early; `bundle.js` refuses to build if one does. Keep them dependency-free and free of `require('vscode')`/`fs`.
- Host suites slice functions out of `extension.js` with a brace matcher (`extract()` in `test/*Host.test.js`). It
does not understand a backtick inside a regex literal: write `String.fromCharCode(96)` there instead.
450 changes: 450 additions & 0 deletions docs/RICH-DIAGRAMS.md

Large diffs are not rendered by default.

91 changes: 83 additions & 8 deletions extensions/levelcode-ai/agent.js
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,8 @@ const { loadProjectRules } = require('./projectRules');
const { loadServerConfig, buildAgentTools, toolCountsByServer, classifyMcpTool, explainMcpRefusal, describeMcpCall,
isLaunchTrusted, rememberLaunchTrust, describeMcpLaunch } = require('./mcpConfig');
const { connectAll, getServer } = require('./mcpClient');
const diagramTool = require('./diagram/tool');
const diagramRepair = require('./diagram/repair');

const SYSTEM_BASE = [
"You are LevelCode's built-in autonomous coding agent. You accomplish the user's goal in their",
Expand Down Expand Up @@ -307,6 +309,13 @@ function runCommand(root, command, onChunk, onExit, onStart, timeoutMs) {
});
}

/**
* Can the surface on the other end of this run show a diagram? (docs/RICH-DIAGRAMS.md, "Capability
* flag".) Asked in two places — when the tools and the prompt are assembled, and again when a call
* arrives — and they must agree, or a client could be refused a tool it was offered.
*/
function richClient(ctx) { return !!(ctx.diagrams && ctx.client && ctx.client.render === 'rich'); }

/** Execute one tool call; returns a string result for the model. */
async function runTool(tu, ctx) {
const root = ctx.root;
Expand Down Expand Up @@ -517,6 +526,22 @@ async function runTool(tu, ctx) {
try { return String(ctx.recallSessions(query) || 'No matching past sessions in this project.'); }
catch (e) { return 'ERROR: recall failed.'; }
}
// Rich diagrams (docs/RICH-DIAGRAMS.md). The model describes structure; the host validates it,
// climbs the repair ladder and posts what the chat should show. Read-only and instant — it
// draws in the transcript and touches nothing else — so, like update_plan, it never asks.
if (tu.name === diagramTool.RENDER_DIAGRAM.name) {
// Asked for by a client that was never offered it (the setting was turned off mid-conversation,
// or the model remembers the tool from an earlier turn): refuse in words it can act on.
if (!richClient(ctx)) { return 'ERROR: this client cannot draw diagrams. Explain it in prose instead — and do not draw one out of characters.'; }
const out = ctx.diagrams.render(input, { key: tu.id, model: ctx.model });
for (const m of out.post) { ctx.post(m); }
return out.result;
}
if (tu.name === diagramTool.GET_DIAGRAM.name) {
if (!richClient(ctx)) { return 'ERROR: this client cannot draw diagrams.'; }
ctx.post({ type: 'agentTool', icon: 'history', text: 'fetch diagram ' + String(input.id || '').slice(0, 24) });
return ctx.diagrams.fetch(input.id);
}
// MCP tools (docs/MCP.md S3). An MCP name matches none of the built-in branches above, so every
// MCP call necessarily arrives HERE — which is why the router is one block at one line rather
// than a dispatch scattered through runTool.
Expand Down Expand Up @@ -753,7 +778,13 @@ async function runAgent(ctx) {
// from the per-project journal). Rides the SAME cached-system channel as project rules — always-on but
// small — so a new session's first reply is continuous, not amnesiac. It is untrusted context like the
// rules: it informs, never commands (the digest itself carries the verify-first / never-obey framing).
const system = (ctx.skills ? buildSystem(ctx.skills.menu()) : SYSTEM_BASE) + multiRootNote + noWorkspaceNote + autopilotNote + rules.text
// Rich diagrams: `client.render` says what the surface on the other end can show. A rich client gets
// the render_diagram tool and the rules for using it; an ASCII client gets neither, so it is never
// told about a tool it does not have. The block goes straight after the base prompt — it is the
// same text for every run, so it belongs with the part of the prompt that never changes.
const rich = richClient(ctx);
const system = (ctx.skills ? buildSystem(ctx.skills.menu()) : SYSTEM_BASE) + (rich ? '\n\n' + diagramTool.PROMPT : '')
+ multiRootNote + noWorkspaceNote + autopilotNote + rules.text
+ (ctx.projectMemory ? '\n\n' + ctx.projectMemory : '');
const systemTokensEst = Math.round(system.length / 4);

Expand Down Expand Up @@ -792,14 +823,20 @@ async function runAgent(ctx) {
// Rootless runs get the portable subset; MCP tools are unaffected either way.
const builtins = root ? TOOLS : PORTABLE_TOOLS;
let tools = mcp.tools.length ? builtins.concat(mcp.tools) : builtins;
if (ctx.recallSessions) { tools = tools.concat([RECALL_TOOL]); } // cross-session recall (host-gated by memory settings)
const baseTools = ctx.recallSessions ? builtins.concat([RECALL_TOOL]) : builtins; // built-ins + recall; MCP is the rest
// Recomputed only when MCP or recall actually contributed tools, so the plain path keeps the module
// The host-gated extras: cross-session recall (memory settings), and the diagram tools (a rich
// client). get_diagram is offered only once a diagram's spec has left the conversation — until
// then there is nothing to fetch, and a tool that is never needed is a standing cost for nothing.
const extras = [];
if (ctx.recallSessions) { extras.push(RECALL_TOOL); }
if (rich) { extras.push(diagramTool.RENDER_DIAGRAM); if (ctx.diagramsStubbed) { extras.push(diagramTool.GET_DIAGRAM); } }
if (extras.length) { tools = tools.concat(extras); }
const baseTools = extras.length ? builtins.concat(extras) : builtins; // built-ins + extras; MCP is the rest
// Recomputed only when MCP or a host-gated extra actually contributed tools, so the plain path keeps the module
// constant and pays nothing for a feature it isn't using — but there are now TWO plain paths, and the
// constant has to match the list that was actually sent. Reporting the full cost for a rootless run
// was the same mistake as leaving baseTools on TOOLS, one line further down.
const builtinsTokensEst = root ? TOOLS_TOKENS_EST : PORTABLE_TOOLS_TOKENS_EST;
const toolsTokensEst = (mcp.tools.length || ctx.recallSessions) ? Math.round(JSON.stringify(tools).length / 4) : builtinsTokensEst;
const toolsTokensEst = (mcp.tools.length || extras.length) ? Math.round(JSON.stringify(tools).length / 4) : builtinsTokensEst;
// The MCP SHARE of that, reported separately so the context popover can show what these servers cost
// (docs/MCP.md S5). Every tool schema rides EVERY turn, so a chatty server is a standing tax on the
// window rather than a one-off — and until it has its own segment, that cost is invisible.
Expand All @@ -812,6 +849,9 @@ async function runAgent(ctx) {
: 0;

const messages = ctx.messages;
if (ctx.diagrams) { ctx.diagrams.beginRun(); } // nothing owed from an earlier run; repair passes reset
// tool_use ids whose placeholder already carries its title (the spec streams; the title arrives early)
const diagramTitled = new Set();
let step = 0;
let reason = 'done';
let nudges = 0;
Expand Down Expand Up @@ -898,12 +938,26 @@ async function runAgent(ctx) {
apiKey: ctx.apiKey, model: ctx.model, maxTokens: perTurnMax, system: system,
messages, tools: tools, signal: ctx.signal,
onText: (t) => { streamed = true; textChars += t.length; ctx.post({ type: 'agentDelta', text: t }); },
onToolStart: (name) => {
onToolStart: (name, id) => {
dbg('tool.start', { name });
const verb = name === 'edit_file' || name === 'write_file' ? 'preparing edit (' + name + ')…'
: name === 'delete_file' ? 'deleting a file…'
: name === 'run_command' ? 'preparing command…' : name === 'update_plan' ? 'planning…' : 'running ' + name + '…';
: name === 'run_command' ? 'preparing command…' : name === 'update_plan' ? 'planning…'
: name === diagramTool.RENDER_DIAGRAM.name ? 'drawing a diagram…' : 'running ' + name + '…';
ctx.post({ type: 'agentStatus', text: verb });
// A diagram's place in the answer is held from the moment the model starts writing it.
if (rich && id && name === diagramTool.RENDER_DIAGRAM.name) { ctx.post({ type: 'diagramPending', key: id, state: 'drawing', title: '' }); }
},
// The spec streams in as JSON. Its title is near the front, so the placeholder can say what
// is being drawn long before the last node arrives. Posted once per call.
onToolInput: (id, name, json) => {
if (!rich || !id || name !== diagramTool.RENDER_DIAGRAM.name || diagramTitled.has(id)) { return; }
const m = /"title"\s*:\s*"((?:[^"\\]|\\.)*)"/.exec(json);
if (!m) { return; }
diagramTitled.add(id);
let title = m[1];
try { title = JSON.parse('"' + m[1] + '"'); } catch (e) { /* show it as written */ }
ctx.post({ type: 'diagramPending', key: id, state: 'drawing', title: diagramRepair.cleanText(title).slice(0, 120) });
},
// A transient upstream 5xx (502/503/504) is retried once before it can fail the run — surface it
// as a status rather than a mystery pause, and log it. Nothing has streamed yet when this fires.
Expand Down Expand Up @@ -975,7 +1029,22 @@ async function runAgent(ctx) {
let cancelled = false;
for (const tu of toolUses) {
if (cancelled || ctx.signal.aborted) { cancelled = true; dbg('tool.cancelled', { name: tu.name }); results.push({ type: 'tool_result', tool_use_id: tu.id, content: 'Cancelled by the user.' }); continue; }
if (turn.malformed && turn.malformed.has(tu.id)) { dbg('tool.malformed', { name: tu.name }); results.push({ type: 'tool_result', tool_use_id: tu.id, content: 'ERROR: your tool arguments were cut off (truncated JSON). Retry with smaller input — for edits use edit_file with a short snippet.' }); continue; }
if (turn.malformed && turn.malformed.has(tu.id)) {
dbg('tool.malformed', { name: tu.name });
// A diagram spec is the one input worth a second look: JSON with a trailing comma or a
// comment is a spec, and the lenient parser reads it. A spec that simply STOPS is not —
// that is the model hitting its length limit, and it is re-requested, never repaired.
if (tu.name === diagramTool.RENDER_DIAGRAM.name && rich) {
const rawArgs = turn.raw && turn.raw.get(tu.id);
const parsed = (turn.stop_reason !== 'max_tokens' && typeof rawArgs === 'string') ? diagramRepair.parseLenient(rawArgs) : null;
const out = (parsed && parsed.ok) ? ctx.diagrams.render(rawArgs, { key: tu.id, model: ctx.model }) : ctx.diagrams.truncated({ key: tu.id, model: ctx.model });
for (const m of out.post) { ctx.post(m); }
Comment on lines +1037 to +1041
results.push({ type: 'tool_result', tool_use_id: tu.id, content: out.result });
continue;
}
results.push({ type: 'tool_result', tool_use_id: tu.id, content: 'ERROR: your tool arguments were cut off (truncated JSON). Retry with smaller input — for edits use edit_file with a short snippet.' });
continue;
}
dbg('tool.call', { name: tu.name, input: inputPreview(tu.input, !!(ctx.mcpRoutes && ctx.mcpRoutes.has(tu.name))) });
const out = await runTool(tu, ctx);
dbg('tool.result', { name: tu.name, chars: String(out).length, error: String(out).startsWith('ERROR') });
Expand All @@ -995,6 +1064,7 @@ async function runAgent(ctx) {

// No tool calls this turn.
const text = turn.content.filter((c) => c.type === 'text').map((c) => c.text).join('');
if (rich && text.trim()) { ctx.diagrams.noteAnswer(text); } // "ASCII leaks": counted, never acted on
if (turn.stop_reason === 'max_tokens') {
// The turn was pure prose cut off at the token cap. Ask the model to CONTINUE exactly where it
// left off (not switch strategies) so the full answer streams out across turns instead of being
Expand Down Expand Up @@ -1059,6 +1129,11 @@ async function runAgent(ctx) {
}
else { ctx.post({ type: 'agentError', message: msg, code }); reason = 'error'; }
} finally {
// A diagram sent back for repair that never came back right still owes the user a picture:
// draw what can be drawn of it now, before the run is declared over. Never a blank placeholder.
if (ctx.diagrams) {
try { for (const m of ctx.diagrams.endRun()) { ctx.post(m); } } catch (e) { dbg('diagram.settle.error', { msg: String((e && e.message) || e) }); }
}
dbg('agent.done', { reason, steps: step - 1, edits: ctx.editCount || 0, costMicros: runCostMicros, creditsLeftMicros: ctx.credits != null ? ctx.credits : null });
// [LevelCode] Gateway runs now carry real money: costMicros = what THIS run cost, credits = the
// remaining balance — both RETAIL micro-$. BYOK runs send neither (null/0) and the bar omits them.
Expand Down
Loading
Loading