Code Coliseum is the command center where autonomous AI engineers test, attack, repair and certify software before it reaches production.
A cinematic 3D verification arena, wired to real infrastructure.
React · TypeScript · React Three Fiber · Three.js · Drei · Framer Motion · Zustand · Vite
Codex / OpenAI · Greptile · Modal · AWS · Stripe
| 1. Architecture | How it works and where each layer plugs in |
| 2. What this is | The problem, and the shape of the answer |
| 3. Quickstart | Running in under a minute |
| 4. The experience | Every screen, in order |
| 5. The battle, beat by beat | What actually happens |
| 6. The three pull requests | Every run is data |
| 7. Wiring the integrations | Keys, endpoints, fallback behaviour |
| 8. The verification report | /report |
| 9. Module reference | Every file, what it does |
| 10. State and the battle director | The engine |
| 11. Inside the 3D scene | How each visual works |
| 12. Extending it | Add a PR, an agent, a district |
| 13. Console handles | Driving it live |
| 14. Demo and recording | Presenting it |
| 15. How this was verified | What was actually tested |
| 16. Specification traceability | Spec → implementation |
| 17. Troubleshooting | When something misbehaves |
A pull request does not get reviewed. It gets put through six stages, and every stage is powered by a real service.
┌──────────────────────────────────────────────────────┐
PULL REQUEST ─────▶│ COMMAND CENTER · triage · queue · human decisions │
└───────────────────────────┬──────────────────────────┘
│
╔════════════════════════════════════════════════▼═══════════════════════════════════════════════╗
║ VERIFICATION PIPELINE ║
╟──────────────┬──────────────┬──────────────┬──────────────┬──────────────┬─────────────────────╢
║ 1 · INDEX │ 2 · ATTACK │ 3 · REPAIR │ 4 · SIMULATE │ 5 · DECIDE │ 6 · CERTIFY ║
╟──────────────┼──────────────┼──────────────┼──────────────┼──────────────┼─────────────────────╢
║ GREPTILE │ GREPTILE │ CODEX │ MODAL │ CODEX │ MODAL + STRIPE ║
║ │ │ (OpenAI) │ │ (OpenAI) │ ║
║ indexes the │ variant │ reasons │ 10,000-user │ states the │ re-runs whatever ║
║ whole repo │ analysis │ over every │ load inside │ architecture│ failed, then the ║
║ before any │ across the │ finding, │ a real │ risk with a │ skill that ║
║ agent moves │ repo graph │ authors the │ isolated │ confidence │ certified the run ║
║ │ │ patches │ container │ score │ is billed ║
╟──────────────┴──────────────┴──────────────┴──────────────┴──────┬───────┴─────────────────────╢
║ escalates│to a HUMAN ║
╚═══════════════════════════════════════════════════════════════════▼═════════════════════════════╝
│
┌────────────────────────▼────────────────────────┐
│ AWS · the scale layer the verified build ships │
│ onto — region reachability and latency probed │
│ live at run time, no credentials required │
└────────────────────────┬────────────────────────┘
│
┌────────────────────────▼────────────────────────┐
│ REPORT · findings · decision record · replay │
└─────────────────────────────────────────────────┘
| Layer | Role | Where it lands in the run | Code |
|---|---|---|---|
| Greptile | Repository intelligence | Indexes the codebase before the fleet moves, then powers variant analysis. This is why the system finds the same defect in four other services — three of which the PR never touched. | greptileScan() |
| Codex / OpenAI | Reasoning + patch authoring | Reasons over every finding as a system, authors the fixes, and produces the recommendation and confidence score that get escalated to a human. | openaiAnalyze() |
| Modal | Execution sandbox | Every attack and the entire production simulation run in a real isolated container. Returns genuine per-stage p99 and error rates. | modalSimulate() · modal/simulate.py |
| AWS | Scale infrastructure | The deployment footprint the certified build ships onto. Region reachability and round-trip latency measured live — works with no key. | awsRegions() |
| Stripe | Skill marketplace | Verification capability is metered, not bundled. The skill that certified this run is pulled from the live catalogue. | stripeSkills() |
Full setup, endpoints and payloads in §7.
┌──────────────────────────────────────────┐
BROWSER │ React 19 · one persistent <Canvas> │
│ │
src/ui/ ─────┤ DOM overlays (Framer Motion) │
│ Hook · Boot · CommandCenter · HUD │
│ Simulation · DecisionGate · Debrief │
│ │
src/arena/ ─────┤ R3F scene graph (Three.js) │
│ Arena · City · CodeCore · Ships │
│ SponsorInfra · Effects · CameraRig │
│ │
src/state/ ─────┤ Zustand store — single source of truth │
│ battle.ts drives it as a timeline │
└──────────────────┬───────────────────────┘
│ fetch · hard timeout · guaranteed fallback
┌──────────────────▼───────────────────────┐
VITE SERVER │ server/sponsors.ts holds the API keys │
│ server/report.ts serves /report │
└──────────────────┬───────────────────────┘
│
Greptile · Codex/OpenAI · Modal · AWS · Stripe
Keys never reach the browser. Every integration call is made server-side from server/sponsors.ts.
1 · One store, everything derives from it. battle.ts is the only thing that writes the timeline. Every component — a camera rig, a building's crack opacity, a HUD bar, an overlay — reads state and reacts. No component tells another what to do, so a single phase change moves the camera, the ships, the city and six overlays with no coordination code.
2 · The battle is a script, not a state machine. runBattle() is a plain async function of awaits and put() calls. You can read the entire experience top to bottom like prose, and retime any beat by changing one number. A cancellation token makes every restart clean.
3 · The canvas never unmounts. "Views" are overlays plus a camera position, which is why the world stays alive underneath every screen and transitions feel continuous rather than like page loads.
Every team is now shipping AI-written code faster than any human can review it. The bottleneck moved. It is no longer writing software — it is proving it.
Code Coliseum is what that proof looks like when you take it seriously and give it a face.
A pull request arrives in a holographic mission queue. You send it into a cyberpunk coliseum where a fleet of AI ships hunts it for bugs, vulnerabilities and bottlenecks. A reasoning agent rebuilds the broken code. The repaired build is then put under synthetic production load in a real sandbox — and when it still fails, the system escalates exactly one question to a human, with a recommendation and a confidence score. You approve it. The agent applies the change, re-runs what failed, and certifies the branch.
Then the developer gets a report: every problem, where it is, the fix, what the AI recommended, what a human chose, and a replayable timeline of the entire run.
The thesis, in one line: AI handles the execution. Humans decide the handful of things that actually need judgement.
It is not a code review dashboard. The 3D world is the interface; the intelligence layer is the product.
| Runtime | ~2 minutes 45 seconds, end to end, unattended or driven |
| Views | Hook → Boot → Title → Command Center → Arena → Debrief → Final Vision |
| AI agents | 4 ships, each with a distinct role, hull and reputation |
| Districts | 5 services with health, security, coverage, risk and dependencies |
| Missions | 3 fully-specified pull requests, each with its own battle |
| Infrastructure | 5 layers, real API calls, per-layer LIVE/SIM badging |
| Code | ~4,700 lines TS/TSX · 828 lines CSS · 255-line integration server · 72-line Modal function · 1,206-line report |
| Assets | Zero. Every mesh, texture, particle and logo mark is procedural |
git clone https://github.com/ayushozha/code-arena.git
cd code-arena
npm install
npm run devOpen http://localhost:5173, go fullscreen, and let the opening play.
| ⚔ ENTER COMMAND CENTER | Drive it yourself. Launch any of the three pull requests. |
| ▶ AUTO DEMO · 2:30 | It runs itself, auto-approves the decision gate after 22s, and lands on the closing card. |
The verification report lives at http://localhost:5173/report.
cp .env.example .env # fill in whatever you have, leave the rest blankEvery key you add flips that layer from SIM to LIVE on screen. Anything blank runs its deterministic profile instead. AWS is live with no key at all.
| Command | Does |
|---|---|
npm run dev |
Dev server on :5173, including the integration API and /report |
npm run build |
Typecheck (tsc -b) then production build to dist/ |
npm run preview |
Serve the production build, integrations and /report included |
npm run lint |
oxlint |
Requirements: Node 20+ (developed on 26), a WebGL2-capable browser, and a GPU you would not be embarrassed to run a game on.
Note on naming. The product is Code Coliseum. Internal identifiers still use
arena— the package name, thewindow.arenaconsole handle,ARENA-…run ids and the on-screen wordmark. Those are left alone deliberately so nothing breaks; the README uses the product name.
Seven views. The 3D canvas is mounted once and never unmounts — every "screen" is an overlay over the same living world, and the camera flies between them.
A terminal writes a whole payment system in seconds — API, database, tests, DONE. Then the screen turns: BUT… Who verified it? That is the entire pitch, delivered in twelve seconds before a single feature is shown. SKIP ▸ jumps to boot.
Systems come online one at a time — arena, repository world, AI agents, then each infrastructure layer by name. Live integration status is fetched here, so the LIVE/SIM badges are truthful by the time you reach the command center.
The glitching wordmark, the four agents, the five infrastructure layers, and the two ways in.
A tactical overhead of the whole arena with the management layer on top:
- Active missions — three pull requests, risk-scored, each launchable
- AI triage — which PR goes first, why, and which agents are being deployed
- Infrastructure — all five layers with live status and LIVE/SIM badge
- AI fleet · reputation — lifetime stats per agent
- Engineering leaderboard — developers by bugs prevented and quality score, plus AI agent rankings
Click any building for its intelligence panel.
The battle — see §5.
The developer-facing report: score strip, problems found and fixed with locations and suggested fixes, production simulation results, the decision record, rankings, and a replay of every beat — timestamped, colour-coded and camera-linked.
The camera pulls away over the running city:
AI doesn't just write code. AI proves code deserves to exist.
Driven entirely by src/state/battle.ts. Phases advance intro → spawn → attack → repair → simulate → decision → verify → score → victory, and the camera, HUD, ships and city all react to the phase.
Opening — the PR enters, the code core materialises at 100% integrity, the repository city rises: five districts, six dependency links.
Repository scan · Greptile — the scanning fleet sweeps and indexes the city district by district before any agent moves. This is why variants get found — the agents attack with a map of the codebase, not just the diff.
Sandbox provisioning · Modal — eight pods light up around the arena. Everything from here runs in isolation.
Fleet launch — four ships fly in from outside, bank into their turns, take station.
Attack phase — each attack is: agent closes on its district → optional infrastructure beat → strike. On impact the building shakes, cracks and blows instanced debris across the platform, the camera shakes, a damage card slides in, health drops.
The Security Hunter is the important one: it finds a vulnerability, then runs variant analysis — the dependency graph turns red and the same defect is struck in three more districts. Finding the bug is table stakes; finding every other place that bug lives is the product.
Repair phase · Codex — the reasoning core ignites and analyses all findings as a system. Patches are authored district by district; each one pulls the debris back in and reseals the structure.
Production simulation · Modal — 10,000 users, high traffic, database stress, latency injection, with real p99 and error-rate readouts. For PR #4921 three stages pass and Payment Retry fails. Every patch was correct. Every test passed. It still falls over — because the problem is not in the diff, it is in the architecture the diff was written into.
Human decision — the arena stops. A gate states the recommendation and confidence, shows both options with their risk, and waits for ✓ APPROVE or ⟲ REQUEST CHANGES. Auto-approves after 22 seconds so unattended loops complete. Approve, and the agent applies the change, re-runs the failed stage, and it passes.
Verification — the branch arrives carrying a claim, "feature complete", and the verification agent does not take it on trust. Four proofs: Tests, Security, Performance, Regression.
Score, victory, debrief — scoreboard, core transformation, fireworks, camera pull-away, report.
Every battle is data. src/state/prs.ts holds three complete missions. battle.ts runs whichever you launch.
| #4921 | #4922 | #4923 | |
|---|---|---|---|
| Title | Payment Refund System | Authentication Update | Database Migration |
| Risk | HIGH | MEDIUM | CRITICAL |
| Diff | 34 files · +1,208 −417 | 18 files · +604 −233 | 52 files · +2,416 −1,988 |
| Attacks | 4 | 3 | 4 |
| Health floor | 72% | 79% | 41% |
| Variants | 5 | 2 | 4 |
| Simulation | 1 failure — Payment Retry | all pass | 2 failures |
| Decision | Architecture change · 94% | Rollout strategy · 89% | Deployment window · 97% |
| Final score | 96 / 94 / 91 / 100 | 98 / 91 / 95 / 100 | 93 / 89 / 86 / 100 |
#4921 is the demo run. #4922 shows the system only escalates when it needs to — its simulation passes clean and no human decision is requested. #4923 is the worst case.
Each layer calls its real API when credentials are present, badges itself LIVE or SIM on screen, and falls back to a deterministic profile when it can't.
POST https://api.openai.com/v1/chat/completions
{ model, response_format: { type: "json_object" }, messages: [ findings + service context ] }
→ { headline, confidence, recommendation, reason, verdict }
OPENAI_API_KEY=sk-...
OPENAI_MODEL=gpt-4o-mini # override with whatever you have access toPOST https://api.greptile.com/v2/query
Authorization: Bearer $GREPTILE_API_KEY X-GitHub-Token: $GITHUB_TOKEN
{ messages, repositories: [{ remote: "github", repository, branch }], genius: true }
→ { message, sources: [{ filepath, linestart }] }
GREPTILE_API_KEY=...
GITHUB_TOKEN=ghp_...
GREPTILE_REPOSITORY=owner/repo
GREPTILE_BRANCH=mainpip install modal && modal setup
modal deploy modal/simulate.py # prints your web URLMODAL_SIM_ENDPOINT=https://<workspace>--code-arena-sim-run.modal.run
MODAL_SIM_TOKEN= # optional bearerReturned stages override the scripted profile per stage name, so a live container can change the outcome of the run.
GET https://dynamodb.{region}.amazonaws.com/ → { regions: [{ region, ms, up }], active, p50 }
An unauthenticated request returns a 4xx, which still proves the endpoint is reachable and yields a genuine round-trip time.
AWS_REGIONS=us-east-1,eu-west-1,ap-south-1GET https://api.stripe.com/v1/prices?limit=6&active=true&expand[]=data.product
STRIPE_API_KEY=rk_... # a restricted read-only key is enoughEvery layer is independently degradable. A missing key, an expired token, a rate limit or dead wifi drops that one layer to its deterministic profile, marks it SIM on screen, and the pipeline continues to a verdict. Server timeout 12s, client 14s.
There is no configuration in which the pipeline hangs or hard-fails on an external service. That is a design constraint, not a happy accident — a demo that depends on venue wifi is not a demo.
curl -s localhost:5173/api/status | jq # see exactly what's live| Route | Returns |
|---|---|
GET /api/status |
Per-layer mode and detail |
POST /api/openai/analyze |
Reasoning verdict |
POST /api/greptile/scan |
Repository query result |
POST /api/modal/simulate |
Load simulation stages |
GET /api/aws/regions |
Live region latencies |
GET /api/stripe/skills |
Skill catalogue |
http://localhost:5173/report — a standalone engineering report for the run. Self-contained: inline styles, inline SVG charts, no build step, no runtime dependencies.
| Section | Contains |
|---|---|
| The diff | 34 files, per-file churn bars, language mix, blast radius |
| What actually happened | SVG code-health trace (100 → 72 → 100) and all 21 beats |
| What the agents changed | Per-agent stats, then six patches with before/after diffs |
| Security | Nine findings with file:line, CWE, CVSS — plus an SVG variant-spread map |
| Production simulation | Stage results, first run vs. re-run |
| What the human approved | The escalation, both options, and a full audit row |
| Infrastructure | Every layer's role, results, and exact endpoint + payload |
| Verification | Four proofs and what each covered |
| Impact | What didn't reach production, engineering record, agent precision |
| Appendix | Run metadata, district states, reproducibility, glossary |
Served by server/report.ts, because Vite's SPA fallback would otherwise hand the arena to any extensionless path.
| File | Lines | Responsibility |
|---|---|---|
Scene.tsx |
57 | Canvas, lighting, tone mapping, postprocessing (bloom, chromatic aberration, noise, vignette) |
Arena.tsx |
194 | Coliseum: platform, grid, rim light, 28 pillars, energy rings, 34 data streams, 7 code panels, dome |
City.tsx |
300 | Five districts: tiered towers, health labels, crack overlay, debris physics, click-through inspection, dependency arcs |
CodeCore.tsx |
164 | The repository crystal — dormant / healthy / attacked / repairing / verified |
Ships.tsx |
313 | Four AI ships with flight physics — fly in, bank, orbit, strike, return — plus attack beams |
SponsorInfra.tsx |
258 | Codex reasoning core, Greptile scan wave, Modal pods, AWS towers, Stripe pylons |
Effects.tsx |
145 | Impact bursts and victory fireworks |
CameraRig.tsx |
190 | Every shot in the film — per-view and per-phase framing, damping, shake, replay focus |
config.ts |
183 | Palette, agents, layers, districts, dependency graph, building intel |
textures.ts |
90 | Canvas-generated grid, code and window textures |
| File | Lines | Responsibility |
|---|---|---|
battle.ts |
533 | The director: boot, battle, simulation, decision gate, replay, auto-demo |
store.ts |
306 | Zustand store, all state and mutators |
prs.ts |
331 | The three pull requests, fully specified |
live.ts |
66 | API client with timeout and fallback |
| File | Lines | Responsibility |
|---|---|---|
Overlays.tsx |
203 | Title card, phase banners, infrastructure beats, attack cards, verification, scoreboard, victory |
HUD.tsx |
177 | PR card, infrastructure telemetry rack, fleet roster, arena feed, phase timeline, demo clock |
Debrief.tsx |
178 | The developer report |
CommandCenter.tsx |
162 | Mission queue, AI triage, infrastructure, fleet reputation, leaderboards |
Inspector.tsx |
86 | Per-district panel — changes, vulnerabilities, recommendations, dependencies |
DecisionGate.tsx |
79 | The human decision gate |
Simulation.tsx |
72 | Production simulation panel |
Hook.tsx |
66 | The opening terminal |
Boot.tsx |
60 | Boot sequence |
Logos.tsx |
60 | Infrastructure marks, drawn as SVG |
Replay.tsx |
38 | Replay timeline |
Finale.tsx |
28 | Closing card |
| File | Purpose |
|---|---|
server/sponsors.ts |
Live integration layer — holds the keys, one handler per layer |
server/report.ts |
Serves /report |
modal/simulate.py |
Deployable Modal load-simulation endpoint |
public/report/index.html |
The verification report |
src/index.css |
828 lines — the entire neon design system |
DEMO.md |
Click-by-click demo runbook with recovery steps |
script.md |
Word-for-word video narration script |
view: 'hook' | 'boot' | 'intro' | 'command' | 'arena' | 'debrief' | 'finale'
phase: 'idle' | 'intro' | 'spawn' | 'attack' | 'repair'
| 'simulate' | 'decision' | 'verify' | 'score' | 'victory'
core: 'dormant' | 'healthy' | 'attacked' | 'repairing' | 'verified'
agent: 'offline' | 'inbound' | 'standby' | 'scanning'
| 'attacking' | 'repairing' | 'returning' | 'done'Plus health and sub-stats, per-district state, infrastructure status, simulation results, the decision gate, the replay log and the leaderboard.
await sponsorBeat('greptile', [...], sleep) // card in, log line, card out
s().hit(target, colour) // impact burst + camera shake
s().damageBuilding(target, amount, finding) // health down, cracks, debris, marker
put({ health: to, core: 'attacked', attackCard: {...} })
mark('Bug Hunter · double refund', 'hit', target) // recorded for the replay
await sleep(3200)Retiming the demo is editing sleep() values. Adding a beat is adding four lines.
A module-level token guards every sleep(). stopAll() increments it, so any in-flight timeline rejects at its next await and unwinds through a single catch. Restarting mid-run is always clean.
const choice = await awaitGate(my) // polls gateChoice; auto-approves at 22sThe timeline genuinely stops. This is the one place the system yields control, and it is the point of the whole product.
Everything is procedural — no GLTF, no CDN, nothing to fail to load mid-demo.
Damage physics. Each district owns an InstancedMesh of 16 chunks. On a hit they launch outward on a ballistic arc with gravity and spin, and settle. On repair they interpolate back to origin and scale to nothing — particles running in reverse, which reads instantly as "rebuilding". The tower squashes on impact and blooms on repair; the crack wireframe's opacity is driven directly by health.
Ship flight. Ships lerp toward a destination that depends on status. Heading comes from the velocity vector and bank angle from velocity magnitude — so they lean into turns for free.
Camera. One rig owns every shot. Attack framing orbits on the target's bearing, between the city ring and the pillar ring, so the target is foreground and the core sits framed behind it — and the shot can never clip through a building. Damping is frame-rate independent (1 - 0.0012^dt) and impacts inject decaying shake.
Bloom and readability. Emissive materials use toneMapped={false} with high emissive intensity so they punch through the bloom threshold, while the platform stays deliberately dark. That is the difference between "neon" and "washed out" — the palette glows, not the lighting.
Append a PRDef to PRS in src/state/prs.ts. It automatically gets a queue row, triage reasoning, its own battle, simulation, decision and report.
{
key: 'pr4924', num: 4924, title: 'Search Reindex', service: 'SEARCH SERVICE',
branch: 'feat/reindex', risk: 'MEDIUM', files: 12, added: 380, removed: 90, author: 'you',
scan: { understanding: 97, related: 2, files: 1842 },
attacks: [ /* agent, target district, damage, stat hit, optional variants + beat */ ],
fixes: [ ['database', 'reindex batched'] ],
problems: [ /* location + suggested fix — these become the report */ ],
triage: { rank: 3, why: '…', deploy: ['bug', 'perf'] },
simulation: { users: 10000, /* … */ stages: [ /* passes: false triggers the human gate */ ] },
decision: { headline, recommendation, confidence, options, recommended, reason, onApprove, onChanges },
scores: { correctness: 97, security: 96, performance: 93, verification: 100 },
bugs: 5, points: 300,
}A stage with passes: false is what summons the human decision gate. That is the only switch.
Append to BUILDINGS in config.ts, add an entry to BUILDING_INTEL, wire it into DEPENDENCIES. The city, links, scan sweep, inspector and camera all pick it up.
Add to AGENTS in config.ts with a colour, spawn angle, powering layer and reputation block, build a hull component in Ships.tsx, register it in the Ships() group.
Every sleep(n) in battle.ts is milliseconds. The auto-demo's pre-roll is the wait(9000) in runDemo().
In dev, window.arena is exposed for driving the thing live — useful mid-presentation.
arena.demo() // run the 2:30 auto demo
arena.command() // jump to the command center
arena.battle('pr4923') // launch a specific mission
arena.boot() // replay the boot sequence
arena.store.getState() // inspect everything
arena.store.getState().set({ view: 'finale' }) // jump to the closing card
arena.store.getState().set({ inspecting: 'payments' }) // open a district panelDEMO.md— the presenter's runbook. Beat-by-beat, what to say and when to click, live-vs-simulated setup, and a recovery table for when something wedges on stage.script.md— the word-for-word video narration, timed to the arena's own pacing, with the three pauses that carry it and a 60-second cut.
| Time | Beat |
|---|---|
| 0:00 – 0:30 | The problem — terminal, then who verified it? |
| 0:30 – 0:50 | Activation and the command center |
| 0:50 – 1:40 | The battle — scan, fleet, attacks, the variant chain |
| 1:40 – 2:00 | Repair, then the production simulation failure |
| 2:00 – 2:20 | The human decision, verification, the report |
| 2:20 – 2:30 | Final vision |
If you're recording: wire Codex/OpenAI and Modal first. Two minutes of setup, and it is the difference between "nice visualisation" and "this is real."
Every stage of this build was driven in a real headless Chrome — boot, command center, full battle, decision gate clicked, replay interaction, report rendering — with console and page errors captured. Final state: 0 errors, simulation passed, 20 replay events recorded, 620 points awarded.
Bugs that pass caught, none of which were visible from reading the code:
| Bug | Why it mattered |
|---|---|
| Attack camera swept through buildings | The hero shot of the whole demo was a close-up of a roof |
Framer Motion writes an inline transform |
Silently killed CSS transform-centering on the banner, attack card, sponsor card, verification panel, scoreboard and command-center header |
.sim class collision |
The simulation panel's class also matched em.sim, flinging LIVE/SIM badges across the page |
| Ship name tags ballooning | Drei Html scales inversely with distance; a ship near the camera produced full-screen ghost text |
| Arena feed overflowing its panel | Log lines escaped the container |
| Code panels washing out the attack shot | Additive planes sat between camera and city |
| AWS labels dominating the closing shot | The finale camera flies past the towers |
| Roof caps wider than the tiers below | Buildings read as flat slabs |
Typecheck (tsc -b) and production build are clean.
Five specification documents live alongside the code. Everything in the first four is implemented.
| Spec | Status |
|---|---|
Code_Arena_3D_Arena_Specification.md |
✅ Arena, code core, repository city, agents, attack/repair, health, verification, victory |
Code_Arena_PRD_v2_Arena_Evolution_Upgrade.md |
✅ Command center, PR queue, infrastructure visualisation, AI ships, physics, feedback, decisions, rankings, loading |
Code_Arena_Enhancement_Addendum_Specification.md |
✅ Repository intelligence overlays + click-through, agent reputation, production simulation, human decision layer, battle replay |
Code_Arena_2_5_Minute_Demo_Flow_Specification.md |
✅ Problem hook, activation, mission arrival, battle, repair + simulation, command center reveal, final vision |
Code_Arena_Hype_Demo_Flow_Final.md |
⬜ Not yet implemented |
Beyond the specs: real API integration with LIVE/SIM badging and guaranteed fallback, the /report verification report, and the demo and script documents.
| Symptom | Cause and fix |
|---|---|
/report shows the arena |
The dev server needs a restart after a vite.config.ts change — the route is a plugin middleware. |
| All layers say SIM | Expected with no .env. curl localhost:5173/api/status shows which key each layer is missing. AWS should still say LIVE. |
| A layer says SIM despite a key | The reason is in the API response — bad key, wrong repo slug, or an undeployed Modal endpoint. Hit the route directly to see it. |
| OpenAI returns an error | OPENAI_MODEL is probably a model your key can't reach. Set one you have access to; the run falls back safely either way. |
| Low frame rate | Close other GPU-heavy tabs. Bloom and postprocessing are the expensive part; dpr is capped at 1.8 and drei's AdaptiveDpr degrades under load. |
| Battle seems stuck | It's the decision gate waiting for you. It auto-approves after 22s, or click ✓ APPROVE. |
| Demo wedged mid-run | arena.command() in the console, or ⌂ COMMAND CENTER at the bottom of the arena. |
| Fonts look wrong | Orbitron and JetBrains Mono come from Google Fonts. Offline, the monospace fallback is used and everything still works. |