Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🏛 CODE COLISEUM

AI can write code. But who verifies it?

Code Coliseum is the command center where autonomous AI engineers test, attack, repair and certify software before it reaches production.

A cinematic 3D verification arena, wired to real infrastructure.

React · TypeScript · React Three Fiber · Three.js · Drei · Framer Motion · Zustand · Vite

Codex / OpenAI · Greptile · Modal · AWS · Stripe


Contents

1. Architecture How it works and where each layer plugs in
2. What this is The problem, and the shape of the answer
3. Quickstart Running in under a minute
4. The experience Every screen, in order
5. The battle, beat by beat What actually happens
6. The three pull requests Every run is data
7. Wiring the integrations Keys, endpoints, fallback behaviour
8. The verification report /report
9. Module reference Every file, what it does
10. State and the battle director The engine
11. Inside the 3D scene How each visual works
12. Extending it Add a PR, an agent, a district
13. Console handles Driving it live
14. Demo and recording Presenting it
15. How this was verified What was actually tested
16. Specification traceability Spec → implementation
17. Troubleshooting When something misbehaves

1. Architecture

The verification pipeline

A pull request does not get reviewed. It gets put through six stages, and every stage is powered by a real service.

                      ┌──────────────────────────────────────────────────────┐
   PULL REQUEST ─────▶│   COMMAND CENTER · triage · queue · human decisions   │
                      └───────────────────────────┬──────────────────────────┘
                                                  │
 ╔════════════════════════════════════════════════▼═══════════════════════════════════════════════╗
 ║                                  VERIFICATION PIPELINE                                          ║
 ╟──────────────┬──────────────┬──────────────┬──────────────┬──────────────┬─────────────────────╢
 ║  1 · INDEX   │  2 · ATTACK  │  3 · REPAIR  │ 4 · SIMULATE │  5 · DECIDE  │    6 · CERTIFY      ║
 ╟──────────────┼──────────────┼──────────────┼──────────────┼──────────────┼─────────────────────╢
 ║  GREPTILE    │  GREPTILE    │    CODEX     │    MODAL     │    CODEX     │  MODAL + STRIPE     ║
 ║              │              │   (OpenAI)   │              │   (OpenAI)   │                     ║
 ║  indexes the │  variant     │  reasons     │  10,000-user │  states the  │  re-runs whatever   ║
 ║  whole repo  │  analysis    │  over every  │  load inside │  architecture│  failed, then the   ║
 ║  before any  │  across the  │  finding,    │  a real      │  risk with a │  skill that         ║
 ║  agent moves │  repo graph  │  authors the │  isolated    │  confidence  │  certified the run  ║
 ║              │              │  patches     │  container   │  score       │  is billed          ║
 ╟──────────────┴──────────────┴──────────────┴──────────────┴──────┬───────┴─────────────────────╢
 ║                                                          escalates│to a HUMAN                   ║
 ╚═══════════════════════════════════════════════════════════════════▼═════════════════════════════╝
                                                  │
                         ┌────────────────────────▼────────────────────────┐
                         │  AWS · the scale layer the verified build ships │
                         │  onto — region reachability and latency probed  │
                         │  live at run time, no credentials required      │
                         └────────────────────────┬────────────────────────┘
                                                  │
                         ┌────────────────────────▼────────────────────────┐
                         │  REPORT · findings · decision record · replay   │
                         └─────────────────────────────────────────────────┘

Who does what

Layer Role Where it lands in the run Code
Greptile Repository intelligence Indexes the codebase before the fleet moves, then powers variant analysis. This is why the system finds the same defect in four other services — three of which the PR never touched. greptileScan()
Codex / OpenAI Reasoning + patch authoring Reasons over every finding as a system, authors the fixes, and produces the recommendation and confidence score that get escalated to a human. openaiAnalyze()
Modal Execution sandbox Every attack and the entire production simulation run in a real isolated container. Returns genuine per-stage p99 and error rates. modalSimulate() · modal/simulate.py
AWS Scale infrastructure The deployment footprint the certified build ships onto. Region reachability and round-trip latency measured live — works with no key. awsRegions()
Stripe Skill marketplace Verification capability is metered, not bundled. The skill that certified this run is pulled from the live catalogue. stripeSkills()

Full setup, endpoints and payloads in §7.

Runtime architecture

                    ┌──────────────────────────────────────────┐
   BROWSER          │  React 19 · one persistent <Canvas>      │
                    │                                          │
   src/ui/     ─────┤  DOM overlays (Framer Motion)            │
                    │    Hook · Boot · CommandCenter · HUD      │
                    │    Simulation · DecisionGate · Debrief    │
                    │                                          │
   src/arena/  ─────┤  R3F scene graph (Three.js)              │
                    │    Arena · City · CodeCore · Ships        │
                    │    SponsorInfra · Effects · CameraRig     │
                    │                                          │
   src/state/  ─────┤  Zustand store — single source of truth  │
                    │    battle.ts drives it as a timeline      │
                    └──────────────────┬───────────────────────┘
                                       │  fetch · hard timeout · guaranteed fallback
                    ┌──────────────────▼───────────────────────┐
   VITE SERVER      │  server/sponsors.ts   holds the API keys │
                    │  server/report.ts     serves /report     │
                    └──────────────────┬───────────────────────┘
                                       │
              Greptile · Codex/OpenAI · Modal · AWS · Stripe

Keys never reach the browser. Every integration call is made server-side from server/sponsors.ts.

The three ideas that hold it together

1 · One store, everything derives from it. battle.ts is the only thing that writes the timeline. Every component — a camera rig, a building's crack opacity, a HUD bar, an overlay — reads state and reacts. No component tells another what to do, so a single phase change moves the camera, the ships, the city and six overlays with no coordination code.

2 · The battle is a script, not a state machine. runBattle() is a plain async function of awaits and put() calls. You can read the entire experience top to bottom like prose, and retime any beat by changing one number. A cancellation token makes every restart clean.

3 · The canvas never unmounts. "Views" are overlays plus a camera position, which is why the world stays alive underneath every screen and transitions feel continuous rather than like page loads.


2. What this is

Every team is now shipping AI-written code faster than any human can review it. The bottleneck moved. It is no longer writing software — it is proving it.

Code Coliseum is what that proof looks like when you take it seriously and give it a face.

A pull request arrives in a holographic mission queue. You send it into a cyberpunk coliseum where a fleet of AI ships hunts it for bugs, vulnerabilities and bottlenecks. A reasoning agent rebuilds the broken code. The repaired build is then put under synthetic production load in a real sandbox — and when it still fails, the system escalates exactly one question to a human, with a recommendation and a confidence score. You approve it. The agent applies the change, re-runs what failed, and certifies the branch.

Then the developer gets a report: every problem, where it is, the fix, what the AI recommended, what a human chose, and a replayable timeline of the entire run.

The thesis, in one line: AI handles the execution. Humans decide the handful of things that actually need judgement.

It is not a code review dashboard. The 3D world is the interface; the intelligence layer is the product.

Runtime ~2 minutes 45 seconds, end to end, unattended or driven
Views Hook → Boot → Title → Command Center → Arena → Debrief → Final Vision
AI agents 4 ships, each with a distinct role, hull and reputation
Districts 5 services with health, security, coverage, risk and dependencies
Missions 3 fully-specified pull requests, each with its own battle
Infrastructure 5 layers, real API calls, per-layer LIVE/SIM badging
Code ~4,700 lines TS/TSX · 828 lines CSS · 255-line integration server · 72-line Modal function · 1,206-line report
Assets Zero. Every mesh, texture, particle and logo mark is procedural

3. Quickstart

git clone https://github.com/ayushozha/code-arena.git
cd code-arena
npm install
npm run dev

Open http://localhost:5173, go fullscreen, and let the opening play.

⚔ ENTER COMMAND CENTER Drive it yourself. Launch any of the three pull requests.
▶ AUTO DEMO · 2:30 It runs itself, auto-approves the decision gate after 22s, and lands on the closing card.

The verification report lives at http://localhost:5173/report.

Optional: wire the live integrations

cp .env.example .env      # fill in whatever you have, leave the rest blank

Every key you add flips that layer from SIM to LIVE on screen. Anything blank runs its deterministic profile instead. AWS is live with no key at all.

Scripts

Command Does
npm run dev Dev server on :5173, including the integration API and /report
npm run build Typecheck (tsc -b) then production build to dist/
npm run preview Serve the production build, integrations and /report included
npm run lint oxlint

Requirements: Node 20+ (developed on 26), a WebGL2-capable browser, and a GPU you would not be embarrassed to run a game on.

Note on naming. The product is Code Coliseum. Internal identifiers still use arena — the package name, the window.arena console handle, ARENA-… run ids and the on-screen wordmark. Those are left alone deliberately so nothing breaks; the README uses the product name.


4. The experience

Seven views. The 3D canvas is mounted once and never unmounts — every "screen" is an overlay over the same living world, and the camera flies between them.

Hook · src/ui/Hook.tsx

A terminal writes a whole payment system in seconds — API, database, tests, DONE. Then the screen turns: BUT… Who verified it? That is the entire pitch, delivered in twelve seconds before a single feature is shown. SKIP ▸ jumps to boot.

Boot · src/ui/Boot.tsx

Systems come online one at a time — arena, repository world, AI agents, then each infrastructure layer by name. Live integration status is fetched here, so the LIVE/SIM badges are truthful by the time you reach the command center.

Title card · src/ui/Overlays.tsx

The glitching wordmark, the four agents, the five infrastructure layers, and the two ways in.

Command Center · src/ui/CommandCenter.tsx

A tactical overhead of the whole arena with the management layer on top:

  • Active missions — three pull requests, risk-scored, each launchable
  • AI triage — which PR goes first, why, and which agents are being deployed
  • Infrastructure — all five layers with live status and LIVE/SIM badge
  • AI fleet · reputation — lifetime stats per agent
  • Engineering leaderboard — developers by bugs prevented and quality score, plus AI agent rankings

Click any building for its intelligence panel.

Arena

The battle — see §5.

Debrief · src/ui/Debrief.tsx

The developer-facing report: score strip, problems found and fixed with locations and suggested fixes, production simulation results, the decision record, rankings, and a replay of every beat — timestamped, colour-coded and camera-linked.

Final Vision · src/ui/Finale.tsx

The camera pulls away over the running city:

AI doesn't just write code. AI proves code deserves to exist.


5. The battle, beat by beat

Driven entirely by src/state/battle.ts. Phases advance intro → spawn → attack → repair → simulate → decision → verify → score → victory, and the camera, HUD, ships and city all react to the phase.

Opening — the PR enters, the code core materialises at 100% integrity, the repository city rises: five districts, six dependency links.

Repository scan · Greptile — the scanning fleet sweeps and indexes the city district by district before any agent moves. This is why variants get found — the agents attack with a map of the codebase, not just the diff.

Sandbox provisioning · Modal — eight pods light up around the arena. Everything from here runs in isolation.

Fleet launch — four ships fly in from outside, bank into their turns, take station.

Attack phase — each attack is: agent closes on its district → optional infrastructure beat → strike. On impact the building shakes, cracks and blows instanced debris across the platform, the camera shakes, a damage card slides in, health drops.

The Security Hunter is the important one: it finds a vulnerability, then runs variant analysis — the dependency graph turns red and the same defect is struck in three more districts. Finding the bug is table stakes; finding every other place that bug lives is the product.

Repair phase · Codex — the reasoning core ignites and analyses all findings as a system. Patches are authored district by district; each one pulls the debris back in and reseals the structure.

Production simulation · Modal — 10,000 users, high traffic, database stress, latency injection, with real p99 and error-rate readouts. For PR #4921 three stages pass and Payment Retry fails. Every patch was correct. Every test passed. It still falls over — because the problem is not in the diff, it is in the architecture the diff was written into.

Human decision — the arena stops. A gate states the recommendation and confidence, shows both options with their risk, and waits for ✓ APPROVE or ⟲ REQUEST CHANGES. Auto-approves after 22 seconds so unattended loops complete. Approve, and the agent applies the change, re-runs the failed stage, and it passes.

Verification — the branch arrives carrying a claim, "feature complete", and the verification agent does not take it on trust. Four proofs: Tests, Security, Performance, Regression.

Score, victory, debrief — scoreboard, core transformation, fireworks, camera pull-away, report.


6. The three pull requests

Every battle is data. src/state/prs.ts holds three complete missions. battle.ts runs whichever you launch.

#4921 #4922 #4923
Title Payment Refund System Authentication Update Database Migration
Risk HIGH MEDIUM CRITICAL
Diff 34 files · +1,208 −417 18 files · +604 −233 52 files · +2,416 −1,988
Attacks 4 3 4
Health floor 72% 79% 41%
Variants 5 2 4
Simulation 1 failure — Payment Retry all pass 2 failures
Decision Architecture change · 94% Rollout strategy · 89% Deployment window · 97%
Final score 96 / 94 / 91 / 100 98 / 91 / 95 / 100 93 / 89 / 86 / 100

#4921 is the demo run. #4922 shows the system only escalates when it needs to — its simulation passes clean and no human decision is requested. #4923 is the worst case.


7. Wiring the integrations

Each layer calls its real API when credentials are present, badges itself LIVE or SIM on screen, and falls back to a deterministic profile when it can't.

Codex / OpenAI · reasoning and patch authoring

POST https://api.openai.com/v1/chat/completions
{ model, response_format: { type: "json_object" }, messages: [ findings + service context ] }
→ { headline, confidence, recommendation, reason, verdict }
OPENAI_API_KEY=sk-...
OPENAI_MODEL=gpt-4o-mini      # override with whatever you have access to

Greptile · repository intelligence

POST https://api.greptile.com/v2/query
Authorization: Bearer $GREPTILE_API_KEY   X-GitHub-Token: $GITHUB_TOKEN
{ messages, repositories: [{ remote: "github", repository, branch }], genius: true }
→ { message, sources: [{ filepath, linestart }] }
GREPTILE_API_KEY=...
GITHUB_TOKEN=ghp_...
GREPTILE_REPOSITORY=owner/repo
GREPTILE_BRANCH=main

Modal · execution sandbox

pip install modal && modal setup
modal deploy modal/simulate.py       # prints your web URL
MODAL_SIM_ENDPOINT=https://<workspace>--code-arena-sim-run.modal.run
MODAL_SIM_TOKEN=                     # optional bearer

Returned stages override the scripted profile per stage name, so a live container can change the outcome of the run.

AWS · scale layer — live with no key

GET https://dynamodb.{region}.amazonaws.com/   → { regions: [{ region, ms, up }], active, p50 }

An unauthenticated request returns a 4xx, which still proves the endpoint is reachable and yields a genuine round-trip time.

AWS_REGIONS=us-east-1,eu-west-1,ap-south-1

Stripe · skill marketplace

GET https://api.stripe.com/v1/prices?limit=6&active=true&expand[]=data.product
STRIPE_API_KEY=rk_...                # a restricted read-only key is enough

Failure posture

Every layer is independently degradable. A missing key, an expired token, a rate limit or dead wifi drops that one layer to its deterministic profile, marks it SIM on screen, and the pipeline continues to a verdict. Server timeout 12s, client 14s.

There is no configuration in which the pipeline hangs or hard-fails on an external service. That is a design constraint, not a happy accident — a demo that depends on venue wifi is not a demo.

curl -s localhost:5173/api/status | jq      # see exactly what's live
Route Returns
GET /api/status Per-layer mode and detail
POST /api/openai/analyze Reasoning verdict
POST /api/greptile/scan Repository query result
POST /api/modal/simulate Load simulation stages
GET /api/aws/regions Live region latencies
GET /api/stripe/skills Skill catalogue

8. The verification report

http://localhost:5173/report — a standalone engineering report for the run. Self-contained: inline styles, inline SVG charts, no build step, no runtime dependencies.

Section Contains
The diff 34 files, per-file churn bars, language mix, blast radius
What actually happened SVG code-health trace (100 → 72 → 100) and all 21 beats
What the agents changed Per-agent stats, then six patches with before/after diffs
Security Nine findings with file:line, CWE, CVSS — plus an SVG variant-spread map
Production simulation Stage results, first run vs. re-run
What the human approved The escalation, both options, and a full audit row
Infrastructure Every layer's role, results, and exact endpoint + payload
Verification Four proofs and what each covered
Impact What didn't reach production, engineering record, agent precision
Appendix Run metadata, district states, reproducibility, glossary

Served by server/report.ts, because Vite's SPA fallback would otherwise hand the arena to any extensionless path.


9. Module reference

src/arena/ — the 3D world

File Lines Responsibility
Scene.tsx 57 Canvas, lighting, tone mapping, postprocessing (bloom, chromatic aberration, noise, vignette)
Arena.tsx 194 Coliseum: platform, grid, rim light, 28 pillars, energy rings, 34 data streams, 7 code panels, dome
City.tsx 300 Five districts: tiered towers, health labels, crack overlay, debris physics, click-through inspection, dependency arcs
CodeCore.tsx 164 The repository crystal — dormant / healthy / attacked / repairing / verified
Ships.tsx 313 Four AI ships with flight physics — fly in, bank, orbit, strike, return — plus attack beams
SponsorInfra.tsx 258 Codex reasoning core, Greptile scan wave, Modal pods, AWS towers, Stripe pylons
Effects.tsx 145 Impact bursts and victory fireworks
CameraRig.tsx 190 Every shot in the film — per-view and per-phase framing, damping, shake, replay focus
config.ts 183 Palette, agents, layers, districts, dependency graph, building intel
textures.ts 90 Canvas-generated grid, code and window textures

src/state/ — the engine

File Lines Responsibility
battle.ts 533 The director: boot, battle, simulation, decision gate, replay, auto-demo
store.ts 306 Zustand store, all state and mutators
prs.ts 331 The three pull requests, fully specified
live.ts 66 API client with timeout and fallback

src/ui/ — the overlays

File Lines Responsibility
Overlays.tsx 203 Title card, phase banners, infrastructure beats, attack cards, verification, scoreboard, victory
HUD.tsx 177 PR card, infrastructure telemetry rack, fleet roster, arena feed, phase timeline, demo clock
Debrief.tsx 178 The developer report
CommandCenter.tsx 162 Mission queue, AI triage, infrastructure, fleet reputation, leaderboards
Inspector.tsx 86 Per-district panel — changes, vulnerabilities, recommendations, dependencies
DecisionGate.tsx 79 The human decision gate
Simulation.tsx 72 Production simulation panel
Hook.tsx 66 The opening terminal
Boot.tsx 60 Boot sequence
Logos.tsx 60 Infrastructure marks, drawn as SVG
Replay.tsx 38 Replay timeline
Finale.tsx 28 Closing card

Server, infrastructure and docs

File Purpose
server/sponsors.ts Live integration layer — holds the keys, one handler per layer
server/report.ts Serves /report
modal/simulate.py Deployable Modal load-simulation endpoint
public/report/index.html The verification report
src/index.css 828 lines — the entire neon design system
DEMO.md Click-by-click demo runbook with recovery steps
script.md Word-for-word video narration script

10. State and the battle director

view:    'hook' | 'boot' | 'intro' | 'command' | 'arena' | 'debrief' | 'finale'
phase:   'idle' | 'intro' | 'spawn' | 'attack' | 'repair'
       | 'simulate' | 'decision' | 'verify' | 'score' | 'victory'
core:    'dormant' | 'healthy' | 'attacked' | 'repairing' | 'verified'
agent:   'offline' | 'inbound' | 'standby' | 'scanning'
       | 'attacking' | 'repairing' | 'returning' | 'done'

Plus health and sub-stats, per-district state, infrastructure status, simulation results, the decision gate, the replay log and the leaderboard.

Writing the timeline

await sponsorBeat('greptile', [...], sleep)      // card in, log line, card out
s().hit(target, colour)                           // impact burst + camera shake
s().damageBuilding(target, amount, finding)       // health down, cracks, debris, marker
put({ health: to, core: 'attacked', attackCard: {...} })
mark('Bug Hunter · double refund', 'hit', target) // recorded for the replay
await sleep(3200)

Retiming the demo is editing sleep() values. Adding a beat is adding four lines.

Cancellation

A module-level token guards every sleep(). stopAll() increments it, so any in-flight timeline rejects at its next await and unwinds through a single catch. Restarting mid-run is always clean.

Blocking on a human

const choice = await awaitGate(my)   // polls gateChoice; auto-approves at 22s

The timeline genuinely stops. This is the one place the system yields control, and it is the point of the whole product.


11. Inside the 3D scene

Everything is procedural — no GLTF, no CDN, nothing to fail to load mid-demo.

Damage physics. Each district owns an InstancedMesh of 16 chunks. On a hit they launch outward on a ballistic arc with gravity and spin, and settle. On repair they interpolate back to origin and scale to nothing — particles running in reverse, which reads instantly as "rebuilding". The tower squashes on impact and blooms on repair; the crack wireframe's opacity is driven directly by health.

Ship flight. Ships lerp toward a destination that depends on status. Heading comes from the velocity vector and bank angle from velocity magnitude — so they lean into turns for free.

Camera. One rig owns every shot. Attack framing orbits on the target's bearing, between the city ring and the pillar ring, so the target is foreground and the core sits framed behind it — and the shot can never clip through a building. Damping is frame-rate independent (1 - 0.0012^dt) and impacts inject decaying shake.

Bloom and readability. Emissive materials use toneMapped={false} with high emissive intensity so they punch through the bloom threshold, while the platform stays deliberately dark. That is the difference between "neon" and "washed out" — the palette glows, not the lighting.


12. Extending it

Add a pull request

Append a PRDef to PRS in src/state/prs.ts. It automatically gets a queue row, triage reasoning, its own battle, simulation, decision and report.

{
  key: 'pr4924', num: 4924, title: 'Search Reindex', service: 'SEARCH SERVICE',
  branch: 'feat/reindex', risk: 'MEDIUM', files: 12, added: 380, removed: 90, author: 'you',
  scan: { understanding: 97, related: 2, files: 1842 },
  attacks: [ /* agent, target district, damage, stat hit, optional variants + beat */ ],
  fixes: [ ['database', 'reindex batched'] ],
  problems: [ /* location + suggested fix — these become the report */ ],
  triage: { rank: 3, why: '…', deploy: ['bug', 'perf'] },
  simulation: { users: 10000, /* … */ stages: [ /* passes: false triggers the human gate */ ] },
  decision: { headline, recommendation, confidence, options, recommended, reason, onApprove, onChanges },
  scores: { correctness: 97, security: 96, performance: 93, verification: 100 },
  bugs: 5, points: 300,
}

A stage with passes: false is what summons the human decision gate. That is the only switch.

Add a district

Append to BUILDINGS in config.ts, add an entry to BUILDING_INTEL, wire it into DEPENDENCIES. The city, links, scan sweep, inspector and camera all pick it up.

Add an agent

Add to AGENTS in config.ts with a colour, spawn angle, powering layer and reputation block, build a hull component in Ships.tsx, register it in the Ships() group.

Retime the demo

Every sleep(n) in battle.ts is milliseconds. The auto-demo's pre-roll is the wait(9000) in runDemo().


13. Console handles

In dev, window.arena is exposed for driving the thing live — useful mid-presentation.

arena.demo()                       // run the 2:30 auto demo
arena.command()                    // jump to the command center
arena.battle('pr4923')             // launch a specific mission
arena.boot()                       // replay the boot sequence
arena.store.getState()             // inspect everything
arena.store.getState().set({ view: 'finale' })          // jump to the closing card
arena.store.getState().set({ inspecting: 'payments' })  // open a district panel

14. Demo and recording

  • DEMO.md — the presenter's runbook. Beat-by-beat, what to say and when to click, live-vs-simulated setup, and a recovery table for when something wedges on stage.
  • script.md — the word-for-word video narration, timed to the arena's own pacing, with the three pauses that carry it and a 60-second cut.
Time Beat
0:00 – 0:30 The problem — terminal, then who verified it?
0:30 – 0:50 Activation and the command center
0:50 – 1:40 The battle — scan, fleet, attacks, the variant chain
1:40 – 2:00 Repair, then the production simulation failure
2:00 – 2:20 The human decision, verification, the report
2:20 – 2:30 Final vision

If you're recording: wire Codex/OpenAI and Modal first. Two minutes of setup, and it is the difference between "nice visualisation" and "this is real."


15. How this was verified

Every stage of this build was driven in a real headless Chrome — boot, command center, full battle, decision gate clicked, replay interaction, report rendering — with console and page errors captured. Final state: 0 errors, simulation passed, 20 replay events recorded, 620 points awarded.

Bugs that pass caught, none of which were visible from reading the code:

Bug Why it mattered
Attack camera swept through buildings The hero shot of the whole demo was a close-up of a roof
Framer Motion writes an inline transform Silently killed CSS transform-centering on the banner, attack card, sponsor card, verification panel, scoreboard and command-center header
.sim class collision The simulation panel's class also matched em.sim, flinging LIVE/SIM badges across the page
Ship name tags ballooning Drei Html scales inversely with distance; a ship near the camera produced full-screen ghost text
Arena feed overflowing its panel Log lines escaped the container
Code panels washing out the attack shot Additive planes sat between camera and city
AWS labels dominating the closing shot The finale camera flies past the towers
Roof caps wider than the tiers below Buildings read as flat slabs

Typecheck (tsc -b) and production build are clean.


16. Specification traceability

Five specification documents live alongside the code. Everything in the first four is implemented.

Spec Status
Code_Arena_3D_Arena_Specification.md ✅ Arena, code core, repository city, agents, attack/repair, health, verification, victory
Code_Arena_PRD_v2_Arena_Evolution_Upgrade.md ✅ Command center, PR queue, infrastructure visualisation, AI ships, physics, feedback, decisions, rankings, loading
Code_Arena_Enhancement_Addendum_Specification.md ✅ Repository intelligence overlays + click-through, agent reputation, production simulation, human decision layer, battle replay
Code_Arena_2_5_Minute_Demo_Flow_Specification.md ✅ Problem hook, activation, mission arrival, battle, repair + simulation, command center reveal, final vision
Code_Arena_Hype_Demo_Flow_Final.md ⬜ Not yet implemented

Beyond the specs: real API integration with LIVE/SIM badging and guaranteed fallback, the /report verification report, and the demo and script documents.


17. Troubleshooting

Symptom Cause and fix
/report shows the arena The dev server needs a restart after a vite.config.ts change — the route is a plugin middleware.
All layers say SIM Expected with no .env. curl localhost:5173/api/status shows which key each layer is missing. AWS should still say LIVE.
A layer says SIM despite a key The reason is in the API response — bad key, wrong repo slug, or an undeployed Modal endpoint. Hit the route directly to see it.
OpenAI returns an error OPENAI_MODEL is probably a model your key can't reach. Set one you have access to; the run falls back safely either way.
Low frame rate Close other GPU-heavy tabs. Bloom and postprocessing are the expensive part; dpr is capped at 1.8 and drei's AdaptiveDpr degrades under load.
Battle seems stuck It's the decision gate waiting for you. It auto-approves after 22s, or click ✓ APPROVE.
Demo wedged mid-run arena.command() in the console, or ⌂ COMMAND CENTER at the bottom of the arena.
Fonts look wrong Orbitron and JetBrains Mono come from Google Fonts. Offline, the monospace fallback is used and everything still works.

Software engineering is moving from humans reviewing code

to AI systems proving software quality.

Built for the Greptile hackathon.

About

Code Coliseum — an experimental 3D interface for AI code review, repair, and verification. Built with React Three Fiber.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages