Skip to content

Add local iOS and Android natural-language E2E testing - #1

Open
perixtar wants to merge 29 commits into
mainfrom
feat/mobile-testing
Open

perixtar wants to merge 29 commits into
mainfrom
feat/mobile-testing

Conversation

@perixtar

@perixtar perixtar commented Sep 19, 2026 •

Copy link
Copy Markdown
Owner

Native apps could not be tested with the web-only jev-e2e runner. This adds a local iOS Simulator and Android Emulator alpha using the same Case/Goal/Step/Expect contract. jev-e2e devices finds a virtual device, plan reviews the cases, and run --platform ios|android --app ... --device ... executes them through pinned agent-device@0.21.6. Jev selects observed controls, while independent accessibility assertions decide PASS, FAIL, or BLOCKED. The optional OpenRouter prose planner is skipped with --planner off; successful flows can be replayed, with Jev repairing stale controls only when needed.

Native fixtures remain local, the runner checks app/device identity before dispatch, and inaccessible or ambiguous evidence blocks. The CLI and local workbench share cancellation, reports, limits, and opt-in recording. Private screens are excluded from shareable artifacts; the live preview is atomically published only after verification. Saved build-path replays pin the installed bundle/package identity. The README and native guide include setup, explicit examples, an owned React Native shopping fixture, a large-phone ten-second video, and real-time recorded segments with capture cuts disclosed.

Validation:

  • Controlled matrix at measured implementation SHA 96a275e8c452fbc40ad5b52d83c538664edc0b64: five flows × ten first attempts × healthy/fault/replay per platform, 300 cases total. Each platform had 50/50 healthy PASS, 50/50 correct seeded-fault FAIL and zero false PASS, and 50/50 replay PASS. Warm healthy medians: iOS 27.20 s, Android 10.57 s; total billed Jev cost $0.055746222. Both replays made zero prose-planner calls; iOS needed 30 Jev repair calls. The 40-case prose cohort compiled 40/40 at $0.0248056. Source trials, registered-cohort boundaries, and limitations are in MOBILE_TEST_RESULTS.md and the sanitized JSON.
  • Real pinned-SDK ownership, typing, switch, recording, privacy, and cancellation gates passed on both devices. A clean packed CLI on this PR runtime returned healthy PASS/exit 0 and seeded invalid-login FAIL/exit 1 on each device, with CLI/JSON/HTML agreement. A packed install without the optional SDK compiled an explicit plan offline and gave actionable missing-SDK guidance in doctor.
  • Final candidate f5390630bc44f6ca8229eee7db9c9f06c7d1b5f9: npm run check, full npm test (55/55 on macOS and 55/55 in a clean Linux Playwright container, including real Chromium flows), isolated Linux command-deadline test (1/1), focused native/preview tests (29/29 on the unchanged runtime), and git diff --check passed. Both GitHub push and PR checks passed on this exact SHA, including npm pack. The 10-second and 58.934-second videos were decoded and inspected; the short file's MP4 header reports exactly 10.000 s. The published 300-case data and filmed runs are historical measurements at 96a, not a claim that all 300 were rerun after the review fixes. The last two commits after packed device smoke change only mock test setup; runtime, package, data, and media bytes are unchanged.

Independent regression review: a separate reviewer checked exact SHA f5390630bc44f6ca8229eee7db9c9f06c7d1b5f9 against latest target origin/main SHA 509f6376797563cf8ce140a3eb21b74a82c4f668, including the complete 54-file diff, related callers/consumers, contracts, app/device identity, ownership, state transitions, concurrency/cancellation, authentication/privacy, failure paths, UI/report/evaluator, package, and media. The reviewer independently ran the full 55-test suite, focused 29-test native/preview set, typecheck, diff checks, and cross-checked all 60 public trials against the private originals; the final test-only deltas were checked against this full audit. Findings fixed before the final SHA: preview publication race; short private Goal literals and descriptor binding; empty-cart Goal requiring a matching final assertion; build-path replay identity pin; benchmark-cohort provenance wording; Linux mock platform selection; and a mock worker whose absent IPC handle could cancel an unreferenced deadline timer. No material finding remains.

Remaining scope: local virtual devices and accessible controls only; physical phones, hosted execution, custom canvas/complex WebViews, biometrics, and payments need separate validation. These controlled results do not establish reliability on arbitrary apps. Replay on iOS was not model-free; the real-time video omits credential and other unsafe screens rather than showing every second of the full run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant