Skip to main content
Visor keeps navigation knowledge on the host so AI agents can reuse what they already learned about a mobile app.

Install the repository skill

Install the skill into the mobile app repository so compatible coding agents receive the same map and interaction rules:
This creates .agents/skills/visor-discovery/ and skills-lock.json. Commit those files with the app repository. Keep .visor/ ignored because it contains private machine-local maps and runtime state. For a complete installation and initial discovery run, paste the prompt from Agent setup into your coding agent.

Delegate device work to Navigator

The setup prompt also copies project-scoped Navigator definitions to .codex/agents/navigator.toml and .claude/agents/navigator.md. Add root AGENTS.md and CLAUDE.md guidance that delegates every device observation and action to Navigator. The parent agent should describe the desired outcome and current permission boundary; Navigator should own Visor commands, recovery, live verification, and map maintenance. Navigator reads the installed visor-discovery skill before it touches the device. It reuses known actions and routes, repairs typed failures deliberately, and annotates new or changed screens, actions, selectors, transitions, and recovery paths as soon as the live app proves them. It never repairs a map by running generic crawl or editing runtime map files directly.

Two storage layers

Visor separates agent context from runtime evidence:
  • The runtime index stores source fingerprints, elements, variants, and edge evidence that Visor needs to drive Appium.
  • Compact agent memory stores semantic screens, meaningful actions, executable destination selectors, known routes, reliability, and unresolved gaps.
visor discover returns the current compact slice in data.memory and persists the complete agent file at data.map.agent_path. Do not load the runtime index at data.map.path into context. The compact file excludes raw UI source, screenshots, raw element inventories, form values, dynamic financial text, and identity-like content.

Fresh discovery

Use the repository’s persistent .visor/maps directory for normal work. Preserve it between agent tasks so known screens and routes do not need rediscovery. Before expanding the map, agree on an exploration policy with the user: safe-only, explicitly scoped, or full access for the specific test environment and account. Ask how to handle login or whether you may register a test account. Do not infer credentials, account-creation permission, or authorization for dangerous actions. Use an explicit empty temporary directory only when you need proof that an agent started without prior knowledge:
The response includes an observation_token. Annotate that exact observation without reading the device again:
Populate the map through AI-assisted, screen-by-screen discovery: interpret compact memory, annotate the exact observation, choose the next action under the agreed permission policy, execute it, and checkpoint the resulting stable screen. Never use discover --crawl to populate or expand semantic memory. Generic crawl remains a low-level diagnostic compatibility tool; it cannot assign product meaning, honor user-specific forbidden actions, or handle authentication interactively. Keep each action’s factual safety classification. Explicit full test access can authorize an agent to execute a risky or destructive action, but it does not turn that action into a safe route step. Execute authorized consequential actions directly, one at a time, and observe the result immediately. Copy executable recognizers exactly from compact memory, including whitespace and line breaks. Use screen-specific recognizers for path eligibility, and keep composite or unusually large accessibility containers marked unknown until you verify a precise safe target.

Deterministic routes

Use visor route <plan.json|-> to execute an agent-authored route in one daemon request. A plan can provide multiple ordered paths. Each step includes:
  • a supported command and arguments;
  • safety: "safe";
  • an expected semantic screen;
  • an executable selector and timeout that prove the destination.
Visor rejects unsafe or malformed plans before device selection. It checks each path’s optional starting selector, executes eligible paths in document order, checkpoints every step, and stops after the first complete path. For one known action, reuse the stored command directly. For multiple steps or alternate paths, prefer one route request over separate CLI processes. This keeps the Appium session warm and avoids repeated source capture around known navigation.

Unknown states

When an action reaches an unexpected screen, Visor records a candidate transition and refreshes compact memory. It then tries the next supplied path whose starting selector matches the live screen. If no path applies, the route result returns status: "needs_discovery" with:
  • the exact observation token;
  • compact current-screen memory;
  • unresolved semantic gaps;
  • the agent-memory path;
  • the route-checkpoint path.
You can annotate the unknown screen, add a safe recovery path, and resubmit the plan. Visor does not guess consequential actions. The atomic checkpoint stores the validated plan, completed attempts, current observation token, and next route position. You can inspect it independently after a partial run and construct a recovery plan without reconstructing the failed request from logs.

Typed route outcomes

Each route step reports one outcome:
  • success
  • runtime_failure
  • verification_failure
Each path reports completed, failed, or skipped. Runtime failures do not count as bad locator evidence. If Appium drops a cached session, the daemon recreates it once and reruns the safe plan.

Session lifetime

Visor sets Appium’s command idle timeout to 600 seconds by default so an agent can reason between commands. Set VISOR_APPIUM_NEW_COMMAND_TIMEOUT_SECONDS to change it. Use --no-map or VISOR_NO_MAP=true to bypass map reads and writes for ordinary direct commands. Deterministic routes require map persistence because they checkpoint progress and rediscovery state.