Replay & E2E Testing
Record a session as an .ad script, then run it again with replay or run a folder of scripts as an E2E suite with test.
How it works
You work in two passes:
- Explore: discover elements and act on refs (
snapshot->click @e../fill @e..) while recording. - Replay: run the recorded
.adscript withreplayfor a deterministic run.
Record a replay script
Pass --save-script to open. When you close the session, the script is written:
By default, the script goes to:
To choose the output file, pass a path to --save-script:
- Missing parent directories are created.
- The script is written on the machine running the daemon, so
--save-scriptis rejected when you use a remote daemon. - For a bare file name that could be read as another argument, use
--save-script=workflow.ador a path-like value such as./workflow.ad.
.ad line grammar
A .ad line is the CLI spelling of one command: <command> [positional ...] [flag ...]. Whitespace separates tokens, so a value with a space needs quotes.
- A token quoted with
"or'is one argument. Single quotes keep a"literal, so'id="far-button"'and"id=\"far-button\""are the same selector — write whichever matches how you typed the command at the shell. - Values in double quotes are JSON strings, so escape
\\,\",\t, and\n. Values in single quotes are literal, as at the shell: a backslash keeps its own character, and the only escape is\'for an apostrophe. - A script line keeps only the flags that command records; CLI-only spellings and per-request options are not part of a step. The common flags a script does not carry are
--settle,--verify,scroll --pixels/--duration-ms, and the device-selection flags (--platform,--serial,--device). help <command>prints the flags each command accepts.help scriptingprints this grammar.
To reach an off-screen element, use a stop condition rather than a fixed scroll amount. A fixed amount passes on one screen size and fails on another:
A # only starts a comment at the beginning of a line. A scroll line carries --until and keeps its distance as a positional (scroll down 0.8 --until <selector>). wait carries --raw, --depth <n>, and --scope <selector|@ref> to choose the capture its target is read from. Scripts use only these long spellings: the -d/-s CLI aliases are not flags in a script, so a hand-written line such as wait text -s so funny still waits for that literal text.
Run replay
-
Replay reads
.adscripts. -
The CLI reads script paths on your machine and sends the script, with the Maestro
runFlowincludes it can resolve and read, to the daemon. An include it cannot read is left out, and the run fails if it reaches that include. The samereplayortestcommand works against a local or a remote daemon without copying files. A script that doesn't exist on your machine fails at once, naming the path you typed. -
A script that does not end in
closeleaves its session open. For a script that does end inclose, pass--keep-sessionto skip only that final action and keep working in the same session:Earlier
closeactions still run.testdoes not accept the flag because each suite attempt cleans up its own session. -
press,click, andlongpresssteps wait up to 2 seconds for their target to appear before they fail, so a step recorded against a screen that was still loading passes once the target shows up. The wait covers only a target that is not on screen yet: a target that is covered, off-screen, or matched by more than one element fails at once, as it does live. The replay output does not show the wait for a step that passed; run with--debugto see it in the diagnostics. -
When the target never appears, replay stops with
REPLAY_DIVERGENCE, anderror.details.readinesssays how long the step waited and how many times it looked (waitedMs,polls,end). For a step recorded with a target annotation (the# agent-device:target-v1line above it),error.details.divergence.kindisselector-missand the step is never sent. For a step without an annotation,error.details.reasonisselector_not_found, as for a live command. If the app shows an empty accessibility tree during that wait,error.details.reasoniscapture_sparseinstead; take a snapshot to see where the app is. -
A step whose read names one element (
iswith a predicate other thanexists/absent, orget attrs) that reaches dispatch with an ambiguous selector fails withAMBIGUOUS_MATCHas the divergence cause — the same code the live command reports, so an ambiguous recorded screen is not replayed as a missing one. An annotated step can also stop earlier, before the command runs: when pre-dispatch target binding cannot confirm the recorded identity (its identity set has more than one member with no sibling or viewport signal isolating one, or no fresh snapshot can be captured to verify against), the divergence cause isIDENTITY_UNVERIFIABLEinstead. Both codes say the recorded screen changed; onlyAMBIGUOUS_MATCHproves your selector matched more than one element.
Run Maestro compatibility flows
Pass --maestro to replay or test to run Maestro YAML flows. Only the subset below is supported:
Supported subset:
- Flows:
launchApp(withclearState,permissions, and Apple-only launch arguments;permissionsapply after state clearing but before launch, and alaunchAppwithoutpermissionstouches nothing — there is no silentall: allowdefault);setPermissions(mid-flow permission grants, denials, and resets;allresolves in the backend — one simctl call on iOS, the declared permissions on Android — with specific entries overriding after it);runFlowfile/inline;runFlow.whenandrepeat.whileconditions (platform,visible,notVisible, andtrue, all re-evaluated before everyrepeatiteration);onFlowStart/onFlowComplete;repeatwithtimes,while, or both; and retry. - Interactions:
tapOn,doubleTapOn,longPressOn,inputTexton the focused element,eraseText,openLink,hideKeyboard, basicpressKey, andback; selector targets poll until available and support recursiveindex,childOf,above,below,leftOf,rightOf,containsChild,containsDescendants, points, andoptional; outer command labels are metadata, not target selectors. - Assertions and navigation:
assertVisible,assertNotVisible,assertTrue(literal values and${VAR}lookups only;"","false","0","null", and"undefined"are falsy, everything else is truthy),extendedWaitUntil,scroll,scrollUntilVisible, absolute/percentage/targetswipe,takeScreenshot,waitForAnimationToEnd,clearState, andstopApp. - Scripts: ordered
runScriptfile/env scripts withhttp.post,json, andoutputvariables;evalScriptinline expressions run flow-scoped JavaScript and writeoutput.*leaves for later steps.
Boundaries:
- Permissions: every entry is one
settings permissioncall, applied in order withallfirst; the step stops at the first entry the selected platform refuses, earlier entries stay applied, and the error names what landed. Android’s only allow level is while-in-use, solocation: inuseandlocation: nevermeanallowanddenythere, whilelocation: alwaysandphotos: limitedare Apple-only and fail. On iOS, which service a runtime changes issimctl privacy’s own verdict: current runtimes refuse a targetednotificationschange and leave notifications untouched underall. - Runtime: iOS and Android only;
launchApp.clearStateand standaloneclearStatesupport Android and iOS simulators, launch arguments are Apple-only, and other standalone device utility/state commands are unsupported. - Expressions:
evalScriptand conditiontrue:fields (runFlow.whenandrepeat.whileshare one evaluator) are evaluated as JavaScript (flowenvand prioroutputleaves are string-typed); atrue:field that is a boolean, amaestro.platformcomparison, or plain literal text after${VAR}lookups is decided without JavaScript, with theassertTruefalsy table for literal text. Other fields stay literal or${VAR}lookup-only —assertTruesupports literals and bare lookups, and other expression-shaped payloads fail loud. - Environment: flow
envis the default,AD_VAR_*overrides it, and CLI-e KEY=VALUEwins over both. - Failure diagnostics: resolved targets and
runFlowpaths are rendered, whileinputTextpayloads remain hidden; do not place secrets in diagnostic identifiers. - Trust:
runScript,evalScript, and JavaScript conditiontrue:fields execute flow scripts in-process vianode:vm, which is not a security sandbox;runScriptmay makehttp.postnetwork requests and its output keys cannot contain a dot.evalScriptandtrue:fields that need JavaScript are refused outright for a flow accepted over the daemon’s remote HTTP surface, since that context can escape to the host. - Errors and tracking: unsupported commands and fields fail with source context when available; open a focused issue only when implementation work is planned.
- Session takeover:
--keep-sessionis a native.adreplay option and is rejected for Maestro YAML.
ADR 0015 lists the deliberate deviations from Maestro. If a missing feature matters for your suite, open a focused issue with a small flow snippet.
Export .ad scripts to Maestro YAML
To run a recorded flow with Maestro, export the .ad script to Maestro YAML:
replay export only converts the file: it does not start the daemon or contact a device. Without --out, it prints the YAML to stdout.
Each open <appId> exports with an explicit launchApp.appId, so a flow can switch between apps and return to the original app. The first app remains the flow's default appId; relaunch options and app-specific deep links stay attached to their authored targets.
Deep links, including schemes without // such as tel: and mailto:, export as openLink. A standalone open tel:+15551234567 emits only the link command; open com.example.app mailto:[email protected] emits the app launch followed by the link.
Export is strict. It writes Maestro YAML for compatible flow actions such as app launch, taps, long press, text input, keyboard dismiss/enter, back, home, text visibility assertions, coordinate swipes, basic scroll, screenshots, and .ad env directives. home exports as pressKey: Home, so flows that visit the home screen and reopen the app can be exported. Agent-only inspection or maintenance actions such as snapshot, get, record, trace, settings, and unsupported selector shapes fail with the source line and action instead of being silently dropped. Known semantic differences are reported as warnings; for example, .ad fill exports as tapOn plus inputText, which may append text in Maestro rather than replacing existing field contents. Native .ad label= selectors export as Maestro text: selectors and warn because Maestro text matching is broader than label-only matching. A strict wait absent is reported as unsupported rather than mapped to Maestro's more lenient notVisible condition.
Run a lightweight .ad suite
testdiscovers.adfiles from files, directories, or globs and runs them serially.- Quote relative globs to expand them on the caller from its working directory, including when the directory name contains glob characters such as
[or{. A missing file input without glob characters reports an error. - The
context platform=...header inside each.adfile decides which platform that file runs on. --platformfilters which files run; with a filter, files without acontext platform=header are skipped. When filtering leaves no runnable sources, the no-match error reports how many sources were skipped for having nocontext platform=header versus how many declared another platform. Add the header to run a file under a filter, or omit--platformand let the selected device decide.- Set
context timeout=...andcontext retries=...per script; CLI flags override them. Retries are capped at3, and duplicate keys in the context header fail fast instead of silently overriding each other. - By default, suite artifacts are written under
.agent-device/test-artifacts/<run-id>/.... Each attempt writesreplay.ad,result.txt, andreplay-timing.ndjson. Failed attempts also keep copied logs and artifact files when the replay produced them. - Copied diagnostic artifacts receive numbered filenames when their names collide with another artifact, a replay source, timing trace, or attempt manifest.
result.txtlists the retained names incopiedArtifacts. replay-timing.ndjsonrecords attempt, cleanup, and per-step start/stop events with durations. Upload it from CI even for passing runs when comparing local and CI performance.- When an attempt hits its timeout, it is marked failed and the replay gets a short grace period to stop before the session is cleaned up.
- The default text reporter streams live progress on stderr while a suite runs, then prints the final summary, failed tests, and passed-on-retry flaky tests. Use
--verboseto include step traces in completed-test progress output. - The default reporter prints a
Warnings:section after the summary when any test accumulated composable warnings — for example a Maestro step withoptional: truethat was skipped — whether that test passed or failed. A failingreplayrun repeats the warnings it accumulated asWarning:lines after the error.--jsoncarries the same strings in each test result'swarningsarray. --reporteris repeatable. Built-ins aredefaultfor the console summary andjunit:<path>for JUnit XML. Passing any explicit reporter list replaces the implicit default reporter, so include--reporter defaultwhen you also want terminal output.--report-junit <path>is an alias for--reporter junit:<path>.- JUnit reports preserve legal Unicode and whitespace, and replace characters forbidden by XML 1.0 (such as terminal ESC or NUL) with
U+FFFD(�) so CI parsers can read the report. JSON and other reporters retain the original suite values. - When
--fail-fastand retries are both set, the current test still consumes its retries before the suite stops.
Custom test reporters
A custom reporter formats suite output. It runs in the local CLI process, not the daemon, and can render both live progress and the final result.
A reporter module can export a reporter object, reporter, createReporter, or a default factory. Factories receive a load context. Reporter hooks receive replay test objects and an IO context with stdout and stderr streams:
For a live terminal reporter that prints each completed test as an emoji, title, and duration:
TypeScript reporters use the same object shape; compile them to JavaScript before passing them to --reporter:
The CLI loads reporter modules with Node dynamic import(). Use .mjs or .js files at runtime; for TypeScript, compile the reporter to JavaScript before passing it to --reporter. Loading .ts files directly depends on Node's type-stripping behavior and is not part of the supported reporter contract.
The live hooks onSuiteStart, onTestStart, onTestStep, and onTestResult run while the suite is running; reporters do not receive generic command progress. Live hooks run as events arrive and are not awaited, so keep their work synchronous and move anything async to onSuiteEnd, which the CLI awaits before exiting. onSuiteEnd receives the final suite result. getExitCode can only raise the suite exit code, never lower it: the highest reporter-provided code wins and failed tests still exit with 1 when no reporter raises it further, so a reporter cannot mask a failing suite. Return an integer from 0 to 255, or undefined to leave the exit code unchanged. Other values fail with INVALID_ARGS; in particular, codes such as 256 are rejected before they can wrap to a successful process exit.
Parametrise .ad scripts
Substitute ${VAR} tokens in .ad scripts using values from the CLI, shell env, script-local env directives, or built-ins.
Precedence
Built-ins
replay and test provide these built-ins in the reserved AD_* namespace.
AD_PLATFORM- matchescontext platform=...or the selected platform when availableAD_SESSION- active session nameAD_FILENAME- path of the running.adfileAD_DEVICE- device identifier (when--deviceis set)AD_ARTIFACTS- attempt artifacts directory (when running undertest)
User-defined keys starting with AD_ are rejected in env, -e, and shell imports such as AD_VAR_AD_FOO, so built-ins cannot be overridden.
Substitution happens inside parsed string values. It does not create extra arguments, so quote selectors or text values that contain spaces:
Fallback and escape
${VAR:-default} yields default when VAR is unset.
\${APP} emits a literal ${APP} with no substitution.
Recipes
Run one flow against two app variants in CI:
Tune timings locally without editing the script:
Extract a reusable selector. Before:
After:
Quote ${VAR} inside selector expressions so the whole expression is treated as a single argument.
Notes
AD_VAR_*values come from the shell that runs the CLI, so they apply the same way whether the daemon runs locally or remotely.- Fallbacks do not nest:
${A:-${B}}is not supported. - An unresolved
${VAR}fails with afile:linereference, so a misspelled variable name stops the run.
Replay divergence and resume
A failing replay/test step returns a structured REPLAY_DIVERGENCE error. The report is size-bounded and redacted, and carries:
step— the 1-based plan index and its source file/line (through MaestrorunFlowincludes).screen— a fresh post-failure snapshot digest with actionable refs, orunavailablewith a reason/hint when capture failed or was sparse (never a stale tree).suggestions— up to 5 ranked, re-resolved candidates for the failing selector (id match ranks above role+label, which ranks above label-only), each with abasisyou can inspect before acting.resume— whether and how to continue without re-running the script from the top.
Text output prints a compact summary of the same fields; --json/MCP carry the full object.
Resume a failed replay
replay --from <n> --plan-digest <sha256> resumes at plan step n, not after it, skipping 1..n-1 without executing them. Both flags come from a divergence report's resume field — from is the failed step, planDigest is the digest of the exact unchanged plan that produced it.
Choose one recovery workflow:
- Change the replay script. Review the suggestion, edit the selector or include, then run a fresh full
replay ./flow.ad. The old digest is intentionally invalid after any plan edit; do not combine it with the edited script. A later divergence supplies a new digest. - Keep the replay plan unchanged. Repair app/device state so the reported failed step can succeed when retried, then resume with the report's unchanged
fromandplanDigest. If you manually complete the failed action itself, the reportedfromwill execute it again; only do that when repeating the action is safe.
The unchanged-plan resume loop is:
- Run
replay ./flow.ad. On failure, readresumefrom the divergence. - Leave the script, includes, platform, and target unchanged. Repair app state yourself so the failed step can be retried safely. Resume does not restore app state; it only skips the earlier steps.
replay ./flow.ad --from <resume.from> --plan-digest <resume.planDigest>.
For Maestro flows, --from counts steps in the top-level plan. A runFlow with no condition, or with
a condition that resolves before the run, is flattened into its commands (or dropped when the condition
is false). A control step decided at run time (runFlow, repeat, or retry) counts as one step, and
you cannot resume at a command nested inside it. As with .ad replay, restore any state and
environment values the skipped steps would have set before you resume.
Passing --plan-digest that no longer matches the current script — because you edited it, an include changed, or platform-conditioned expansion differs — fails INVALID_ARGS before any action; run a fresh full replay to get a new digest. --from is replay-only; test rejects it (a suite run must stay full and deterministic).
--update/-u (retired)
--update/-u does not rewrite .ad files. The flag is accepted and does nothing: every replay divergence carries ranked suggestions whether or not you pass it. Review a suggestion, then edit the .ad file yourself if it's right. ADR 0012 explains why automatic rewriting was retired.
Troubleshooting
- Replay fails after UI/layout changes:
- Read the divergence report's
suggestionsand repair the selector by hand; there is no automated rewrite. Because the edit changes the plan digest, run a fresh full replay instead of using the old resume flags.
- Read the divergence report's
- Repeated re-runs are slow or the app is stateful, but the script is still correct:
- Repair app state and resume with the unchanged
--from/--plan-digest. See Resume a failed replay.
- Repair app state and resume with the unchanged
- Replay file parse error:
- Validate quoting in
.adlines (unclosed double quotes are rejected). A selector with a space needs quotes, andhelp scriptingstates how each quote form decodes.
- Validate quoting in
- A replay passes on one device size and fails on another because a target was off screen:
- The script used a fixed
scrollamount. Replace it with the stop condition,scroll down --until <selector>, which repeats until the element is actually on screen.
- The script used a fixed
- A
pressorclickstep fails because its target was not found, but the element is on the screenshot:- A
selector-missdivergence, orerror.details.readiness.end: expired, means the element was not in the accessibility tree for the whole 2-second wait: the selector is wrong for this build, or the element is not exposed to accessibility.readiness.end: sparsemeans the app showed an empty tree: the screen was mid-transition or the app had left. Add awaitstep for a landmark on the new screen before the press.
- A
- Maestro compatibility flow fails on unsupported syntax:
- Check ADR 0015. If the missing feature matters to your suite, open a focused issue with a small flow snippet.
