Node.js API
Use createAgentDeviceClient() for typed, deterministic device automation from Node.js instead of shelling out to the CLI.
Building an agent? Start with the dedicated AI SDK or Eve integration.
Runnable examples
The repository includes runnable, typechecked Node.js examples. They import the same published agent-device/* entry points your code uses:
Each example typechecks against the SDK through its own tsconfig.json. Build the package with pnpm build, then run an example with Node:
API reference
Public entry points:
agent-devicecreateAgentDeviceClient(options?)createLocalArtifactAdapter(options?)AppError,isAgentDeviceError(error),normalizeAgentDeviceError(error)centerOfRect(rect)- types:
AgentDeviceClient,AgentDeviceClientConfig,AgentDeviceDevice - types: every option and result type of the client's methods, among them
AppOpenOptions,AppOpenResult,CaptureSnapshotOptions,CaptureSnapshotResult,CaptureScreenshotResult,PressOptions - types:
SnapshotNode,RawSnapshotNode,SnapshotState,Rect,Point - types:
NormalizedError,AppErrorCode,KnownAppErrorCode,AppErrorDetails,ErrorCause
agent-device/iocreateLocalArtifactAdapter(options?)- types:
ArtifactAdapter,ArtifactDescriptor,CreateTempFileOptions,FileInputRef,FileOutputRef,LocalArtifactAdapterOptions,OutputVisibility,ReserveOutputOptions,ReservedOutputFile,ResolveInputOptions,ResolvedInputFile,TemporaryFile
agent-device/metrobuildBundleUrl(baseUrl, platform)normalizeBaseUrl(baseUrl)resolveRuntimeTransport(runtime)prepareMetroRuntime(options?),reloadMetro(options?),stopMetroTunnel(options)- types:
MetroBridgeDescriptor,MetroTunnelRequestMessage,MetroTunnelResponseMessage
agent-device/batchrunBatch(req, sessionName, invoke)
agent-device/remote-configresolveRemoteConfigProfile(options)- types:
RemoteConfigProfile
agent-device/contractscenterOfRect(rect)defaultHintForCode(code),normalizeError(error)- types:
DaemonError,DaemonInstallSource,DaemonRequest,DaemonResponse,DaemonResponseData,JsonRpcId,JsonRpcRequestEnvelope,LeaseBackend,SessionRuntimeHints
agent-device/selectorsparseSelectorChain(expression)tryParseSelectorChain(expression)resolveSelectorChain(nodes, chain, options)findSelectorChainMatch(nodes, chain, options)listSelectorChainMatches(nodes, chain, options)formatSelectorFailure(chain, diagnostics, options)isNodeVisible(node)isSelectorToken(token)isNodeEditable(node, platform)- types:
SelectorChain,SelectorDiagnostics
agent-device/findersfindBestMatchesByLocator(nodes, locator, query, requireRectOrOptions)parseFindArgs(args)- types:
FindMatchOptions
agent-device/install-sourceARCHIVE_EXTENSIONSisTrustedInstallSourceUrl(sourceUrl)(deprecated; install sources are not gated on it)validateDownloadSourceUrl(url)- types:
MaterializeInstallSource
agent-device/artifactsresolveAndroidArchivePackageName(archivePath)
agent-device/android-adbcreateAndroidPortReverseManager(provider)/createAndroidPortReverseManager(executor, { noRebind })captureAndroidLogcatWithAdb(executor, options?)readAndroidClipboardWithAdb(executor)/writeAndroidClipboardWithAdb(executor, text)getAndroidKeyboardStatusWithAdb(executor)/dismissAndroidKeyboardWithAdb(executor)openAndroidAppWithAdb(executor, packageName)forceStopAndroidAppWithAdb(executor, packageName)listAndroidAppsWithAdb(executor)getAndroidAppStateWithAdb(executor)- types:
AndroidAdbExecutor,AndroidAdbExecutorOptions,AndroidAdbProvider,AndroidKeyboardState,AndroidKeyboardDismissResult,AndroidPortReverseEndpoint
agent-device/limrunnew LimrunRuntime(options)runtime.getDeviceSession(device)- types:
LimrunRuntimeOptions,LimrunDeviceSession,LimrunAndroidDeviceSession,LimrunIosDeviceSession,LimrunIosCommandExecution
agent-device/plugins/webdriver- experimental
WebDriverPluginOptionsfor providers using the shared engine.
- experimental
agent-device/plugins- experimental factory context:
ProviderPluginHost; see provider plugins.
- experimental factory context:
agent-device/ai-sdkcreateAgentDeviceTools(options)- types:
AgentDeviceToolSet,AgentDeviceTools,CreateAgentDeviceToolsOptions
Basic usage
This example opens an app, takes a snapshot, interacts, handles typed errors, and always closes the session. Run it from examples/sdk/client-session.ts.
client.devices.capabilities() returns { device, availableCommands }, using the same capability matrix as the CLI. Use it to check which command names the selected target supports.
For direct iOS simulator app launches, client.apps.open({ app, platform: 'ios', launchConsole: './artifacts/app.console.log' }) captures launch-time
stdout/stderr. The option mirrors open --launch-console and is not valid for URL opens or non-simulator targets.
client.sessions.stateDir() mirrors session state-dir and returns the resolved daemon state directory without starting
or contacting the daemon. Pass { stateDir } to resolve an explicit override the same way the CLI resolves --state-dir.
client.sessions.artifacts({ provider, providerSessionId }) mirrors artifacts --provider ... --provider-session ... and returns provider-hosted cloudArtifacts.
Use it for BrowserStack, AWS Device Farm, or TestMu AI session videos/logs after a cloud session has stopped; client.sessions.close() returns the ID as closed.provider?.providerSessionId. Omit providerSessionId when an embedding host has registered a provider runtime that can infer the active lease. Limrun does not expose provider artifacts through this command.
Device cloud sessions
You drive Limrun, BrowserStack, AWS Device Farm, and TestMu AI devices with the same typed client methods. Use the CLI connect flow to persist connection state locally. Pass provider settings in the client config when your Node.js integration already owns credentials and provider selectors.
apps.open also returns a response-level selection record describing whether the target came
from an explicit selector, an existing session, one local booted/bootable candidate, the one booted
simulator with the app installed, or one provider-owned candidate. Ambiguous requests fail with
structured retry selectors instead of silently retargeting.
To fetch provider-hosted video and log URLs after close, pass closed.provider?.providerSessionId to client.sessions.artifacts(). See the BrowserStack, AWS Device Farm, TestMu AI, and Limrun guides for provider-specific setup.
Web sessions
Pass platform: 'web' to drive a browser session with the same command methods. Web automation
requires Node 24+. No client method sets up the managed web backend: run agent-device web setup
before first use, in the same state directory the client uses. Run agent-device web doctor to check
backend health.
MCP tools use the same command contracts, so they can also target platform: 'web' after setup;
setup and doctor stay CLI-only. Web network inspection returns the standard network result shape
without request or response bodies.
Android ADB providers
Use agent-device/android-adb when your bridge owns Android device access and you want
agent-device's behavior for ADB operations. Executors receive the arguments that follow adb, so a
remote bridge can route the same argument arrays through an ADB tunnel, websocket API, or another
remote transport.
The helpers take an executor directly. Use captureAndroidLogcatWithAdb(executor, options?) for a
bounded logcat capture.
A provider can also expose reverse to own port reversal. Plain executors do not advertise reverse
support; call createAndroidPortReverseManager(providerOrExecutor)
only when the provider supports adb reverse argument semantics. The manager makes duplicate setup
idempotent for the same owner and rejects conflicting owners for the same local endpoint. For a
device that other adb clients also drive, pass an executor with { noRebind: true }: the manager
runs adb reverse --no-rebind and never replaces an existing device mapping, including one it
created. When adb reverse --list shows the mapping, the refusal fails with COMMAND_FAILED and
details.reason: 'android_port_reverse_rebind_refused'. Otherwise it fails as an ordinary adb error.
The device shell re-parses whatever follows shell or exec-out, so those commands are built for you:
every dynamic word is rendered for the quoting its transport applies before it reaches the device. adb
forwards words verbatim, so a word is single-quoted; hdc wraps each element it sends in double quotes,
where $, a backquote, and " stay live, so a word is escaped for that context instead. An array that
begins with shell or exec-out and did not come from those builders is refused with INVALID_ARGS and
details.reason: 'unguarded-device-shell-argv' instead of being dispatched. A bridge that composes its
own device commands calls runAdbShell(executor, words, options?) or
runAdbExecOut(executor, words, options?) from agent-device/android-adb, passing each value as its
own word; runAndroidShell(device, words, options?) and runAndroidExecOut(device, words, options?)
resolve the executor from a device instead.
Command methods
Use client.command.<method>() for command-level device actions. These calls go through the same daemon transport as the higher-level client methods, so they carry session metadata and tenant/run/lease fields, return normalized daemon errors, and handle remote artifacts.
Results are daemon-shaped objects with typed known fields and follow the CLI's command semantics.
A failed interaction rejects with the same error the CLI prints. Read error.details.dispatched before you retry; Retry after a failed command explains the values.
Every client call that dispatches to the daemon accepts signal?: AbortSignal to cancel that one call:
With the built-in transport, a signal that is already aborted rejects the call without sending anything (error.details.dispatched: 'no'). Aborting while the request is in flight closes that request's connection, the daemon marks the request canceled, and the promise rejects with the typed canceled-request error (error.details.reason: 'request_canceled', error.details.dispatched: 'unknown'). A custom transport receives the signal on its context and may cancel differently. The daemon and the session stay alive for other requests. An abort is not a timeout: it never triggers the runner cleanup or daemon reset that a timeout does. Cancellation covers only the daemon request, so a response-artifact download already underway is not stopped, and a canceled one-shot replay still runs the existing cleanup that can tear down a daemon this client started.
fold accepts either pose or keyframes. Keyframes use linear interpolation at roughly 60 updates per second; repeat an angle to hold it. Timestamps must start at zero and increase strictly, with 2–64 frames and a final timestamp no greater than 60,000ms. Angles must be finite and between 0° and 180°. The final timestamp bounds motion, excluding helper preparation and final hinge verification. A custom final angle is verified within 0.5°; interior angles must also settle. Cancellation stops the motion at its current angle. Re-snapshot afterwards, including after interrupted motion.
press, click, and longpress take readinessTimeoutMs. With it, the command waits up to that many milliseconds for a target that is not on screen yet, then performs the requested interaction. Without it, the command looks once and fails at once, which is the right choice for an agent that most often misses because the selector is wrong. Use it in scripted flows, where a step can run before the screen finishes rendering:
The wait is capped at 2 seconds and covers only a target that has not appeared. When the target is still missing after the wait, the error carries error.details.readiness with waitedMs, polls, and end (expired or stalled). A capture that shows an empty accessibility tree ends the wait at once with capture_sparse and readiness.end: sparse. When the command had to wait and then succeeded, the result carries data.readiness with polls and waitedMs. A command that found its target on the first look has no readiness field. A covered, off-screen, or ambiguous target fails at once, and a screen that stays unreadable for the whole wait fails with its own error; neither carries readiness. readinessTimeoutMs is not an MCP tool argument and has no CLI flag.
Vega OS client support is VVD-only and covers device discovery, app open/close, back, home, and tvRemote. Physical Fire TV, capture, selector, install, logging, and performance methods report unsupported for Vega targets.
Supported command methods:
waitalertappStatebackhomeorientationappSwitcheractionButtonfoldkeyboardclipboardtvRemotereactNativedoctorprepareviewport
rotate() is a deprecated alias for orientation(); use orientation().
Domain client methods:
client.devices.list(),capabilities(),boot(),shutdown()client.sessions.list(),stateDir(),close(),saveScript(),artifacts()client.apps.install(),reinstall(),installFromSource(),list(),open(),close(),push(),triggerEvent()client.materializations.release()client.leases.allocate(),heartbeat(),release(). PassretainOnClose: truetoallocate()when you release the lease yourself; sessionclosethen leaves it active untilrelease(), expiry, or daemon shutdown. The returned lease hasretainOnClose: trueonly when the daemon honored it. Keep heartbeating a retained lease between sessions: it expires after itsttlMslike any other lease.client.metro.prepare(),reload()client.capture.snapshot(),screenshot(),diff()client.interactions.click(),press(),longPress(),swipe(),pan(),drag(),fling(),swipeGesture(),focus(),type(),fill(),scroll(),pinch(),rotateGesture(),transformGesture(),get(),is(),find()client.replay.run()andclient.replay.test()client.batch.run()client.observability.perf(options),logs(),events(),network(), andaudio()client.debug.symbols()client.recording.record()andclient.recording.trace()client.settings.update()
client.devices.list() returns AgentDeviceDevice entries. Their optional model and osVersion fields describe the hardware and OS when discovery reports them; see Device discovery for the sources.
client.capture.snapshot() carries an optional viewport: { width, height } beside nodes: the box those rects are measured in, in the same coordinate space and orientation. Scale and clip against it instead of inferring the screen from the largest rect. When no box was measured, the field is absent rather than zero; see Coordinates and viewport for what each producer answers with.
client.observability.events({ cursor, limit }) reads the session event timeline as paged JSON entries. Use nextCursor from the previous page to continue from the daemon-owned events.ndjson file without replaying already uploaded/displayed events. Cursors are absolute and survive the file's size rotation; a cursor older than the retained window rejects with COMMAND_FAILED, details.reason: "EVENT_LOG_CURSOR_EXPIRED", and details.earliestCursor to resume from.
The timeline leaves out user-entered content such as typed text, selector values, and payloads; see Sessions for what it keeps.
client.observability.audio() mirrors audio probe start|status|stop. Use it to collect compact RMS/peak dBFS buckets while other session actions continue:
Pass platform: 'web', 'macos', 'ios' (simulators), or 'android' (emulators on macOS hosts). See Audio probes for what each platform samples, permissions, and unsupported targets.
client.observability.perf() requires an options object with an explicit area; a call without one fails with an error that names the replacement. The option shapes mirror the perf CLI forms:
- Frame health:
{ area: 'frames' } - Memory:
{ area: 'memory', action: 'sample' }, or{ area: 'memory', action: 'snapshot', kind: 'android-hprof' | 'memgraph', out }to write an artifact to disk - CPU profile:
{ area: 'cpu', subject: 'profile', action: 'start' | 'stop' | 'report', kind: 'xctrace' | 'simpleperf', out } - Trace:
{ area: 'trace', action: 'start' | 'stop', kind: 'perfetto' | 'xctrace', out }
To start an xctrace capture, also pass a template such as 'Time Profiler' or 'Animation Hitches'.
On Android, each successful perf({ area: 'frames' }) read and each open of an app resets the frame window, so call it, perform a transition or gesture, then call it again to inspect that window. See Performance snapshots for result fields and platform support.
client.recording.record({ action: 'start', path, quality: 'medium' }) starts a recording with medium output quality.
client.capture.screenshot({ path, scale: 0.3 }) captures a screenshot at 30% of its original width and height. scale accepts 0.01 through 1.
client.batch.run({ steps }) accepts structured steps:
{ command: 'open', input: { app: 'settings' } }. Step input uses the same fields as the
matching client command. Daemon-shaped positionals/flags steps are internal to the daemon and not
part of the client API.
Batch orchestration for custom transports
Use agent-device/batch when a bridge or in-process runner receives daemon-shaped requests but owns command dispatch itself. The helper applies the same validation, inherited flags, serial execution, partial results, and error envelopes as the daemon batch command.
Full example from examples/sdk/batch-orchestration.ts:
Android installFromSource()
On Android, a successful installFromSource() response includes the identity you need to relaunch the app:
packageNamelaunchTarget
If the daemon cannot determine the installed app's identity, the request fails instead of returning an empty success payload.
URL source rules
installFromSource() limits URL sources:
- Private and loopback hosts are blocked by default.
- URL sources from any public host may point directly to an installable, including a bare iOS
.ipa, or to a.zip,.tar,.tar.gz, or.tgzarchive containing exactly one. - For existing reachable artifact URLs, use
source: { kind: 'url', url: ... }. - For local artifacts, use
source: { kind: 'path', path: ... }or the CLIinstall/reinstallcommands. - When a remote daemon resolves CI artifacts server-side, pass a GitHub Actions artifact source:
Remote daemons may also support { kind: 'github-actions-artifact', owner, repo, artifactName } or { kind: 'github-actions-artifact', owner, repo, runId, artifactName }. The local client passes these payloads through unchanged; it does not authenticate with GitHub or download the artifact.
Android .apk and .aab URL sources resolve package identity from the downloaded install artifact. Archive URLs may contain one installable .apk, .aab, .ipa, or iOS .app, including inside nested archives.
Remote Metro helpers
Use agent-device/remote-config for profile loading and path resolution, agent-device/metro for Metro preparation, reload, and tunnel lifecycle, and agent-device/contracts when a server consumer needs daemon request or runtime contract types. For bridged remote Metro, proxyBaseUrl is the bridge origin and publicBaseUrl is optional; the bridge descriptor supplies cloud iOS wildcard HTTPS hints and Android runtime-route hints. reloadMetro() calls Metro's /reload endpoint, matching the terminal r reload path for connected React Native apps.
Selector helpers
Use agent-device/selectors to parse and match selector expressions in a remote daemon or bridge. The role= term matches the platform-neutral kind vocabulary that snapshot --json publishes (see Snapshots): pass the nodes as captured, and a node whose kind is text always matches role=text. Separately, pre-reconciliation leaf spellings (statictext, edittext, textview, …) still match during a deprecation window, each on the nodes that actually carried that class. A legacy spelling matches beside such a node's kind, not in place of it — role=webarea resolves a macOS-helper AXWebArea node whose kind is axwebarea, for example. Matching stays platform-aware because editability checks differ by backend.
listSelectorChainMatches(nodes, chain, options) returns every node the winning selector alternative matches, in snapshot order, plus that alternative and its index — the same first-match domain findSelectorChainMatch uses, without uniqueness refusal, so a runner applies its own strictness to the same nodes the CLI matched. resolveSelectorChain can name a later alternative: by default it refuses an ambiguous one and keeps walking, so the indices agree when it passes requireUnique: false, when the first matching alternative is unique, or when disambiguateAmbiguous: true resolves that alternative in place. It returns null when no alternative matches. options is { platform, requireRect? }; the matched SnapshotNode objects are the ones passed in.
