Skip to content
You're viewing the v2 docs. Looking for v1?Go to v1 docs
Docs
Testsigma
Popular questions
↑↓ to navigate↵ to selectesc to close
Book a demo

Addon SDK reference

The second argument to execute() is the execution context. Which fields it carries depends on the platforms declared in applicationType and the permissions requested. Requesting nothing still gives you logging, a scratch directory, runtime variables, and an abort signal.

This reference covers the Modern (TypeScript) SDK. Classic addons use Java and Selenium directly.

FieldAvailablePurpose
ctx.ui.browserPlatforms include WEB or MOBILE_WEBDrive the page under test
ctx.ui.mobilePlatforms include ANDROID or IOSDrive the app under test
ctx.runtimeAlwaysRead and write runtime variables by name
ctx.loggerAlwaysStructured logging: debug, info, warn, error
ctx.tmpDirAlwaysPer-step scratch directory, auto-cleaned, always readable and writable
ctx.signalAlwaysAbortSignal that fires if the run is cancelled
ctx.executionAlwaysRead-only metadata: sessionId, applicationType, result identifiers
ctx.aiWith permissions.aiPlatform-brokered LLM generation
ctx.ocrWith permissions.ocrPlatform-brokered OCR and image finding
ctx.spawnWith permissions.spawnHost-brokered subprocess, local agent only

Your addon runs in a worker while the browser lives in the host process. That has 3 consequences, and each one produces a bug that looks like something else.

  • Await everything, including calls that are synchronous in Playwright, such as page.url(). Composing without awaiting is still free, so page.getByRole("form").locator("input") reaches the host only once, when you finally await it
  • Await every assertion: generic matchers return a promise here. An un-awaited failing assertion becomes an unhandled rejection, and the step passes when it should not
  • No closures: function arguments are rejected. Anything evaluating in the browser or on the device takes a string instead, including evaluate("el => el.textContent"), waitForFunction, $$eval, and WebdriverIO’s execute
MethodNotes
locate(elementRef | selector)Returns a locator handle for a declared element or an ad-hoc selector
goto(url, opts?)waitUntil: "load" | "domcontentloaded" | "networkidle"
screenshot(opts?){ fullPage?: boolean }, returns a Buffer
title() and url()Current page title and URL
evaluate(scriptString, args?)Has to be a string. Function values cannot be transferred
waitForSelector(sel, opts?){ state?: "attached" | "visible" | "hidden", timeout? }

A locator handle supports click(), fill(text), textContent(), getAttribute(name), isVisible(), count(), boundingBox(), and string-based evaluate().

locate() accepts a declared ElementRef or an ad-hoc { strategy, value } literal. The strategies are accessibilityId, id, xpath, className, androidUiAutomator for Android, and iosPredicate or iosClassChain for iOS.

SurfaceMethods
Element handletap(), longPress({ms?}), fill(text), clear(), text(), getAttribute(name), isVisible(), boundingBox(), swipe(direction), count()
Devicescreenshot(), getContexts(), getCurrentContext(), switchContext(name), getOrientation(), setOrientation(...), hideKeyboard()

Prefer ctx.ui.mobile.locate(...) over driver.$(...) where it suffices. locate() applies UiAutomator resourceId wrapping, the iOS class-name predicate workaround, and xpath normalization that the raw driver does not.

WebdriverIO’s waitUntil is unavailable, because its condition callback runs in Node and no string form can replace it. Use ctx.ui.mobile.waitUntil, which re-evaluates a driver expression on the device each interval.

await ctx.ui.mobile.waitUntil(driver.$("~ok").isDisplayed(), { timeout: 10_000 });

Pass the expression without awaiting it. Awaiting first would check it once, and is refused with an explanation. The options are WebdriverIO’s timeout, interval, and timeoutMsg.

ctx.ui.browser and ctx.ui.browser.page are the same object at 2 levels of abstraction. Where the adapter does not reach far enough, every session also exposes the real automation object and the real assertion library.

PlatformAutomationAssertions
Webctx.ui.browser.page, a Playwright Pagectx.ui.browser.expect
Mobilectx.ui.mobile.driver, a WebdriverIO driverctx.ui.mobile.expect

No permission declaration is needed for any of them. These are the genuine libraries rather than a curated subset, so every matcher, .not, .soft, and the generic value matchers work as documented.

Reach for them for frames, keyboard and mouse, dialogs, downloads, and accessibility snapshots on web, or raw Appium commands on mobile.

Asserting the shape of a form or menu in one call beats a pile of per-element checks. Matching is a subset check, so unrelated markup can change freely, and it auto-retries like any Playwright assertion.

const page = ctx.ui.browser.page;
const expect = ctx.ui.browser.expect;
await expect(page.getByRole("form", { name: "Filters" })).toMatchAriaSnapshot(`
- heading "Filters"
- textbox "Search"
- button "Apply"
`);
const snapshot = await page.getByRole("form", { name: "Filters" }).ariaSnapshot();

Your addon is one step inside someone else’s session, so anything that outlives the step or hands over credentials is refused with a clear error.

GroupMembers
Persistent callbackson, once, route, exposeFunction, addInitScript
Session lifecycleclose, pause, setDefaultTimeout, deleteSession, reloadSession, addCommand
Escapes that outlive the stepnewPage, newCDPSession, tracing, browser
Credentials, mobileoptions, requestedCapabilities

What remains available is worth knowing. One-shot waitForEvent(...) covers dialogs, downloads, and popups. capabilities stays available, so you can branch on platformVersion. And page.context() and page.request both work, so cookies, storageState(), and API calls are all fine. Only the dangerous members inside them are refused.

const v = await ctx.runtime.get("orderId"); // undefined when unset
await ctx.runtime.set("orderId", "ORD-1042");
await ctx.runtime.set("token", jwt, { isEncrypted: true });
const answer = await ctx.ai.invoke("Summarize this receipt", {
files: ["receipt.png"],
});

File paths are relative to ctx.tmpDir, or within your fs allowlist.

MethodReturns
extractTextFromPage()Recognized text spans with bounding boxes for the current page
extractTextFromImage(path)The same, for an image file you supply
extractTextFromElement(elementRef)The same, for one element. Web only
findImage(refImage, opts?){ isFound, x1, y1, x2, y2 }

findImage accepts threshold, scale, and occurrence options.

Was this page helpful?