For AI agents and LLMs: a machine-readable index is available at llms.txt. A plain-Markdown version of any documentation page is available by appending .md to its URL.
Skip to main content

Screen Reader Automation: Executor Hooks

Real Device

Executor Hooks let your Appium script drive the screen reader on a TestMu AI real device. Through the lambda_executor interface you can turn TalkBack or VoiceOver on and off, move focus with the same gestures a screen reader user makes, and read back which element has focus and exactly what was spoken for it. Because every call returns data to your script, you can write assertions and turn screen reader validation into a CI check.

The hooks are language-agnostic. Anything that can call executeScript on an Appium session can use them.

When to use this​

Use Executor Hooks when you need a deterministic pass or fail on specific controls or flows: the primary button on every screen speaks a sensible name, a form field announces its label and error, a checkout flow can be completed with screen reader navigation alone, and the traversal order on a critical screen has not regressed. For broad, assertion-free coverage across the whole suite, use Auto Report. Both can run in the same session.

Prerequisites​

  • An Appium test project targeting TestMu AI real devices (Android 11 or later, iOS 15 or later).
  • LT_USERNAME / LT_ACCESS_KEY available to the process.
  • Screen Reader Automation enabled for your organization. See Screen Reader Automation (Overview).
  • The screen reader capability for your platform set on the session, as shown below. It makes sure a device that supports screen reader control is allocated.

Capabilities​

Set accessibility: true and the platform's screen reader capability inside LT:Options.

CapabilityPathTypePlatformDescription
accessibilityLT:Options.accessibilitybooleanBothMaster switch. Must be true for the executor hooks to be accepted.
talkBackLT:Options.accessibilityOptions.talkBackbooleanAndroidAllocates a device with TalkBack control and enables the TalkBack executor commands.
voiceOverLT:Options.accessibilityOptions.voiceOverbooleaniOSAllocates a device with VoiceOver control and enables the VoiceOver executor commands.
{
"LT:Options": {
"platformName": "Android",
"deviceName": "Pixel 8",
"platformVersion": "14",
"isRealMobile": true,
"app": "lt://APP1234567890",
"accessibility": true,
"accessibilityOptions": {
"talkBack": true
}
}
}

To also generate a Screen Reader Report in the same session, add the screenReader block from Auto Report next to these keys.

Executor contract​

Every hook is a lambda_executor call made through executeScript. The payload is a JSON object with an action and, where needed, an arguments object.

driver.execute_script('lambda_executor: {"action": "screenReaderGesture", "arguments": {"gesture": "navigate_next"}}')

The return value of executeScript is the command's response, so assign it to a variable when you want to assert on it.

Command reference​

PlatformActionArgumentsWhat it does
AndroidscreenReaderenable: "true" or "false"Turns TalkBack on or off.
AndroidscreenReaderGesturegesture: a name from Supported gesturesPerforms a TalkBack gesture and returns the element that received focus with its spoken output.
AndroidscreenReaderSpokenDescriptionresourceId: the element's resource IDReturns the spoken output for a specific element without moving focus.
iOSvoiceOverToggleenable: "true" or "false"Turns VoiceOver on or off.
iOSvoiceOverGesturegesture: a name from Supported gesturesPerforms a VoiceOver gesture and returns the element that received focus with its spoken output.
iOSgetVoiceOverElementnoneReturns the currently focused element and its spoken output without moving focus.

Gesture names, argument names and the response shape are the same on both platforms, so a test can target Android and iOS with the platform-specific action names swapped and nothing else changed.

Android (TalkBack)​

Enable or disable TalkBack​

Turn TalkBack on before any other TalkBack command. Turn it off when you are done with screen reader checks so the rest of the test runs without it.

// Enable TalkBack
driver.executeScript("lambda_executor: {\"action\": \"screenReader\", \"arguments\": {\"enable\": \"true\"}}");

// ... screen reader checks ...

// Disable TalkBack
driver.executeScript("lambda_executor: {\"action\": \"screenReader\", \"arguments\": {\"enable\": \"false\"}}");

Perform a gesture​

Each gesture call moves focus the way a swipe or tap would, waits for TalkBack to speak, and returns the newly focused element together with the text that was spoken.

Object result = driver.executeScript(
"lambda_executor: {\"action\": \"screenReaderGesture\", \"arguments\": {\"gesture\": \"navigate_next\"}}");

Get the spoken output for an element​

Use screenReaderSpokenDescription to read what TalkBack says for a specific element, identified by its resource ID, without moving focus. This is the simplest way to assert on a control's accessible name.

Object spoken = driver.executeScript(
"lambda_executor: {\"action\": \"screenReaderSpokenDescription\", \"arguments\": {\"resourceId\": \"com.example.app:id/sign_in\"}}");

The response lists the spoken output keyed both by resource ID and by the element's on-screen rectangle:

{
"spoken_description": {
"by_resource_id": {
"com.example.app:id/sign_in": ["Sign in, Button, Double tap to activate"]
},
"by_rect": {
"70 1171 1010 1297": ["Sign in, Button, Double tap to activate"]
}
}
}

The value is an array because a single element can be announced in more than one segment, for example label, role and hint.

iOS (VoiceOver)​

Enable or disable VoiceOver​

Turn VoiceOver on before any other VoiceOver command. The first enable in a session can take a few seconds longer than later calls.

// Enable VoiceOver
driver.executeScript("lambda_executor: {\"action\": \"voiceOverToggle\", \"arguments\": {\"enable\": \"true\"}}");

// ... screen reader checks ...

// Disable VoiceOver
driver.executeScript("lambda_executor: {\"action\": \"voiceOverToggle\", \"arguments\": {\"enable\": \"false\"}}");

Perform a gesture​

voiceOverGesture accepts every cross-platform gesture plus the iOS-only gestures listed under Supported gestures, and returns the newly focused element with its spoken output.

Object result = driver.executeScript(
"lambda_executor: {\"action\": \"voiceOverGesture\", \"arguments\": {\"gesture\": \"navigate_next\"}}");

Get the focused element​

getVoiceOverElement returns the element that currently has VoiceOver focus and the most recent announcement for it, without moving focus. Call it after a gesture, or after your own Appium interaction, to check what VoiceOver landed on.

Object focused = driver.executeScript("lambda_executor: {\"action\": \"getVoiceOverElement\"}");

Supported gestures​

Gesture names are case-sensitive. The cross-platform set works with both screenReaderGesture and voiceOverGesture; the iOS-only set works with voiceOverGesture.

GestureWhat it doesTalkBack equivalentVoiceOver equivalentPlatform
navigate_nextMove focus to the next elementSwipe rightSwipe rightAndroid, iOS
navigate_previousMove focus to the previous elementSwipe leftSwipe leftAndroid, iOS
activate_itemActivate the focused elementDouble tapDouble tapAndroid, iOS
scroll_downScroll the current container downTwo-finger swipe upThree-finger swipe upAndroid, iOS
scroll_upScroll the current container upTwo-finger swipe downThree-finger swipe downAndroid, iOS
backNavigate back or dismissBack gestureTwo-finger scrubAndroid, iOS
homeReturn to the home screenHome gestureHomeAndroid, iOS
navigate_firstMove focus to the first element on the screen–Four-finger tap, top of screeniOS
navigate_lastMove focus to the last element on the screen–Four-finger tap, bottom of screeniOS
read_from_topRead continuously from the top of the screen–Two-finger swipe upiOS
read_from_currentRead continuously from the focused element–Two-finger swipe downiOS
rotor_nextMove to the next item for the current rotor setting–Swipe downiOS
rotor_previousMove to the previous item for the current rotor setting–Swipe upiOS
pause_speechPause or resume speech–Two-finger tapiOS

read_from_top and read_from_current start continuous reading. Use pause_speech to stop it before the next gesture.

Response shape​

Gesture commands and getVoiceOverElement return an element_description for the element that has focus. The same field names are used on Android and iOS.

FieldDescription
resourceIdThe element's resource ID on Android, or its accessibility identifier on iOS.
classNameThe native class of the element, for example android.widget.Button or UIButton.
textThe element's visible text or accessibility label.
contentDescriptionThe content description on Android, or the accessibility hint on iOS.
boundsThe element's on-screen rectangle.
spokenOutputThe text the screen reader spoke when the element received focus.
propertiesAdditional accessibility state, for example whether the element is clickable, enabled, checked or focusable.

An illustrative response after navigate_next lands on a sign-in button:

{
"element_description": {
"resourceId": "com.example.app:id/sign_in",
"className": "android.widget.Button",
"text": "Sign in",
"contentDescription": "",
"bounds": "[70,1171][1010,1297]",
"spokenOutput": "Sign in, Button, Double tap to activate",
"properties": {
"clickable": true,
"enabled": true,
"focusable": true
}
}
}

A gesture that produces no focus change, for example navigate_next at the end of a list, returns success with the previously focused element rather than throwing.

Writing assertions​

Executor Hooks are most useful when the returned data feeds an assertion. Three checks cover most needs.

Spoken output for a control. Confirm the primary action announces a descriptive name and role.

driver.execute_script('lambda_executor: {"action": "screenReader", "arguments": {"enable": "true"}}')

out = driver.execute_script('lambda_executor: {"action": "screenReaderSpokenDescription", "arguments": {"resourceId": "com.example.app:id/sign_in"}}')
actual = out["spoken_description"]["by_resource_id"]["com.example.app:id/sign_in"][0]
expected = "Sign in, Button, Double tap to activate"

assert actual == expected, f"Expected '{expected}', got '{actual}'"

driver.execute_script('lambda_executor: {"action": "screenReader", "arguments": {"enable": "false"}}')

Focusability and traversal order. Walk the screen with navigate_next and compare the sequence of focused elements with the order you expect.

driver.execute_script('lambda_executor: {"action": "voiceOverToggle", "arguments": {"enable": "true"}}')

expected_order = ["email_field", "password_field", "sign_in_button"]
actual_order = []

for _ in expected_order:
result = driver.execute_script('lambda_executor: {"action": "voiceOverGesture", "arguments": {"gesture": "navigate_next"}}')
element = result["element_description"]
actual_order.append(element["resourceId"])
assert element["spokenOutput"].strip(), f"{element['resourceId']} was focused but nothing was spoken"

assert actual_order == expected_order, f"Traversal order changed: {actual_order}"

driver.execute_script('lambda_executor: {"action": "voiceOverToggle", "arguments": {"enable": "false"}}')

Completing a flow with the screen reader. Use navigate_next until the focused element is the control you want, then activate_item, and assert on the next screen. This proves the flow can be finished by a screen reader user, not just that each control has a label.

Execution rules​

  • Enable first. Every gesture and spoken-output command requires the screen reader to be on. Calling one while it is off returns: "Screen reader is not enabled for this session. Call the enable command first or set the screen reader capability."
  • Enable and disable are idempotent. Repeated calls in the same state return success and are logged as no-ops.
  • Commands are sequential. Each call waits for the gesture to settle and the speech to be captured before returning. Do not fire screen reader commands in parallel on one session.
  • Speech capture has a timeout. If no speech is captured within the gesture timeout, the command returns with the reason "No spoken output captured within gestureTimeout for gesture '<name>'." rather than a partial result.
  • Unknown locators are explicit. screenReaderSpokenDescription with a resource ID that matches nothing on the current screen returns "No element matched locator '<locator>' on the current screen."
  • Elements with no metadata are still reported. If the element exists but exposes no accessibility metadata, the response flags it so a missing label is debuggable instead of silent.
  • Platform mismatch is rejected. Android actions on an iOS device, or VoiceOver actions on an Android device, return an error instead of being ignored.
  • The screen reader is turned off at session end. You do not have to disable it yourself, but doing so keeps the rest of a long test faster.

Best practices​

  • Wait for a stable screen before a gesture. Spinners and animations change what receives focus. Put your explicit waits before the first screen reader command on each screen.
  • Assert on spokenOutput, not on text. The spoken output is what a user hears. It includes role, state and hint, which is exactly what static rules miss.
  • Prefer screenReaderSpokenDescription for single controls. It does not move focus, so it can be called in any order without affecting a traversal check.
  • Keep traversal checks short. Assert the order of the five or six controls that matter on a screen rather than walking every element. Long walks are what Auto Report is for.
  • Disable the screen reader in a finally block. If an assertion fails mid-check, the rest of the test still runs without TalkBack or VoiceOver slowing it down.
  • Use the same gesture names on both platforms. Write the check once and swap only the action names (screenReader / voiceOverToggle, screenReaderGesture / voiceOverGesture) per platform.

Combining with Auto Report​

Executor Hooks and Auto Report can be enabled in the same session. Set both the platform capability (talkBack or voiceOver) and the screenReader block with autoReport: true. Your executor calls control the screen reader as described here, and the report is still generated for every screen the test reaches. The gestures your script performs do not add extra entries to the report.

Troubleshooting​

SymptomWhat to check
Hook returns an entitlement errorScreen Reader Automation is not enabled for your organization. Contact support.
Hook returns "Screen reader is not enabled for this session."Call screenReader or voiceOverToggle with enable: "true" before other commands, and make sure accessibility: true and the platform capability are in LT:Options.
Gesture returns the same element every timeFocus is at the end of the list or inside a container that needs scroll_down first. Check spokenOutput to confirm where focus is.
Spoken output is empty for an element that has a labelThe element may not be focusable by the screen reader, so it is skipped. Check the properties in the response and the rule repository guidance for focusable containers.
iOS session cannot be allocatedThe voiceOver capability requires a device that supports VoiceOver control. Widen the device or OS selection, or ask support which devices in your pool support it.
Command fails with a platform mismatch errorThe action name belongs to the other platform. Use screenReader* actions on Android and voiceOver* / getVoiceOverElement on iOS.

Terminal First Testing With Kane CLI

Natural language browser & mobile app tests right from terminal.

×
Schedule Your Personal Demo
Kane CLI terminal

Help and Support

Related Articles