Codex flagged that the previous catch-all skipped root-frame failures
too, masking a stale main-frame id as a successful (but empty) AX tree
and breaking interactiveness/form-value serialization that depend on
node.ax_node. Re-raise on the root frame so the retry/empty-DOM path
runs; continue to skip detached child iframes.
Codex review noted the original `previous_end - previous_start` formula,
while semantically misnamed, is load-bearing for replay correctness:
rerun_history uses it as a heuristic delay before replaying each step,
giving slow pages time to settle. Switching to the literal "gap" between
steps (~0 because step_start_time is set at the top of step()) silently
regresses replay on slow pages. #4484 needs a more careful redesign.
- StepMetadata.step_interval was computing the previous step's duration
instead of the gap between steps. Use current_start - previous_end so
rerun_history's "saved interval includes LLM time" comment actually
holds (closes#4484). Existing test encoded the bug as expected; updated
to assert gap semantics.
- _get_ax_tree_for_all_frames called asyncio.gather without
return_exceptions=True, so a single iframe detaching between
Page.getFrameTree and Accessibility.getFullAXTree (TOCTOU race common
with ad/widget iframes) discarded AX data from every other frame
including the main document, leaving the agent blind. Now skips dead
frames and merges surviving trees (closes#4778).
Three targeted fixes for pages with 5k+ elements:
1. build_snapshot_lookup: convert isClickable list to set before loop
_parse_rare_boolean_data used `index in list` (O(n) per call).
Called once per node = O(n²). Now O(1) via set lookup.
20k elements: 14,160ms → 2,973ms. 100k: 356s → 9s.
2. RectUnionPure: add _MAX_RECTS=5000 safety cap
Paint order rect union fragments exponentially with overlapping
layers. Uncapped, 20k elements took 372s. Capped: 4.7s.
Degrades gracefully (less aggressive filtering, same correctness).
3. Skip JS listener detection on pages with >10k elements
querySelectorAll('*') + per-element DOM.describeNode took 2.3s at 20k.
Elements still detected via accessibility tree + heuristics.
Combined effect at 20k elements: ~400s → ~16s (25x faster).
Normal pages (<5k elements) are completely unaffected.
_parse_rare_boolean_data used `index in list` (O(n) per call) on the
isClickable rare boolean data. Called once per node, this was O(n²)
total — the #1 bottleneck in the entire pipeline.
Fix: convert the list to a set once before the loop. O(1) per lookup.
Before → After:
5k elements: 1,788ms → 768ms (2.3x)
20k elements: 16,819ms → 2,911ms (5.8x)
100k elements: 356,629ms → 9,224ms (38.7x)
This single fix makes the full pipeline 2-3x faster at every scale:
5k: 8.0s → ~3.8s
20k: 29.7s → ~16s
100k: impossible → ~20s (Chrome-limited, not Python-limited)
Two additional performance fixes for heavy pages:
1. Hoist get_or_create_cdp_session() outside _construct_enhanced_node
Previously called once PER DOM NODE inside the recursive tree
construction. On a 100k-element page, this was 100k+ async
operations. Now resolved once before recursion starts.
2. Add _MAX_RECTS=5000 safety cap to RectUnionPure
The paint order rect union can fragment exponentially with many
overlapping translucent layers (each add() splits up to 4 rects).
Cap prevents memory/CPU explosion on complex pages.
Also: expanded stress test suite to 15 pages (up to 132k elements)
including shadow DOM + iframe combos, overlapping layers, cross-origin
iframes, and a 100k flat element test. All 15 pass.
Pages with very large DOMs (e.g. Stimulsoft designer with 20,000+
elements) cause the browser state capture to time out, making the
agent unable to interact with the browser.
Three targeted fixes:
1. Skip JS listener detection on heavy pages (>10k elements)
The querySelectorAll('*') + getEventListeners() loop followed by
individual DOM.describeNode CDP calls for each listener element is
O(n) and can take 10s+ alone on heavy pages.
2. Batch DOM.describeNode calls (chunks of 50)
Previously all calls fired at once via asyncio.gather, flooding the
CDP WebSocket and causing timeouts on concurrent operations.
3. Adaptive CDP timeouts based on page complexity
- >15k elements: 25s initial / 10s retry (was 10s/2s)
- >5k elements: 15s initial / 5s retry
- Normal pages: unchanged 10s/2s
Get pixel location of each iframe element that's out of the viewport, convert from pixels to number of page lengths, and provide that to the LLM as context
Elements in iframes beyond the 1000px viewport threshold are filtered from selector_map, causing agents to scroll blindly to find them. Add hints to the LLM context indicating hidden content
uses CDP Runtime.evaluate with includeCommandLineAPI to access getEventListeners() and mark elements with click/mousedown/pointerdown listeners before DOM snapshot capture
- Added 'row', 'cell', 'gridcell' to interactive_roles in ClickableElementDetector
- Added 'row', 'cell', 'gridcell' to interactive_ax_roles in ClickableElementDetector
- Verified logic with isolated test case
- Changed the type of the all_frames parameter to allow for None, enabling lazy fetching of cross-origin iframes only when necessary.
- Updated comments to clarify the behavior of all_frames during DOM tree construction.
- Improved the `create_task_with_error_handling` function to allow for better exception logging and retrieval based on the `suppress_exceptions` parameter.
- Updated `SessionManager` to implement event-driven recovery for stale agent focus, replacing polling with efficient event handling.
- Refactored session retrieval methods to ensure focus validation and recovery are handled automatically.
- Enhanced logging for recovery processes and session management to provide clearer insights into state changes and errors.
- Adjusted various watchdogs to utilize the new session management methods, ensuring they correctly handle focus validation and session retrieval.