Expose held pointer primitives in browser automation
Feature request
Please expose low-level held-pointer primitives for Codex browser automation, especially in the Codex desktop app's in-app browser.
Useful API shapes would be one of:
mouseDown(x, y),mouseMove(x, y),mouseUp(x, y)beginDrag(...),moveDrag(...),endDrag(...)dragHold(path)that leaves the pointer down so the agent can inspect/screenshot before releasing
Why this matters
Many interactive browser experiences have important transient state while the pointer is held down: drag-and-drop editors, drawing/canvas tools, WebGL apps, maps, games, timeline editors, design tools, sliders, scrubbers, resize handles, and sortable lists.
A concrete example is a Unity WebGL game where dragging a card from the hand into the world shows a real-time placement preview: range circles, blocked placement text, interaction highlights, and move score deltas. The agent needs to inspect that held state before deciding where to release.
The current in-app browser CUA surface appears to provide high-level input methods such as click, double_click, drag, move, scroll, keypress, and type. drag works for all-in-one interactions, but it is atomic: press, move, release. The agent only gets to observe after release, which misses the held-pointer UI state.
Expected behavior
Codex should be able to perform a general held-pointer workflow:
- press down at a chosen screen coordinate or target element
- move while keeping the pointer/button held down
- take a screenshot or inspect visible state before release
- continue moving while still held, if needed
- release at the chosen coordinate
This should work for canvas/WebGL content as well as ordinary DOM-backed drag interactions.
Impact
This would make Codex much better at manually QA-ing browser games, WebGL tools, drawing/canvas apps, drag-and-drop editors, design tools, and any interface where hover/held pointer state is the behavior under test.