How I use it
I work across the complete agent loop: Playwright and Chrome CDP for control, DOM, screenshot and accessibility-tree context for perception, set-of-marks scanning for grounding, and ReAct-style planning for the next action.
Changes are evaluated on fixed episodes with paired seeds, exact McNemar tests, per-arm cost tracking and a pre-committed rule to revert anything without a positive pooled result.
Evidence in the work
- Moved a fixed 45-episode MiniWoB++ regression set from 33/45 to 43/45 through seven paired experiments.
- Kept three general-mechanism wins and cleanly reverted four candidates; every retained change received a regression test.
- Established a fixed-slice WebVoyager baseline and rejected all three architecture candidates when they missed the non-regression gate.

