Set a useful boundary for exploration.
Choose the target, allowed browser actions, synthetic inputs and assertions that define expected behaviour. Exploration works within that policy and the run’s execution limits. The report records approved assertion observations and action history, giving your team a basis for investigating what the agent encountered without treating every model suggestion as a confirmed defect.
Turn supported observations into reproducible findings.
Finding verification uses an approved assertion and retained native evidence. A fresh isolated attempt must reproduce the same failure before confirmation. Supported browser sequences can replay the recorded steps through the selected assertion. After review, supported findings can produce portable Playwright regression drafts that preserve the relevant observed behaviour.
- Inspect the actions and assertions behind an observation.
- Distinguish a suspected failure from a reproduced finding.
- Keep the original session while investigating a repair.
Choose an appropriate journey for an evaluation.
Start with a bounded journey and a clear expected result. Browser tooling, approved target access and a qualified runner determine available execution. Some observations cannot support deterministic replay, including unsupported actions or evidence that has been redacted beyond comparison. Those cases require review. Production actions and human accessibility assessments retain their separate approval requirements.