# The Verify agent

Understand how Verify plans, dispatches, runs, and evaluates browser journeys.

## Follow one run

The Verify agent is a focused browser worker for one selected scenario and persona. It observes the page, uses bounded browser actions to carry out the authored journey, and reports what happened. It does not edit your source code or run in a checkout.

1. **Plan:** Verify pins the ready environment revision, scenario-pack digest, policy, selected journeys, and runner image. Change runs use diff-aware selection; manual, PR, and scheduled runs select the full pack.
2. **Dispatch:** Composal creates attempts and admits them within the run's concurrency and organization limits. Shared setup, when declared in the pack, completes before dependent attempts.
3. **Execute:** Roder coordinates each hosted attempt and starts its pinned browser sandbox on **Blaxel**. The agent works through the scenario using browser observations and the permitted actions. Each attempt has a duration and tool budget.
4. **Evaluate:** Typed checks determine assertion outcomes. Verify aggregates the attempts, performs configured cleanup, and publishes the run result and evidence.

The model can help navigate a dynamic interface, but its explanation alone is not a passing assertion. A recorded pass requires the authored checks to succeed. Run pins let you relate evidence to the environment, scenario version, and runner that produced it.

## Launch without opening the dashboard

The CLI and Verify MCP can start a remote browser run directly. List or register an environment, import an immutable scenario pack, create an impact plan, and launch using the returned plan ID and digest. For a manual URL-only run, no repository or Change is required:

```sh
com verify sweeps plan --org <org> --environment-revision <revision-id> --scenario-pack <pack-id> --json
com verify sweeps run --org <org> --impact-plan <plan-id> --plan-digest <digest> --idempotency-key <stable-key> --json
```

Use `com verify sweeps get` or `watch` to follow the returned sweep ID. The hosted MCP offers `verify_post_sweeps_plan` and `verify_post_sweeps`; the local MCP offers `verify_sweeps` with `plan` and `run` actions. `com verify schedules` and `verify_schedules` in local MCP manage cron runs. Change and pull request plans add a repository. The [Verify run skill](https://composal.ai/skills/com-verify-run.md) has the complete tool workflow.

## Queued and concurrency

**Queued** means the run or attempt was accepted but has not started its browser session. Dispatch, shared setup, available Roder admission, and the requested parallelism can delay an attempt. A configured limit is a ceiling: a pack with one scenario starts one scenario agent even when the limit is 20.

Once admitted, the attempt moves to **Running** and its browser preview appears in **Live** when the runner supplies one. If a run remains queued unexpectedly, inspect its **Logs** and attempt state; a failed dispatch or unavailable runner needs investigation. A queued badge by itself does not establish that Blaxel is running or that the application failed.

## Personas and credentials

A pack can describe anonymous users, members, moderators, and administrators. Give each scenario the persona and start state it needs. For example, an administrator can create a category, one member can start a topic, another can reply, and a moderator can review a report. Stage the shared data between those roles with scenarios or shared setup; one browser attempt does not silently become several signed-in users.

Store test-account values in **Verify → Credentials** and refer to their logical keys from scenarios. Verify supplies only the requested values to an admitted attempt. Prefer synthetic accounts and data, especially when the target URL is production. An external URL registration describes the target; actions still have to fit the environment's policy and the approved test scope.

## Evidence and verdicts

Use **Live** to watch progress, **Replay** to revisit recorded browser steps, and **Results** to see the assertions. **Logs** and browser Console and Network events help distinguish an app error from a sign-in, selector, timing, or runner problem. **Artifacts** hold protected screenshots and recordings; evidence can include test data visible in the browser.

Treat a failed journey as a lead to investigate. An assertion failure with a captured app state supports a product finding; an agent or infrastructure failure may leave the product outcome unjudged. Use [Verify overview](/docs/verify) for the first-run path and [scenario and MCP reference](/docs/verification) for the detailed setup and recording controls.
