Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Computer-use agents: from screen observations to verified outcomes

By Anup Rai15 min readReviewed September 2026

A computer-use agent is a software system that uses a model to choose actions in a graphical interface, executes permitted actions through a controller, and observes the resulting state. Its observations may include screenshots, a browser's document structure, or an operating system's accessibility tree. Pixel-based control is one implementation, not the definition of every computer-use system.

Grounding means connecting an instruction such as “open the claim” to a specific visible element or coordinate. Execution means sending the input. Verification means checking that the intended claim actually opened. These are three different opportunities for error.

Example: a claims operator needs to enter an already reviewed claim into a permitted legacy portal. The agent identifies the correct fields, fills a draft, checks the values, and requests the required submission approval. A controller applies the allowed input. A saved claim number and a fresh read of its values establish success; the model saying “done” does not.

Product references were checked on September 24, 2026. Start with tool-agent architecture for the general model/tool loop and computer-use workflow design for a full operational case.

Choose the interface before choosing the model

Interface How it selects the target Good fit Main limitation
Supported business API Resource ID and documented operation Repeated structured reads/writes The needed operation or permission may be unavailable
Scripted browser automation Role, label, test ID, or other locator Known web workflows Changing semantics and incomplete page metadata
Model-directed semantic UI Model chooses from inspected page/accessibility elements Variable navigation with usable structure Ambiguous labels, incomplete accessibility, model mistakes
Model-directed visual UI Screenshot region and action coordinates Canvas, native dialogs, inaccessible controls Geometry, small text, stale views, visual ambiguity
Hybrid Selects the appropriate surface per step Mixed business workflows More adapters and evidence contracts to maintain

Prefer a supported API when it satisfies the actual requirement and permissions. A visual check may still be necessary to assess a rendered chart or exported document. A native application may expose useful accessibility or automation APIs; “desktop” does not mean “pixels only.”

Playwright supports role/label locators and resolves the current matching element when an action uses the locator. It is not restricted to fragile CSS selectors. Its locator guidance and actionability checks support robust scripts, although a successful click still cannot prove a correct business outcome.

Avoid the misleading comparison

Claim Better interview answer
“Visual agents need almost no maintenance.” They need model, prompt, environment, policy, and regression maintenance.
“Scripts are deterministic, so the workflow always succeeds.” The script can be predictable while the network, page, and account state vary.
“A visual agent works with any GUI.” Test application compatibility, permissions, visibility, input methods, and task complexity.
“Human-like clicks bypass access restrictions.” Automation must use permitted access; blocked login or verification calls for an authorized handoff.
“Computer use is always 100 times slower.” Compare measured end-to-end latency, repair, review, and maintenance on the same task.

Useful selection rule: choose the least ambiguous supported interface that can perform and verify the task. Use a visual agent when missing structure or variable interaction justifies its extra uncertainty and operating cost.

Current implementation options

Option Current documented surface Integration consequence
Claude computer use Desktop screenshot/input client toolset Your application provides execution; shell/editor access is optional
OpenAI computer use Code-driven UI or structured computer actions Current GPT-6 Astra guidance recommends code execution; the structured tool remains supported
Gemini computer use Browser, mobile, and desktop; current guide recommends Gemini 3.8 Flash Implement the selected environment's actions, coordinate mapping, and safety responses
Amazon Nova Act Browser workflows with API integration Useful for scoped enterprise form, extraction, and QA workflows
Microsoft UFO UFO² Windows GUI/API automation; UFO³ Galaxy coordinates devices Distinguish a single-desktop agent from multi-device orchestration

An end-user agent product, a model API, a framework, and a hosted execution service occupy different layers. Do not compare them as interchangeable model names. Benchmark your complete integration rather than treating a vendor's example as a production reliability estimate.

Build an observe–act–verify loop

Architecture / visual model
flowchart TD T[Authorized task and success criteria] --> E[Acquire environment and exclusive input lease] E --> O[Observe current application state] O --> M[Model proposes next bounded action] M --> P[Validate target, authority and remaining budget] P -->|Allowed| X[Execute through controller] P -->|Needs decision| H[Human review or handoff] X --> V[Observe and evaluate postcondition] V -->|Progress, work remains| O V -->|Outcome uncertain| R[Reconcile against application state] V -->|Verified goal| D[Record outcome and release environment] R -->|Resolved, more work| O R -->|Unresolved| H H --> D
Read diagram source
flowchart TD
    T[Authorized task and success criteria] --> E[Acquire environment and exclusive input lease]
    E --> O[Observe current application state]
    O --> M[Model proposes next bounded action]
    M --> P[Validate target, authority and remaining budget]
    P -->|Allowed| X[Execute through controller]
    P -->|Needs decision| H[Human review or handoff]
    X --> V[Observe and evaluate postcondition]
    V -->|Progress, work remains| O
    V -->|Outcome uncertain| R[Reconcile against application state]
    V -->|Verified goal| D[Record outcome and release environment]
    R -->|Resolved, more work| O
    R -->|Unresolved| H
    H --> D

The human branch records its actual result: completed, cancelled, blocked, or unresolved. It must not mark an unresolved task as successful merely because the loop stopped.

What each stage must establish

  1. Task admission: identify the account, allowed application, intended records, permitted data, and success criteria.
  2. Observation: capture the current window/tab identity, location, dimensions, and relevant screen or semantic state.
  3. Proposal: ask for a bounded next action with the evidence it relies on.
  4. Policy: validate the action's scope and any required approval outside the model.
  5. Execution: use a typed handler and the correct active environment; reject unsupported operations.
  6. Postcondition: confirm progress from a new observation, not from the action's acknowledgment alone.
  7. Recovery or completion: reconcile uncertain effects, preserve useful evidence, and release resources.

Short sequences such as focus → type → inspect can reduce inference calls. Keep actions in order when they depend on shared focus or state. Do not run clicks and typing concurrently on one desktop. Split a sequence before an externally consequential action or a point where the next action depends on an unseen result.

Keep provider protocol separate from application policy

For the Claude API, a current request fragment is:

{
  "model": "claude-opus-5-5",
  "max_tokens": 2048,
  "tools": [{ "type": "computer_toolset_20260801" }],
  "messages": [{ "role": "user", "content": "Inspect the open claim draft and report its claim number. Do not change it." }]
}

The toolset requires no beta header. Dispatch its member tool_use calls by name and toolset_name; return corresponding results with the original IDs and toolset name. Ordered batches require an outcome for every call, including calls skipped after a failure. The tool definition supplies no desktop implementation. Opus 5.5 uses this toolset on the Claude API/Google Cloud; Bedrock compatibility differs. See the versioned tool contract.

OpenAI's structured path returns computer_call with ordered actions, followed by a matching computer_call_output. A call marked completed has finished generation, not execution in your application. previous_response_id continues the conversation; it does not restore the browser, cookies, or process variables. See API state and execution.

Keep a provider adapter for these envelopes and a separate application controller for identity, bounds, approval, execution, and evidence. An SDK change should not silently change who can submit a claim.

Screen geometry: make the coordinate contract explicit

A screenshot may be a crop, a resized image, or a device-pixel representation of a display whose input system uses logical coordinates. Store the transform with the observation. Do not infer it later from a model response.

Architecture / visual model
flowchart LR R[Controller region<br/>origin 200,100<br/>size 960 by 600] -->|Resize for observation| I[Image 480 by 300] I -->|Model identifies point| P[Image point 240,150] P -->|Apply stored transform| C[Controller point 680,400] C --> V[Check focus and fresh target<br/>before sending input]
Read diagram source
flowchart LR
    R[Controller region<br/>origin 200,100<br/>size 960 by 600] -->|Resize for observation| I[Image 480 by 300]
    I -->|Model identifies point| P[Image point 240,150]
    P -->|Apply stored transform| C[Controller point 680,400]
    C --> V[Check focus and fresh target<br/>before sending input]

For this axis-aligned example:

controller_x = origin_x + floor(image_x × region_width / image_width)

Use the equivalent formula for y. Here, 200 + floor(240 × 960 / 480) = 680. Crop origin, browser viewport, scroll position, and device scale are different pieces of state. A full-page browser screenshot cannot be clicked using viewport coordinates without locating the relevant viewport position first.

Gemini's current UI actions use normalized coordinates from 0 to 999. Convert them according to that API's documented convention before using a controller that expects pixels. Do not silently send them through a screenshot-pixel adapter. See Gemini action definitions.

Executable exercise: reject a stale or out-of-bounds point

This original helper maps integer image pixels to an axis-aligned controller region. Observation versions and geometry come from the controller, not the model. It performs no input action and grants no permission.

def map_capture_point(x, y, *, image_width, image_height,
                      origin_x, origin_y, region_width, region_height,
                      observation_version, current_version):
    values = (x, y, image_width, image_height, origin_x, origin_y,
              region_width, region_height, observation_version, current_version)
    if any(type(value) is not int for value in values):
        raise ValueError("integer pixel geometry and versions required")
    if min(image_width, image_height, region_width, region_height) <= 0:
        raise ValueError("positive dimensions required")
    if observation_version < 0 or current_version < 0:
        raise ValueError("nonnegative versions required")
    if observation_version != current_version:
        raise ValueError("observe again before using this point")
    if not (0 <= x < image_width and 0 <= y < image_height):
        raise ValueError("point outside captured image")
    return (origin_x + x * region_width // image_width,
            origin_y + y * region_height // image_height)

A controller version can detect known invalidation, such as a resize, navigation, or intervening input. It cannot freeze an independently changing page. A popup can still appear after the check. Use short action sequences, current target checks, and postconditions; do not call this an atomic GUI transaction.

Separate the control plane from the desktop

Architecture / visual model
flowchart TD Q[Task queue] --> O[Orchestrator and policy] O --> L[Durable task, lease and operation records] O --> M[Approved model endpoint] O --> C[Restricted executor channel] subgraph Environment[Dedicated execution environment] C --> A[Browser or desktop input adapter] A --> U[Application and task-scoped session] U --> S[Screen or accessibility observation] end S --> O U -->|Permitted traffic| P[Authorized portal] O --> H[Human review and recovery] O --> R[Protected evidence storage]
Read diagram source
flowchart TD
    Q[Task queue] --> O[Orchestrator and policy]
    O --> L[Durable task, lease and operation records]
    O --> M[Approved model endpoint]
    O --> C[Restricted executor channel]
    subgraph Environment[Dedicated execution environment]
      C --> A[Browser or desktop input adapter]
      A --> U[Application and task-scoped session]
      U --> S[Screen or accessibility observation]
    end
    S --> O
    U -->|Permitted traffic| P[Authorized portal]
    O --> H[Human review and recovery]
    O --> R[Protected evidence storage]
Component Role Isolation/operating requirement
Virtual display/window system Makes a GUI available without a physical monitor Explicit display ID, size, focus, and lifecycle
Browser or native application Runs the target interface Dedicated profile, controlled downloads/extensions, scoped credentials
Input adapter Click/type/scroll/key operations Typed arguments, target checks, cancellation, no arbitrary extra commands
Observation adapter Captures permitted state Correct geometry, redaction, sensitive-data handling
Remote viewer Human inspection or takeover Authenticated access and exclusive input ownership
Orchestrator Task lifecycle and decisions Keep broad provider/control credentials out of the desktop

Xvfb provides an X11 framebuffer; a window manager arranges windows; an input utility injects events. None of those components is a sandbox. Linux containers share the host kernel; a VM adds a separate guest kernel. The required boundary depends on the code, applications, host resources, and threat model. See execution isolation.

A container image alone does not specify runtime security. Configure its user, mounts, privileges, network, device access, resource limits, and cleanup. Avoid privileged mode, host home-directory mounts, or a Docker socket merely to make a tutorial work. Use a tested environment image and startup/readiness checks instead of assuming a package name or fixed sleep starts a healthy desktop.

Managed environments can reduce provisioning work, but persistence is a separate decision. For example, E2B distinguishes running, paused, and killed sandboxes; paused state remains until explicit removal. “The task ended” therefore does not prove credential or filesystem cleanup. See E2B persistence.

Browser versus desktop

A browser context can isolate cookies and storage for a workflow, but it is not an operating-system security boundary. Browser-only execution often needs fewer components than a complete desktop. Native applications may require OS-specific privileges, installation, licensing, or interactive sessions. Neither choice guarantees better speed or reliability: measure the workload.

For a shared workstation, use explicit task scope and human control. For unattended business automation, a dedicated execution environment usually makes attribution, cleanup, concurrency, and recovery easier to reason about.

Recovery: ask what actually happened

Failure Immediate response Why blind retry is wrong
Stale screen or changed layout Observe again and locate the target The previous coordinate may now mean Delete
Wrong field received text Stop, inspect draft, correct only understood changes Repeated typing compounds the error
Unexpected permission/consent dialog Apply the actual permission policy or hand off A dialog is not always a nuisance to dismiss
Session expired/MFA needed Human or supported identity flow The agent cannot invent authentication authority
Submit clicked, response lost Mark outcome unknown and inspect existing records A second submit may create a duplicate
Worker lease expired Fence the old worker from further input Two workers may control one account/task
Repeated action with no progress Stop at a bounded threshold and diagnose Changing wording does not guarantee progress
Model/API failure Preserve environment state and apply bounded retry Restarting a model call need not restart the business operation

A lost submission response

Architecture / visual model
sequenceDiagram participant O as Orchestrator participant D as Durable ledger participant W as Desktop worker participant P as Legacy portal O->>D: Reserve task and record intended submission O->>W: Approved submit for exact draft W->>P: Click Submit once P->>P: Create claim Note over W,P: Connection is lost before confirmation is recorded O->>D: Mark submission outcome unknown O->>W: Open fresh lookup for source claim reference W->>P: Search existing claims P-->>W: Matching record or inconclusive result W-->>O: Record identity and observed fields O->>D: Resolve verified result or hold for human review
Read diagram source
sequenceDiagram
    participant O as Orchestrator
    participant D as Durable ledger
    participant W as Desktop worker
    participant P as Legacy portal
    O->>D: Reserve task and record intended submission
    O->>W: Approved submit for exact draft
    W->>P: Click Submit once
    P->>P: Create claim
    Note over W,P: Connection is lost before confirmation is recorded
    O->>D: Mark submission outcome unknown
    O->>W: Open fresh lookup for source claim reference
    W->>P: Search existing claims
    P-->>W: Matching record or inconclusive result
    W-->>O: Record identity and observed fields
    O->>D: Resolve verified result or hold for human review

A unique operation ID plus atomic dispatch-state checks can prevent your workers from repeatedly dispatching the same operation. The ID alone cannot create idempotency in a remote portal that has no such contract. Use a permitted unique external reference where available and verify its uniqueness semantics. If the existing record cannot be identified reliably, hold the case for human reconciliation. See durable execution.

Security and approvals are part of the workflow

  1. Identity: use a permitted account with the smallest useful business role; keep a record of whose authority applies.
  2. Data: treat typing, uploads, screenshots, clipboard contents, downloads, and notifications as possible data transfers.
  3. Scope: enforce application, account, destination, file, and operation restrictions in the controller and surrounding environment.
  4. Untrusted content: webpage text and images can contain prompt injection. They are observations, not new user instructions.
  5. Consequential actions: bind required approval to the actual record, values, recipient, operation, and freshness window.
  6. Takeover: pause agent input before giving a person the session; resume only after a fresh observation and explicit ownership transfer.
  7. Evidence and cleanup: protect logs/screenshots, set retention, revoke temporary access, and verify resource termination.

A second model reviewing the same screenshot can repeat the first model's mistake. Use deterministic checks where possible, an independent source when available, and human review for unresolved high-impact cases. A prompt-injection warning in the system prompt is one layer; it cannot substitute for scope enforcement. See agent security.

Estimate latency, capacity, and full cost

Latency is a measured sum

For an illustrative ten-turn workflow, assume each turn takes 0.12 seconds to capture, 0.08 seconds to prepare/transfer the observation, 2.5 seconds for inference, and 0.8 seconds for input plus application settling:

10 × (0.12 + 0.08 + 2.5 + 0.8) = 35 seconds.

This excludes startup, login, queueing, review, and repair. The number of model turns need not equal the number of UI actions. Track p50/p95 task time and timeout rate; an average alone hides slow applications and failed workflows.

Interview sizing: 500 claims per day

Assume an eight-hour processing window and a measured mean worker occupancy of two minutes per claim, including ordinary setup and recovery. Keep human waiting outside the active pool when a session can be safely checkpointed.

Quantity Calculation Result
Average arrivals 500 / 480 minutes 1.04 claims/minute
Average occupied workers 1.04 × 2 2.08
Slots at 70% planned occupancy ceil(2.08 / 0.70) 3
Illustrative 3× sustained peak ceil(6.25 / 0.70) 9
Ideal daily capacity of 20 continuously busy workers 20 × 480 / 2 4,800 claims

Three average-load slots and nine peak-load slots are planning estimates, not queue-delay guarantees. Check portal concurrency limits, provider quotas, CPU/memory, session restrictions, burst duration, and p95 occupancy. A permitted account that only allows one active session may be the true bottleneck.

Monthly economics

Assume 11,000 claims/month, $0.35 measured model/tool usage per claim, and $0.10 runtime/storage cost per claim. Suppose 20% need four minutes of review at $45/hour; maintenance takes eight hours at $100/hour.

Component Calculation Monthly cost
Model/tools 11,000 × $0.35 $3,850
Execution/storage 11,000 × $0.10 $1,100
Human review 2,200 × 4/60 × $45 $6,600
Maintenance 8 × $100 $800
Total Sum $12,350
Cost per attempted case $12,350 / 11,000 $1.12

At 40% review, total cost becomes $18,950, or $1.72 per attempted case. These are hypothetical rates, not provider prices. If only 95% of attempted cases become verified completions, the first scenario costs approximately $1.18 per verified completion, including the spending on unsuccessful cases.

Do not estimate screenshot billing from base64 file size. Image processing rules, tool definitions, accumulated context, caching, reasoning, and output affect billed tokens. Keep image history under the selected model's documented limits; do not remove prior context in a way that invalidates its continuation contract.

Optimize without hiding failures

  • Use API or semantic actions for stable, inspectable steps.
  • Batch short dependent UI actions in order, with observation checkpoints.
  • Crop/resize only when the target remains readable and the transform is preserved.
  • Reuse trusted environment templates while isolating task credentials and state.
  • Parallelize separate sessions; keep input ownership exclusive within a session.
  • Report cost per verified outcome and the rate of human rescue alongside throughput.

Evaluate the workflow, not only the next click

OSWorld 2.0 contains 108 long-horizon tasks and 31 self-hosted websites, with checkpoint scoring as well as complete-task outcomes. Its task distribution and horizons differ from the earlier benchmark; scores are not interchangeable. See the benchmark authors' project.

For your own release gate, include:

Test slice What it establishes
Normal workflows Correct final records and field values
UI variants Robustness to layout, labels, scaling, locale, and application updates
Failure injection Behavior after timeout, restart, stale screen, and expired login
Duplicate/unknown submission Reconciliation rather than repeated effects
Unauthorized or injected requests Scope enforcement independent of page content
Human takeover/cancellation No continued input from the displaced worker
Long workflows Budget handling and recovery across accumulated context

Measure complete-task success, false completion, duplicate effects, correction rate, human minutes, and p95 duration. Preserve a held-out task set and repeat trials because model behavior is variable. A demo that clicks correctly once is not a release criterion.

Interview walkthrough: an authorized claims-entry service

These requirements describe a hypothetical system, not a named customer deployment.

Functional requirements

  1. Intake: accept reviewed claim data and a permitted source reference.
  2. Validation: check required fields, types, totals, and source identity before opening the portal.
  3. Draft: populate the correct claim form in an isolated session.
  4. Review: show the exact pending submission and capture required approval.
  5. Submission: perform the authorized action and retain the resulting record identity.
  6. Reconciliation: investigate missing confirmations and suspected duplicates.
  7. Operations: support task status, cancellation, human takeover, and audit lookup.

Non-functional requirements

  1. Correctness: no success label without a verified record and checked critical fields.
  2. Authorization: prevent cross-customer/account access and unapproved actions.
  3. Capacity: process 500 claims/day within the agreed window and portal limits.
  4. Recovery: survive worker loss without blindly repeating a submission.
  5. Privacy: minimize screenshots and protect source documents, credentials, and evidence.
  6. Operability: bound each run and measure review workload and cost per completion.

Basic design: queue → browser worker → confirmation screenshot. It demonstrates the path but leaves extraction errors, approval, duplicates, and failure recovery unresolved.

Detailed design: validate reviewed input, create a durable task and exclusive lease, use a scoped account, fill a draft, compare critical values, obtain required approval, submit once, and reconcile the saved record. The control-plane diagram and submission sequence above implement this design.

Flaw in the basic design Revision Benefit Added cost
OCR error copied into a valid-looking form Field/source validation before entry Prevents confidently submitting bad input Validation and exception handling
Timeout triggers another submit Unknown state plus existing-record lookup Reduces duplicate business effects Ledger and reconciliation workflow
Reused browser leaks another client's data Separate scoped identity and session Better confidentiality and attribution Login/session provisioning overhead
Every unusual screen escalates Add tested semantic steps and bounded visual repair Reduces avoidable review More adapter and regression maintenance
Review is assumed free Model reviewer minutes and queue capacity Exposes the actual operating constraint Staffing and scheduling

Other valid applications include permitted legacy data entry, visual QA, document-layout inspection, and cross-application workflows. Each needs its own definition of success and authority; a claims design is not automatically suitable for publishing content or moving money.

Interview questions and answer notes

  1. Is computer use necessarily pixel-only? No. Screenshots, semantic browser structure, accessibility APIs, and ordinary tools can coexist.
  2. Why can a correct coordinate still click the wrong target? Focus, layout, overlays, scaling, or application state can change after capture.
  3. Does a model's completed tool call prove the UI action ran? No. Generation, execution, and business completion are different events.
  4. When is Playwright a better choice? When a permitted workflow has stable semantic controls and needs repeatable low-overhead execution.
  5. Can two agents speed up one form by typing in parallel? Shared focus and state create races. Parallelize independent sessions instead.
  6. Why not retry a timed-out Submit? The portal may already have committed; reconcile the existing record first.
  7. Is a second vision model enough to verify a claim? No. Errors can correlate; compare authoritative fields and identifiers where possible.
  8. What does a paused environment retain? It depends on the platform and snapshot mode; pause is not a deletion guarantee.
  9. What is missing from cost per model call? Other calls, execution, screenshots/storage, repair, review, maintenance, and failed outcomes.
  10. Why is nine workers only a starting estimate? The average occupancy calculation omits the full queue distribution and external bottlenecks.
  11. What makes an approval stale? Changed record, values, destination, scope, identity, or expiry; revalidate before execution.
  12. What proves the design is ready to expand? Measured complete-task quality, bounded harmful/duplicate effects, workable recovery, and acceptable total economics on representative held-out tasks.

Final summary and closing answer

Remember observe → ground → authorize → act → verify → reconcile. Prefer useful structure, preserve the coordinate contract, isolate execution, and make unknown outcomes visible.

A strong closing answer is: “I would use supported APIs and semantic controls where they work, and visual control for the remaining interface gaps. Each task gets scoped identity, exclusive input ownership, a budget, and a verified final state. Submission uncertainty goes to reconciliation. I would expand automation only after measuring complete-task correctness, human review time, and cost on representative failures as well as normal cases.”

Next: Building tool-use agents.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← OpenClaw: designing a persistent assistant around a trusted gateway
NEXT LESSONBuilding tool-use agents: contracts, execution, and evidence →

Explore the diagram