MMARW / INTELLIGENCE / AI
Desktop control vs in-page browser agents — instrument choice, demo failure modes, and what to prove before the room.AI-assisted publicationAI contributed to the research, drafting, or imagery. MMARW retains editorial responsibility for the published page.
AIIn the context of 2026 agentic workflows, a successful client demo is not defined by a "lucky click path"—a single, unrepeatable sequence where the model happens to navigate a specific UI correctly. A demo that survives professional scrutiny is one that demonstrates repeatable artifact production and robust failure handling.
Clients and stakeholders do not invest in stochastic magic; they invest in predictable utility. To move from a "wow" moment to a "work" moment, a demo must prove two things:
True survivability means the output of the demo is an artifact (a saved file, a database entry, a structured report with a verifiable URL) rather than just a chat transcript showing a series of completed tool calls. If the demo ends with a "Look how it moved the mouse," you have failed. If the demo ends with "Look at this processed CSV in the secure sandbox," you have succeeded.
Choosing the wrong instrument is the primary cause of demo-day latency and failure. Practitioners must distinguish between full desktop control and optimized in-page interaction.
Decision Rule: If the task requires moving data between a local desktop application and a web app, use Computer Use. If the task stays entirely within a browser and requires high precision on web elements, use a browser-agent approach. Note that these can be declared together in a single session; they use independent coordinate frames and are disambiguated by their toolset_name.
Building a demo requires managing the high-latency, iterative cycle of the agent. The interaction follows a strict loop governed by the Messages API:
tool_use block. In the computer toolset, this may contain multiple member tools (e.g., screenshot, left_click, type) sent in a single batch to improve efficiency. The toolset includes ~17 members, including zoom and scroll to manage visual context.tool_use block. It must execute these actions sequentially in the controlled environment (the sandbox).tool_result for each action, matched by a unique ID. This result includes the outcome (success or specific error).Critical Implementation Note: Because batch actions stop at the first failure, your application must be prepared to handle a halt error mid-turn. A demo that hangs during a batch failure is a dead demo.
To ensure an agentic demo does not collapse under the weight of real-world complexity, practitioners must apply strict scoping constraints.
Scope the demo around a single, complex decision point. Do not attempt to show an agent doing "everything." Show the agent making a meaningful choice (e.g., "Based on this email, should we file this as an invoice or a support ticket?").
Define exactly when the agent should stop. If the agent encounters an SSO wall or a CAPTCHA, the demo must have a programmed exit or a transition to a human. An agent spinning its wheels for 30 seconds looking for a button that isn't there is a demo killer.
For any irreversible action (deleting a file, sending an external email, moving money), the demo must include a "Human Approval" step. This demonstrates safety and control—two primary concerns for enterprise clients.
Never demo on a production machine. Use a sandboxed environment with allowlists. This protects you from prompt-injection (where untrusted content on a webpage could steer the agent to execute malicious commands) and ensures the environment is reproducible.
Before entering the room, every builder must verify the following:
screenshot → tool_use → tool_result accounted for in your presentation timing?halt errors mid-batch correctly without crashing the session?Ignore the marketing slides; focus on the bridge between reasoning and execution.
Turn the demo into a decision file — instrument choice, stop rules, human gates, and the artifact you will show. Start free on mmarw.com. Plans: Free / Starter $6 / Professional $24.
MMARW / INTELLIGENCE
This publication was formulated by MMARW. Try the workspace free for your own focused AI work.
| Element-based (selectors/DOM) |
| Best Use Case | Multi-app workflows (e.g., Excel → Slack → Local File) | Web-only tasks (e.g., CRM updates, Web Research) |
| Primary Failure | Coordinate drift / UI scaling issues | DOM changes / Shadow DOM complexity |
| Environment | Requires controlled desktop/VM | Requires browser instance/headless driver |
| Scaling | Heavy (Full OS overhead) | Light (Page-level focus) |
screenshot if necessary, and generates the next set of actions.Every publication is source-led, human-reviewed, and approved by MMARW before it goes live.
Create your own