Tools & Products

OpenAI Agents API with Native Computer Use: Headless Browser Sessions, Sandbox Runtimes, and OSWorld 2.0 Benchmarks

OpenAI released the Agents API with native Computer Use at DevDay 2026. The new tool allows models to capture screen buffers, move cursors, click UI elements, and execute bash scripts inside isolated virtual displays. Achieving 64.2% on OSWorld 2.0, the API provides developers with production infrastructure for end-to-end desktop and browser automation.

By FreakVinci · 2026-09-29 · 17 min read

Expanding Beyond Text: Native OS and Browser Actuation

At DevDay 2026, OpenAI launched the Agents API with native Computer Use capabilities. The system moves beyond traditional structured JSON function calling, equipping models with direct graphical user interface (GUI) actuation primitives: mouse clicks, cursor dragging, keyboard input, screenshot captures, and shell command execution.

While Anthropic pioneered the computer use API concept in late 2024, OpenAI implementation bundles managed cloud browser sessions and sandboxed desktop containers directly into the OpenAI developer dashboard, eliminating the need for self-hosted VNC infrastructure.


Execution Lifecycle and Coordinate Translation

The Agents API handles the visual-action loop through a structured state machine:

Computer Use Execution Loop
┌─────────────────────────────────────────────────────────────┐
│ 1. Capture Virtual Framebuffer (PNG Frame at 1024x768)      │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ 2. Visual Reasoning & Element Localization (GPT-6.1 Sol)    │
│ Identifies target DOM buttons, form fields, and dropdowns   │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ 3. Tool Call Emission                                       │
│ Type: 'computer_2026'                                       │
│ Action: 'click', Coordinate: [450, 312]                     │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ 4. Sandboxed OS Actuation (Wayland / Chromium Container)    │
│ Translates canonical coordinates to physical display pixel  │
└──────────────────────┬──────────────────────────────────────┘
                       ▼
┌─────────────────────────────────────────────────────────────┐
│ 5. Verification Screenshot & Action Loop Completion         │
└─────────────────────────────────────────────────────────────┘
  1. Resolution Normalization: The runtime scales incoming framebuffers to a canonical 1024x768 coordinate space. This reduces vision token consumption by 45% while preserving button click accuracy.
  2. Action Primitive Palette: The API defines 7 core primitives: mouse_move, left_click, right_click, double_click, mouse_drag, type_keys, and hotkey.
  3. Sub-Pixel Anchor Smoothing: To avoid synthetic bot detection, the runner applies natural Bezier-curve mouse trajectories rather than instant coordinate leaps.

Empirical Benchmark: OSWorld 2.0 Results

The OSWorld 2.0 evaluation suite assesses an agent ability to complete multi-step workflows across native desktop tools (LibreOffice, Chrome, VS Code, Thunderbird, and VLC).

Frontier Model OSWorld 2.0 Success Rate Mean Execution Steps Error Recovery Rate Cost Per Task
GPT-6 Astra 65.8% 14.2 steps 78.4% $0.48
GPT-6.1 Sol 64.2% 12.6 steps 76.2% $0.11
Claude Sonnet 5.5 60.9% 15.8 steps 71.0% $0.22
Claude Opus 5.5 58.1% 17.4 steps 68.5% $0.54
DeepSeek-V4.1 Flash 44.8% 21.2 steps 52.3% $0.04

GPT-6.1 Sol achieves 64.2% success on OSWorld 2.0 at an average cost of $0.11 per task—delivering higher task completion than Claude Opus 5.5 at one-fifth the operational expense.


Developer Integration Code

Developers configure computer use through the tools parameter in the OpenAI Agents SDK:

import OpenAI from 'openai';

const openai = new OpenAI();

async function launchBrowserAutomationTask() {
  const session = await openai.agents.sessions.create({
    model: 'gpt-6.1-sol',
    tools: [
      {
        type: 'computer_2026',
        display_width_px: 1024,
        display_height_px: 768,
        environment: 'managed_chromium_sandbox'
      }
    ],
    instructions: 'Navigate to target SaaS portal, export the quarterly CSV billing audit, and format findings into a summary.'
  });

  const run = await openai.agents.runs.create(session.id, {
    task: 'Log into billing.example.com using vault credentials and download the invoice report for Q3 2026.'
  });

  for await (const step of openai.agents.runs.streamSteps(session.id, run.id)) {
    if (step.type === 'tool_call' && step.name === 'computer_2026') {
      console.log(`Agent Action: ${step.args.action} at [${step.args.coordinate?.join(', ')}]`);
    }
  }
}

Security Architecture: Safeguarding Automated Sessions

Executing real operating system operations introduces distinct attack surfaces. OpenAI implements three containment layers:

  1. Ephemeral Ephemeral Containers: MicroVMs terminate immediately upon task completion, purging cookies, local storage, and cached file artifacts.
  2. Domain Whitelisting: Developers can restrict browser network egress to verified corporate domains, eliminating data exfiltration risks through malicious third-party links.
  3. Biometric Human Checkpoint: Attempts to enter payment details or access administrator console routes trigger an interactive push request to the developer, pausing script execution until authenticated.