am-will/codex-skills

gemini-computer-use

Build and run Gemini 2.5 Computer Use browser-control agents with Playwright.

All-time #8846 First seen Jan 23, 2026
8-week activity · all time api

Installation

$ npx skills add am-will/codex-skills --skill gemini-computer-use

Summary

  • Build and run Gemini 2.5 Computer Use browser-control agents with Playwright.
  • Use when a user wants to automate web browser tasks via the Gemini Computer Use model, needs an agent loop (screenshot → function_call → action → function_response), or asks to integrate safety confirmation for risky UI actions.

Similar popular skills

Related neighbors and high-traction skills in the same topics — useful to compare before installing.

Also in this package

Other skills from am-will/codex-skills · top by installs.

npx skills add am-will/codex-skills

Browse all from am-will/codex-skills

More details

Agent compatibility

Declared targets from SKILL.md / docs. Unmarked agents are not listed — the skill may still install via the CLI.

Claude Code Not declared
Cursor Not declared
Codex Not declared
GitHub Copilot Not declared
Windsurf Not declared
Gemini CLI Not declared
Cline Not declared
OpenCode Not declared

Repository health

Stars 1.0K
Default branch main
Open issues 3
Status Active

Package contents

Files included with this skill beyond the listing page.

  • skill md SKILL.md 2,082 B
  • docs SUMMARY.md 1,988 B

History

  1. First seen on skills.sh
  2. First recorded snapshot · 1,200 installs

SKILL.md

Gemini Computer Use

Quick start

  1. Source the env file and set your API key:

``bash cp env.example env.sh $EDITOR env.sh source env.sh ``

  1. Create a virtual environment and install dependencies:

``bash python -m venv .venv source .venv/bin/activate pip install google-genai playwright playwright install chromium ``

  1. Run the agent script with a prompt:

``bash python scripts/computeruseagent.py \ --prompt "Find the latest blog post title on example.com" \ --start-url "https://example.com"; \ --turn-limit 6 ``

Browser selection

  • Default: Playwright's bundled Chromium (no env vars required).
  • Choose a channel (Chrome/Edge) with COMPUTERUSEBROWSER_CHANNEL.
  • Use a custom Chromium-based executable (e.g., Brave) with COMPUTERUSEBROWSER_EXECUTABLE.

If both are set, COMPUTERUSEBROWSER_EXECUTABLE takes precedence.

Core workflow (agent loop)

  1. Capture a screenshot and send the user goal + screenshot to the model.
  2. Parse function_call actions in the response.
  3. Execute each action in Playwright.
  4. If a safetydecision is requireconfirmation, prompt the user before executing.
  5. Send function_response objects containing the latest URL + screenshot.
  6. Repeat until the model returns only text (no actions) or you hit the turn limit.

Operational guidance

  • Run in a sandboxed browser profile or container.
  • Use --exclude to block risky actions you do not want the model to take.
  • Keep the viewport at 1440x900 unless you have a reason to change it.

Resources

  • Script: scripts/computeruseagent.py
  • Reference notes: references/google-computer-use.md
  • Env template: env.example