paperclipai/paperclip

agent-browser

Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

80Collecting
See how to use itView GitHub source
npx skills add https://github.com/paperclipai/paperclip --skill "packages/skills-catalog/catalog/optional/browser/agent-browser"

Quick start

Start using it in three steps

Install it or open the source, trigger it with a clear task, then follow the source workflow.

1

Install the Skill

npx skills add https://github.com/paperclipai/paperclip --skill "packages/skills-catalog/catalog/optional/browser/agent-browser"
2

Describe the task

Use agent-browser to help me with: [describe your task]. Before you begin, tell me what input you need, the steps you will follow, and the expected output.

3

Follow the workflow

No structured workflow was detected; follow the original SKILL.md below.

Continue to the workflow

Direct answers

Answers to review before you install

What is agent-browser?

Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

Who should use agent-browser?

It is relevant to workflows involving Operations.

How do you install agent-browser?

SkillSignal detected this source-specific command: npx skills add https://github.com/paperclipai/paperclip --skill "packages/skills-catalog/catalog/optional/browser/agent-browser". Inspect the repository and command before running it.

Which Agent platforms does it support?

The upstream source does not declare a dedicated Agent platform.

What permissions or risks should you review?

No obvious permission action was detected by the static rules. This is not proof that the Skill is safe.

What are the current evidence limits?

This page combines upstream documentation with deterministic repository, quality, and static-risk signals. It is not described as a manual test or security review.

SkillSignal brief

Decide whether it fits your work first

Drive a real browser to inspect or interact with a web page or app — navigate, take screenshots, read console and network, fill simple forms — for verification tasks, not unattended automation.

Useful in these contexts

Not yet included in a workflow collection

Core capabilities

Operations

Distilled from the source

Understand this Skill in one minute

About 4 min · 10 sections

When it is worth using

  1. You need a screenshot of a deployed page or a local dev server to confirm a UI change.

  2. You need to read JavaScript-rendered content that curl/wget will not see.

  3. A user reports a UI bug and you need to reproduce it interactively to capture console errors, network requests, or layout state.

  4. You need to walk through a short flow (load page, click, observe) to verify acceptance criteria.

Limits and cautions

  1. The page is reachable as static HTML. Use curl/HTTP fetch — it is cheaper, faster, and more reliable.

  2. The task is unattended large-scale scraping. That belongs to a dedicated scraper with rate limits, robots.txt handling, and a real user agent policy — not this skill.

  3. The site is behind authentication you do not own credentials for, or whose terms of service prohibit automation.

  4. The site involves sensitive accounts (banking, healthcare, government) where automation risks lockout or compliance issues.

Repository stars
74,938
Repository forks
13,965
Quality
80/100
Source repository last pushed

Quality breakdown

Based on traceable docs and repository signals; stars are not treated as quality.

80/100
Documentation24/30
Specificity19/25
Maintenance20/20
Trust signals17/25

Compare before choosing

Related Agent Skills and source variants

These links are selected from shared tasks, functions, stacks, platforms, and same-name variants. Compare the source owner, documentation, permissions, and maintenance signals.

agent-browser by nexu-io

Browser automation CLI for AI agents. Use when the user needs to inspect, test, or automate browser behavior: navigating pages, filling forms, clicking buttons, taking screenshots, extracting page data, reading selected Open Design browser-tab context, testing web apps, dogfooding Open Design previews, QA, bug hunts, or reviewing app quality. Prefer local Open Design preview URLs unless the user explicitly asks for external browsing.

design-review by event4u-app

Use when the user says "review the design", "check the UI", or wants a comprehensive UI/UX review. Uses a 7-phase methodology covering interaction, responsiveness, accessibility, and more.

dask by k-dense-ai

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

medchem by k-dense-ai

Medicinal chemistry filters for compound triage. Apply drug-likeness rules (Lipinski, Veber, CNS), structural alert catalogs (PAINS, NIBR, ChEMBL), complexity metrics, and the medchem query language for library filtering.

neurokit2 by k-dense-ai

Use NeuroKit2 to build or audit reproducible research workflows for physiological time-series preprocessing, event/interval analysis, multimodal alignment, variability, and complexity. Trigger when code imports neurokit2 or needs its current APIs, schemas, and method-aware validation—not for diagnosis or device validation.

View original Skill.mdThis page is parsed directly from the repository SKILL.md without editorial rewriting. Collected: Jul 28, 2026 · about 4 min

Agent Browser

Use a controlled browser to verify behavior, capture evidence, or extract information from web pages that a static fetch cannot reach (SPAs, login-gated pages, dynamic content). This skill is about supervised verification, not unattended scraping.

When to use

  • You need a screenshot of a deployed page or a local dev server to confirm a UI change.
  • You need to read JavaScript-rendered content that curl/wget will not see.
  • A user reports a UI bug and you need to reproduce it interactively to capture console errors, network requests, or layout state.
  • You need to walk through a short flow (load page, click, observe) to verify acceptance criteria.

When not to use

  • The page is reachable as static HTML. Use curl/HTTP fetch — it is cheaper, faster, and more reliable.
  • The task is unattended large-scale scraping. That belongs to a dedicated scraper with rate limits, robots.txt handling, and a real user agent policy — not this skill.
  • The site is behind authentication you do not own credentials for, or whose terms of service prohibit automation.
  • The site involves sensitive accounts (banking, healthcare, government) where automation risks lockout or compliance issues.

Before launching the browser

  • Confirm the URL and what state should be true after navigation.
  • Decide what evidence is needed: full-page screenshot, viewport screenshot, console log, network trace, HTML snapshot, extracted text.
  • Decide the viewport size that matters for the task (mobile vs desktop). Default to a desktop size unless the task is mobile-specific.
  • For local dev servers, confirm the server is running and the port is what you expect.

Driving the browser

A typical verification session:

  1. Launch with a real-looking user agent when the target is the public internet; an unrealistic UA flags automation traffic.
  2. Set a sane viewport (e.g., 1366×768 desktop, 390×844 iPhone-ish).
  3. Navigate and wait for the right signal. Prefer waiting for a specific selector or network-idle over arbitrary sleeps.
  4. Capture evidence immediately after the wait condition succeeds, before any interaction perturbs the state.
  5. Interact deliberately. One click at a time, with a wait between actions; re-screenshot after each meaningful state change.
  6. Read the console and network panels for unexpected errors, 4xx/5xx responses, or slow requests.
  7. Close the browser cleanly when done. Long-running browser sessions leak memory and hold ports.

What evidence to record

For a verification task, deliver:

  • A full-page or viewport screenshot of each meaningful state.
  • The console log, filtered to warnings/errors.
  • Any non-2xx network response with the URL, status, and a short response body excerpt.
  • A short narration: "Navigated to X, observed Y, clicked Z, observed W."

For a UI bug repro, also record:

  • The exact reproduction steps the user can follow.
  • Viewport size and (where relevant) device pixel ratio.
  • Whether the bug reproduces on first load vs after interaction.

Login-gated pages

  • Prefer programmatic auth (API token, magic link) over UI login.
  • If UI login is the only path, the user must provide credentials explicitly for this run. Never reuse credentials outside the session.
  • Do not store credentials in the session log, screenshot, or returned output.

Performance and politeness

  • Throttle to one navigation per few seconds when touching shared infra.
  • Respect robots.txt for public sites you are inspecting at any volume.
  • Cancel navigations if a page exceeds a reasonable timeout (e.g., 30s); the page is broken or rate-limiting you.
  • Do not retry forever on failure. Retry once with a longer timeout, then escalate.

Common failure modes

  • Selector not found. Page changed, or you are waiting before render. Take a screenshot to see actual state; adjust the selector.
  • Click does nothing. The element is offscreen, covered by a modal, or in a shadow DOM. Scroll into view or pierce the shadow root.
  • Headless detection. Some sites detect headless Chrome and serve a different page. Use a non-headless mode or a fingerprint-realistic configuration only when authorized.
  • Cross-origin iframe blocking. Iframes you do not own cannot be inspected; the page must offer the data outside the iframe or the task is infeasible.

Anti-patterns

  • Long unsupervised browser sessions that drift from the original task.
  • Scraping behind authentication you do not own.
  • Captioning a screenshot with "looks good" without saying what state was loaded and what selectors confirmed it.
  • Treating a passing screenshot as proof of correctness across viewports you did not actually test.
Skill path
packages/skills-catalog/catalog/optional/browser/agent-browser/SKILL.md
Commit SHA
77979950381a
Repository license
MIT
Data collected