Build workflows
Reading web pages
Let an agent read and use a site as Markdown, following links and filling in forms by number, for a fraction of the tokens a browser snapshot costs.
web.* reads web pages as Markdown and acts on them by number. An agent opens a page, reads it, follows a link or fills in a form, and reads the page that comes back. It is the first thing to reach for when a workflow has to read or use a site: searching, working through a form, following a flow from page to page.
The web tools are available to every account with nothing to set up, under the web. prefix. They are meant for an agent's tools list, and they also work in tool steps.
How an agent sees the page
Every call answers with the page as the Markdown you would write by hand. Navigation, footers, cookie banners and anything the page hides are dropped. Links are numbered and form fields carry their current values:
# Sign up
[input: Email][f1]
[select: Plan = Starter (3 options)][f2]
[checkbox: Send me the newsletter][f3]
[button: Create account][f4]
Already have an account? [Sign in][1].
The next call names a number: web.fill with values: { "1": "ada@example.com", "2": "Team" } and submit: 4, or web.follow with link: 1. Numbers belong to the answer they came from, so an agent always acts on the latest one.
steps:
- key: price
kind: agent
model: balanced
tools: [ web.* ]
prompt: |
Open https://example.com, search for "{{ inputs.product }}" and report the first result's price.
Close the session when you are done.
Pages run their JavaScript in a browser engine written in Ruby, so single-page apps and Hotwire sites work, including what a page changes by itself between calls. web.read shows the page as it is now. A page longer than 40,000 characters comes in parts, and page: 2 reads the next one.
Web tools or browser tools
Browser sessions drive a real Chromium and answer with an accessibility snapshot of every element. That costs many more tokens per page. Use them for what web.* can't do:
- screenshots along the way, a video or a Playwright trace of the session
- signing in with a browser profile
- uploading a file
- a page that needs real layout, such as a canvas, a map or infinite scroll
For a page that only needs reading once, with nothing to click, builtin.fetch_url and builtin.extract are cheaper still.
What a page can reach
Pages load signed out, from the public internet, through the same gateway as screenshots and browser sessions. Addresses on a private network are refused, and so is every request a page makes to one. A workflow with egress: vpn or home sends its pages through that route.
A session keeps its cookies until it closes. An account can have 3 sessions open at once. A session closes when its run finishes, after ten idle minutes, or after 30 minutes in all, so have the agent close it when it is done.