Browser agent: local models (Ollama + browser-use)

Paperful’s browser_agent lane and paperful recover drive browser-use against your session vault Chromium profile. The model must emit tool-shaped actions (click by element index, wait, navigate) — not chat about downloading.

There is no official browser-use benchmark for local Ollama tags. The guidance below is evidence-weighted from browser-use docs, GitHub issues, and community reports (2025–2026), folded for Paperful’s honesty rules and shipped wiring. See ROADMAP § browser-use.

How Paperful uses browser-use (not vanilla quickstart)

Topic

Paperful behaviour

LLM

[browser_agent].model if set, else [llm].model; Ollama → ChatOllama, LiteLLM → ChatLiteLLM. Optional fallback_model: one full retry on miss (not on captcha / bot wall). Tags with / (e.g. openai/gpt-4o-mini) use LiteLLM on the fallback attempt and need llm.allow_remote when provider = "ollama".

Vision

[browser_agent].use_vision = false by default — DOM / element index only. Set true only with a VL-capable tag (qwen2.5vl, etc.); doctor / recover preflight reject text-only tags when vision is on.

Browser

System Chrome (channel="chrome"), same profile as session login; uBlock + cookie extensions on

Success

Valid PDF on disk (min_pdf_bytes, stable size) — not agent done alone

Task

Stay on DOI/landing; no search engines; no purchase; stop on CAPTCHA / paywall / 403

Caps

[browser_agent].max_steps, max_wall_s; agent stopped early when PDF lands

Batch recover

paperful recover --from-last-run replays keys from state/last-run.json (--from-last-run-mode, --limit). Same runner as --item.

Learned playbooks

Agent PDFs append state/fetch-wins.jsonl with promotable click: / rewrite wins and optional steps trace → playbooks propose / promote.

Operator setup: LLM § recover, sessions, config § browser_agent. Install via uv sync --extra browser-agent or a heavy Compose image (Docker); headed vault login stays on the host.

Context, temperature, and Ollama tuning

browser-use sends large DOM snapshots each step. Ollama may default to a small context window on consumer GPUs (community reports cite ~4k under 24 GiB). Agent workloads often want 16k–32k+; 64k when VRAM allows. Paperful does not yet pass num_ctx into ChatOllama — raise context in Modelfile / Ollama env or track ROADMAP for a config knob.

Community configs often use temperature=0 for action selection. browser-use maintainers note that smaller local models struggle with tool-calling (discussion #4261).

Acceptance test (one page, your rights)

Before trusting a tag for a collection run:

  1. Pick one article URL you already may access (OA or campus session in the vault).

  2. paperful recover --item <KEY> --dry-run then run without dry-run (or recover --from-last-run --dry-run after a run that hit the agent).

  3. Pass = PDF under out/ with sane bytes — not a polite done in logs.

  4. Fail = invents a new host, opens a search engine, or claims success with no file (Playwright download caveat).

Paperful already enforces disk verification and aborts on search/support URLs; models that only succeed by hallucinating PDF URLs should be demoted regardless of size.

Deterministic paths beat the agent

Use HTTP / vault Playwright lanes first. Reserve browser_agent for soft UI (cookie banner → “Download PDF”) when metadata lanes and vault retry already failed. See architecture § sources and workflows.

LiteLLM / remote

With provider = "litellm", page text may leave the machine (orange disclaimer). Use for recovery only with allow_remote understood. Reject ollama/… ids on the LiteLLM path — use provider = "ollama" instead. See llm.md.

References