How it works¶
A fill is Grab on the workbench, or run on the CLI. Paperful checks
the record, searches for a PDF, and writes the file to disk. It does not
merge duplicates and it does not rewrite titles or dates. Those jobs sit
around a fill — see Around a fill.
Click-by-click: First fill. The landing page is the short version of this story.
Trust the disk before notes or Ask: gaps → attachments (honesty) →
run (bytes on disk under out/) → you read out/ → then summarize /
ask. attach is the write gate into Zotero. Do not call that step
admit.
Claim |
Quote |
|---|---|
Dry-run shows the hit |
|
Gaps without fetch |
|
Contact without fetch |
|
A fill, in order¶
Pick the slice. A collection, or the whole library. Year and item-type filters are optional. Items that already have an imported PDF are skipped.
Check the record. Before any download, Paperful verifies an existing DOI against Crossref or OpenAlex, and can fill a missing one from a URL, a PubMed id, or a title match. That check is in memory for the search. The original identifier stays in the log. Zotero is not rewritten here.
Look for a free copy first. Open-access indexes, then the item’s own URL, then campus access when you have it. News and blog items can be printed to PDF. Odd hosts use playbooks you write, or recipes learned on this machine (sessions).
Use your library login when you have it. Log in once in Chrome or Edge.
runreuses that campus session through EZProxy. Same pattern for Google Scholar if you opt in. See Using your Scholar or library login.Only try what fits. Each item is sent only to sources that match what the record already knows — a DOI, an arXiv id, a publisher URL. If a site says slow down, Paperful waits. If a site keeps blocking, that source is paused. See Slowing down.
Save on this machine first. The PDF lands under
out/with a provenance stamp (paperful oa:unpaywall,campus:ezproxy, …) and a readable parent line (“Free copy from Unpaywall.”).Attach into Zotero 10+ when you ask. Older Zotero still gets the file on disk. Sparse one-page stubs are dropped; denser one-pagers wait for
attach --allow-short-pdf. Wrong-work PDFs. Not every paywalled or DOI-less item comes back.
Where it looks¶
Default order:
Unpaywall, OpenAlex, arXiv, bioRxiv/medRxiv, Europe PMC, Semantic Scholar, CORE, OpenAIRE
the item’s own URL
campus EZProxy
print-to-PDF for web items
Google Scholar stays off until you opt in; when on, it runs late (not
in the middle of campus/grey). [fetch].order = "list" keeps the sources
array order instead. Open-access sources can run in parallel. EZProxy,
print-to-PDF, and opt-in SerpApi stay one-at-a-time.
Per-item routing is on: a DOI paper is not sent to every site. The log line
trying: … is that shorter lane. Source routing.
Around a fill¶
Hygiene is a separate loop. Nothing is merged or rewritten until you say so.
Before a fetch, gaps counts missing PDFs. refs gap lists works cited
inside those PDFs that are not in the library (always dry-run). coverage
checks DOIs named in a briefing note or file against -C (same dry-run →
ingest-dois loop). run --dry-run prints a Would-hit column (sources in
order) and does not write the library.
Before a messy ingest, dedupe writes a review pack on disk. With
--apply it copies the extra parent’s PDF, notes, and better fields onto
the keeper, then moves that emptied parent to the trash. Dedupe.
After a fill, lint and fix-metadata propose identifier and title
patches on disk; --apply writes them to the catalogue. run never writes
bibliographic fields.
If you want notes, summarize writes a grounded note from the PDF text
(local model, off until you enable it). synthesize reviews those notes for
a collection. Scans need a text layer first. LLM.
paperful all is gaps → run → lint → fix → summarise. Dedupe is not in
that chain; run it before a messy collection. Workflows.
Related verbs (not inside a fill):
Job |
Where |
|---|---|
People you follow |
|
Saved keyword / seed profile |
|
Rank creators in |
|
Ask authors, do not fetch |
Slowing down¶
If a site returns a rate limit, the HTTP client waits and retries. If a site keeps showing a block page or a CAPTCHA (three times by default), Paperful pauses that source, continues with the rest of the lane, and tries one later item. A clean result opens the source again.
Items missed because a source was paused, or because the campus session
expired, are picked up on the next run. Items every applicable source
missed stay closed until you ask to retry them. When ezproxy is configured,
run probes the vault before batch 1. An expired session skips further proxy
wraps; on a terminal, run offers re-login at the next batch boundary and
again after the fetch for items left session expired (ezproxy_relogin,
default on; --no-ezproxy-relogin skips the pauses).
The knobs and error names: Source routing.
Using your Scholar or library login¶
The tool never asks for or stores your institutional password. You log in
once in a headed browser (paperful session login ezproxy or scholar on
the host). run reuses that login until the campus session expires —
typically hours to a few days. Being already logged in in your everyday
browser does not count — Paperful does not read that cookie jar (see
Sessions). Confirm the vault is still live with
paperful doctor --probe or paperful session status --probe (file presence
alone is not enough). When a publisher page is HTML, the same
browser follows the PDF link or download control before giving up. Scholar
fetches use the same browser profile, because Google often keys a CAPTCHA
to the browser, not just cookies.
Those logins live on this machine. Do not commit them or paste them into chat. Sessions · Campus EZProxy.
What you get back¶
A PDF on disk, when a copy could be reached, with a readable source line on the item. Optional: a grounded summary note, then a collection-level review of those notes. A folder copy you can back up or export as RIS, BibTeX, or EndNote XML — not a sync service.
Why this shape: Why Paperful. Commands: Commands.
Architecture of the same flow: Architecture. The same verbs
on localhost HTTP: Workbench. docker compose up serves it at
http://127.0.0.1:8765. Grab fetches to disk; Attach writes Zotero.
Walkthroughs: first fill, grow,
tidy.