Paperful
See how it works Walkthroughs Docs Install Cleaning Service GitHub
Contents Menu Expand Light mode Dark mode Auto light/dark, in light mode Auto light/dark, in dark mode Skip to content
Paperful
Light Logo Dark Logo

Start here

  • How it works
  • Walkthroughs
  • First fill: missing PDFs onto disk
  • Grow the library: topic search and people you follow
  • Tidy a collection: lint, identifiers, duplicates
  • Paperful workbench
  • Why Paperful
  • Docker
  • Zotero
  • Commands
  • Workflows
  • Research pack playbook
  • Dedupe
  • Configuration (config.toml)
  • Paperful architecture
  • Developer guide
  • Quiet mirror (out/)
  • Paperful terms
  • Releases and stability

Using paperful

  • Mendeley
  • EndNote
  • Source routing and circuit breaker
  • Campus EZProxy
  • Research operators
  • Browser sessions (Scholar, EZProxy, publishers)
  • Sci-Hub
  • SerpApi (Google Scholar search)
  • Optional LLM: setup and verbs
  • Ask your library: rag and ask
  • Browser agent: local models (Ollama + browser-use)

Product

  • How Paperful compares
  • How Paperful compares — reference
  • Snowball
  • Author watch lists
  • BBNJ author-site + authorwatch (dogfood)
  • E2E stack (topic + effort)
  • E2E all-in: NBA (reference dogfood)
  • Roadmap
Back to top
View this page

SerpApi (Google Scholar search)¶

SerpApi is a paid search API. Paperful uses only its Google Scholar engine: one HTTP call per remaining item, structured results instead of scraping scholar.google.com yourself.

It is link discovery, not a PDF host. When a result lists a PDF URL (author manuscript, institutional repository, preprint), Paperful downloads that file the same way it would after Unpaywall or a local Scholar hit. Provenance is web:serpapi. The download is still your httpx/browser path; SerpApi does not fetch the PDF.

Paperful never turns this on by itself. There is no silent cloud default.

When it can help¶

Local Scholar is opt-in and easy to block (CAPTCHA, 429, 503). SerpApi is a separate quota and a JSON API, so it can still return [PDF] links after those local lanes stall.

Typical useful cases:

  • Unpaywall / OpenAlex missed a copy that Google indexed (IR, author page, society preprint).

  • Title-only or weak-DOI records where Scholar’s query is a better lookup than a DOI resolver.

  • You already burned local Scholar or the vault browser and still have misses.

Each remaining item after the free/local serial tail can cost one search credit. Hits still have to download; a listed PDF that 403s is a miss.

When it will not help¶

  • Publisher paywalls. A Scholar snippet that points at the article landing page is not an OA PDF. Paperful does not treat SerpApi as a bypass.

  • Soft-blocked downloadpdf URLs. Those still need campus EZProxy, the vault browser, or handoff in your normal session.

  • Places Google does not index, or results that only show HTML.

  • Exhausted SerpApi plan, HTTP 429/402 from their API (Paperful latches the rest of the run as serpapi:skipped(quota)).

  • Expecting it to replace Unpaywall. Most OA PDFs already come from the parallel OA head.

Prefer campus access and OA indexes first. Prefer a logged-in Scholar handoff tab when you want to click through yourself ([handoff].scholar).

Setup¶

  1. Create a SerpApi account and copy the API key.

  2. Put it in a gitignored .env (see .env.example):

    SERPAPI_API_KEY=…
    

    paperful loads that file at startup. Do not put the key in config.toml or commit it.

  3. Opt in:

    [serpapi]
    enabled = true
    max_calls = 20
    

paperful doctor greens when the key is set and does not spend a search to prove it. Amber if enabled is true and the env var is empty.

Order in a run¶

With default [fetch].order = "policy", SerpApi runs in the late serial tail: after OA, campus, grey, local Scholar, and browser_agent, and before any later opt-in grey-zone source. [fetch].order = "list" only uses serpapi if you put it in sources yourself.

The query is the same shape as local Scholar (DOI, else a long title). --try-all still respects the paid cap below.

Caps (max_calls / --serpapi-max)¶

Paid searches are easy to burn on a large --retry-failed or library-wide run. Paperful counts actual API calls this process (not skipped inapplicable items).

Setting

Meaning

[serpapi].max_calls

Default 20. 0 = no cap for every run that uses this config

--serpapi-max N

Override for one run or all (also 0 = unlimited)

When the cap is reached, remaining items log serpapi:skipped(max_calls) and continue through later sources. The quota latch (skipped(quota)) is separate: that fires on 429/402 from SerpApi, not on your configured cap.

# stay inside a small monthly leftover
uv run paperful run -C Inbox --serpapi-max 5

# this process only; ignore the config default
uv run paperful run -C Inbox --serpapi-max 0

What Paperful does not do¶

  • Probe SerpApi on doctor or at startup (that would spend a credit).

  • Use other SerpApi engines (Google web, Google Maps, …).

  • Send your library through SerpApi unless [serpapi].enabled is true and SERPAPI_API_KEY is set.

Next
Optional LLM: setup and verbs
Previous
Sci-Hub
Copyright © 2026, Paperful contributors
Made with Sphinx and @pradyunsg's Furo
On this page
  • SerpApi (Google Scholar search)
    • When it can help
    • When it will not help
    • Setup
    • Order in a run
    • Caps (max_calls / --serpapi-max)
    • What Paperful does not do