ToolsTray

Star a tool to keep it here.

Chat with PDF

Chat with a PDF in ordinary words. Answers cite the page and show a snapshot of the passage they came from.

Downloads the text embedding, language and text recognition models (~595 MB) the first time you run it, then works offline.

Drag & drop your file here, or

Any PDF — digital or scanned. Scans are read on-device with OCR first (a one-time ~30 MB model), so you can ask about those too. A small AI model also downloads once (~22 MB) on first use, then works offline — on any browser, phones included. You always get the matching passages and a snapshot of where each sits in the PDF; a written answer additionally needs a GPU-capable browser (Chrome, Edge, Safari 26, or a recent Firefox).

Getting an answer

  1. Drop the PDF in. The first visit fetches a 22 MB embedding model, cached from then on.
  2. Wait through indexing. The document is cut into 900-character chunks with a 150-character overlap, so a sentence sitting on a boundary survives in one piece, and every chunk is embedded. The status line names the page it is on.
  3. Ask the question the way you would ask a colleague. Press Ask, or press Enter.
  4. Read the answer, then read the cards beneath it. Each shows its page, the passage rendered as a picture of the real page, and the sentence that earned the match picked out in yellow.
  5. Missed the point? Ask again in the document's own vocabulary. Matching happens between your phrasing and the file's, so naming the term the document actually uses shifts the results far more than rewording the question politely.

About this tool

Ask a long PDF a short question and you usually end up scrolling. Chat with PDF turns that around. Type the question the way you would say it out loud, and the tool embeds every passage in the document, scores them against what you asked, and brings back up to three. Each one lands as a card carrying a Page badge, a Copy button, and a cropped picture of that exact spot on the page with the matching sentence highlighted in yellow. Where the browser can reach a graphics card, a small language model then reads those same passages and writes a short answer above them, finishing on a page citation like [p12].

That written answer is worth reading with one eye open. It comes from a 0.6-billion-parameter model, and it is told to work only from the passages it was handed, to treat anything inside them that reads like an instruction as data rather than a command, and to reply with “I couldn't find that in this PDF.” when they do not cover the question. Citations pointing at pages it was never shown get stripped before the text reaches you. It can still misread a clause, which is precisely why the evidence sits underneath it. Two models are involved: retrieval is a 22 MB embedding model that runs anywhere, phones included, and the answer writer is 543 MB and wants WebGPU. Both are fetched once and then work from cache on your own hardware. The PDF never makes the trip.

What a single question actually does

The question box opens pre-filled with the shape it expects: “e.g. What is the refund policy?”. Hand it a 40-page terms document, ask exactly that, and it embeds your question, scores it against every 900-character chunk of the file, then keeps the closest three, dropping any that overlap each other or score far below the best one.

The paragraph that appears above those cards is capped at 192 tokens, which is why it stays a sentence or two and ends with something like [p18]. If page 18 was not among the passages it was given, that citation is deleted before you see it. A citation on screen always points at a card sitting below it.

Good to know

  • It reads what it retrieves, never the whole file at once. A fact assembled from three sections fifty pages apart will not come back whole, and a question like “how many times does this contract mention arbitration” has no chance at all, because counting is a different job from similarity search. Long documents stop at the first 80 pages or 600 chunks, whichever arrives first, and a line above the question box tells you where it cut.
  • The written answer needs WebGPU with 16-bit float support, which in practice means Chrome, Edge, Safari 26 or a recent Firefox on hardware that exposes shader-f16, plus a 543 MB one-time download. Everywhere else the tool drops back to ranked passages and their page snapshots, which is still most of the value. Scans bring in a third model: a file with no text layer anywhere goes through a 30 MB OCR pass before indexing, and faint or handwritten pages come back patchy.

Frequently asked questions

Why am I getting passages but no written answer?

The paragraph at the top is produced by a language model that has to run on your graphics card, and the browser has to hand it one through WebGPU with 16-bit float support. Chrome, Edge, Safari 26 and recent Firefox builds generally manage it; older browsers and plenty of phones do not. When yours cannot, the tool skips that step and goes straight to ranked passages, which work everywhere.

How much should I trust the paragraph at the top?

Treat it as a pointer at the evidence rather than the evidence itself. A 0.6-billion-parameter model is small enough to misread a clause, and it only ever sees three passages of the document. What it cannot do is cite a page it was not given, because those citations are removed on the way out. The passages it used are printed underneath with their page numbers for the times it does get something wrong.

Can I ask questions about a scanned PDF?

Yes. When no page in the file carries a text layer, a 30 MB OCR model reads the pages first and the text it recovers is what gets indexed. Quality tracks the scan: clean printed pages come back well, faded fax paper and handwriting come back thin. If OCR recovers nothing at all, it says so rather than pretending to have indexed something.

What happens if I load a 400-page PDF?

The first 80 pages are indexed and a line appears above the question box saying exactly that. A second ceiling sits at 600 chunks, which a densely typed file can hit before page 80. Anything past the cut has nothing to match against, so with a long document it is worth splitting out the part you care about and loading that instead.

It can't find something I know is in the document. What now?

Retrieval compares meaning, and it does so against 900-character chunks. Ask about termination when the contract only ever says cessation of tenancy and the match comes out weaker than you would expect. Try the document's own wording, or quote a phrase you remember reading. When nothing clears the relevance bar, the nearest passage is shown anyway with a note warning that it may not answer the question.

Missing something in Chat with PDF? Suggest a feature →

Tell someone who needs this.