# Chat with a Document — PDF, Word, Excel or Text, All In Your Browser

> Ask questions about a PDF, Word file, spreadsheet, CSV, Markdown or text file and get answers with the passage each came from.

Chat with Any Document reads a PDF, Word file, spreadsheet, CSV, Markdown, HTML or text file in your browser, splits it into passages, embeds them, and answers questions using the passages closest to what you asked — each one shown with the answer along with where it came from. Retrieval and answering both run on your own device using models that download once and are then cached, so no file is uploaded and no API key or account is required.

**URL:** https://convertto.tech/t/chat-with-document
**Category:** On-Device AI (https://convertto.tech/c/local-ai-tools)
**Privacy:** Runs entirely in the browser; no upload
**Cost:** Free, no sign-up
**Last updated:** 2026-08-01

## Key facts

- **Accepts:** PDF, DOCX, XLSX, XLS, CSV, TSV, TXT, Markdown, HTML, JSON
- **Citations:** Page, sheet and row, or heading, depending on the format
- **Uploads:** None — every stage runs in your browser
- **Cost:** Free, with no API key or account
- **Privacy:** Runs entirely in your browser — nothing is uploaded

## How to use

1. Select your any document file — the file stays on your device and is never uploaded.
2. Enter or paste your question.
3. Choose the answer model.
4. Press Run, then download the result when it is ready.

## FAQ

### Which formats can it read?

PDF, Word (.docx), Excel (.xlsx and .xls), CSV and TSV, Markdown, HTML, JSON and plain text. Legacy .doc files cannot be read in a browser — save them as .docx first. Each format keeps a locator suited to it, so a PDF answer cites a page, a spreadsheet answer cites a sheet and row, and a Word or Markdown answer cites the heading it sits under.

### Can I ask about several documents at once?

Not yet — one document at a time. Combining several would need the passages from each to be labelled by source, and a merged index makes the "answered from" line ambiguous. Merge them into a single PDF first if you need to.

### Is my document uploaded anywhere?

No. The file is read, split, indexed and answered entirely inside this browser tab. The only thing downloaded is the model itself, from Hugging Face, and that happens once and is then cached. Nothing about your document goes the other way — which is the point of using this rather than a service on a contract, a payslip or a medical letter.

### How does it answer questions about a document too long for the model to read?

The document is split into passages and each one is turned into a vector that captures its meaning. Your question is turned into a vector the same way, the closest few passages are found by comparing them, and only those passages are given to the answer model. This is the same retrieval-augmented approach the hosted services use; the difference is that here the retrieval and the answering both happen on your device.

### How accurate is it?

It is a model between 65 and 400 megabytes, which is one to four orders of magnitude smaller than a hosted assistant. It is good at pulling out a fact that is stated plainly in one place and much weaker at questions requiring several parts of the document to be combined, or at anything needing judgement. That is exactly why the passages behind every answer are shown with it — the answer is a shortcut to the right paragraph, not a substitute for reading it.

### What happens if the document does not contain the answer?

It says so rather than inventing one. Two separate checks make that possible: if no passage is even topically close to the question, the question is never put to the model at all, and when the extractive model is used its confidence is thresholded, because that model always returns its best guess and only its low confidence distinguishes a guess from an answer.

### Do I need an API key or an account?

Neither. There is no key to obtain, no sign-up, no quota and no per-question cost. The trade is that the models are small enough to run on your own hardware, and answer quality reflects that.

### Does it work offline?

After the first run, yes. The model is cached by the browser, so a second visit works with no network at all. The first visit has to download it.

## Sources

- [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks](https://arxiv.org/abs/2005.11401) — arXiv
- [transformers.js — running Hugging Face models in the browser](https://huggingface.co/docs/transformers.js) — Hugging Face

## Related tools

- [Chat with Excel & CSV](https://convertto.tech/t/chat-with-excel): Ask questions about a spreadsheet in plain English and get answers with the sheet and row they came from. Runs entirely in your browser.
- [Chat with PDF](https://convertto.tech/t/chat-with-pdf): Ask questions about any PDF and get answers with the page each one came from. The model runs in your browser — the file is never uploaded.
- [Chat with Word Document](https://convertto.tech/t/chat-with-word): Ask questions about a Word document and get answers with the heading each came from. Runs on your device — nothing is uploaded.
- [Semantic Search & Similarity](https://convertto.tech/t/semantic-search-tool): Embed a list of sentences with a real model and rank them by meaning — or find the near-duplicates in a list.
- [Text Chunker for RAG](https://convertto.tech/t/text-chunker-for-rag): Chunk by tokens, sentences, paragraphs, a separator ladder or markdown headings — with overlap, token counts and embedding cost.
- [Embedding Similarity Calculator](https://convertto.tech/t/embedding-similarity-calculator): Paste embedding vectors and get every pairwise cosine, dot product, Euclidean and Manhattan distance, ranked.
- [AI Alt Text Generator](https://convertto.tech/t/ai-alt-text-generator): Describe an image with a captioning model running in your browser — a first draft for the alt attribute you then edit.
- [AI Audio Transcription](https://convertto.tech/t/ai-audio-transcription): Run Whisper locally to turn speech into text, SRT, VTT or a timestamped transcript. Works offline after the first run.
