FileFlow Guides

PRACTICAL FILE GUIDE

Extract PDF pages, export images or recognize scanned text

A PDF can contain page graphics, selectable text, scanned images or a mixture of them. First decide whether you need a smaller PDF, pictures of its pages or editable plain text. These outputs preserve different parts of the document.

Updated October 8, 2026

Open the PDF pages and text tool

Select the pages you actually need

Open the PDF in FileFlow and wait for the preview editor. Mark the pages to keep, use the arrows to arrange them and rotate individual pages when needed. Only selected pages are included in the output.

For example, to share just pages 2 and 3 of a report, deselect the other pages and choose PDF output. Open the downloaded document to confirm the page selection, ordering and orientation before sending it.

PDF output versus image output

Choose PDF when the recipient needs a document containing the selected pages. Choose PNG or JPG output to render those pages as images; FileFlow packages the page images in a ZIP download.

Rendering is different from extracting the original embedded images. It creates a picture of the entire page at the selected resolution. Text and page graphics become pixels, so the resulting image does not retain PDF text selection, links or interactive forms.

DPI changes rendering work and image size

FileFlow offers 72, 150 and 300 DPI for page rendering. Higher requested DPI normally produces more pixels and can improve small text, but it also needs more memory and can create larger downloads. Start with 150 DPI and compare whether your destination needs more.

For PDF OCR, large pages may automatically render below the requested resolution to stay within pixel and memory budgets. A DPI choice is not a promise that every page will be rendered at that value. Smaller selections make large-document workflows easier to check.

Embedded text and OCR are different

TXT reads the text already embedded in a PDF. It works well when the PDF has a usable text layer, but reading order and layout can still differ from the visible page. It does not turn a photograph of words into text.

Choose TXT (OCR) for scanned pages that have no selectable text. Pick the language of the document rather than just the interface language. A mixed-language document can use the optional second recognition language. OCR output is plain text, not an editable copy of the original page layout.

Verify OCR rather than trusting it blindly

PDF OCR accepts no more than 20 selected pages per run, even if the file is small. Break longer documents into selections. Blurry scans, tiny characters, handwriting and complex tables can produce missing or incorrect text.

Review names, dates, reference numbers and amounts against the page image. A plausible word can still be wrong. If a page has a good embedded text layer, try TXT before OCR; recognizing an image of text is unnecessary work in that case.

Keep the original document

Treat page extraction and assembly as creation of a new document. Do not rely on it to preserve digital signatures, bookmarks, forms or every document-level property. Inspect the result in a PDF reader if those features matter.

FileFlow’s processing happens locally in the browser, but browser memory and PDF decode limits can reject complex files before the nominal page or byte limits. A smaller page selection may help; the original remains your reference copy.

Try it in FileFlow

  1. Choose a PDF and select the pages to keep in the preview editor.
  2. Choose PDF, PNG/JPG ZIP, TXT for embedded text, or TXT (OCR) for scans; choose DPI and OCR language when shown.
  3. Process and download; check page order, image clarity or recognized text.

Save pages 2 and 3 as a PDF, or choose TXT (OCR) for a scanned letter with no selectable text.

Limits and privacy

Up to 100 MiB per file, 20 files, 200 MiB per operation and 500 PDF pages, including the combined page count for merging. PDF OCR accepts at most 20 selected pages per run. Complex PDFs may hit memory or decoding limits earlier. Check the downloaded pages.

Processing runs in your browser. File contents, names and recognized text are not sent to FileFlow servers or analytics. Loading the site and OCR models still uses the network. Keep your original until you have checked the download.

This guide page has no file picker and does not process documents. The tool opens in a separate website document. Live advertising is currently disabled.

Frequently asked questions

Why is TXT empty, and when should I use OCR?

TXT reads embedded text and does not recognize a picture of text. Use TXT (OCR) for scans, choose the document language and select no more than 20 pages per run. OCR can lose layout and misread characters; higher requested DPI may be reduced to fit memory limits.

Are my files uploaded?

Processing runs in your browser. File contents, names and recognized text are not sent to FileFlow servers or analytics. Loading the site and OCR models still uses the network. Keep your original until you have checked the download.

Related reading