Extract PDF Links with JavaScript and Turn Them into a Download Batch
A documentation page has the PDFs you need for a project: specifications, reference sheets, and a few sample documents. You want a local copy of the relevant files before starting work. You could open each link and save
A documentation page has the PDFs you need for a project: specifications, reference sheets, and a few sample documents. You want a local copy of the relevant files before starting work.
You could open each link and save each file. You could also write a downloader. For a small, occasional collection, a useful middle ground is to extract the links, review the selection, and use a browser extension to manage the batch.
This walkthrough starts with a JavaScript snippet you can use independently, then shows how to take the resulting URLs through a download workflow.
Disclosure: we make Simple Mass Downloader, the desktop Chrome extension used in the second half of this example.
1. Extract direct PDF links from the page
On a page containing ordinary PDF links, open Chrome DevTools and select the Console. The following snippet reads anchor elements from the current document and prints a deduplicated URL list. It does not make network requests or start downloads.
(() => {
const urls = new Set();
for (const link of document.querySelectorAll("a[href]")) {
try {
const url = new URL(link.getAttribute("href"), document.baseURI);
if (!['http:', 'https:'].includes(url.protocol)) continue;
if (!/\.pdf$/i.test(url.pathname)) continue;
url.hash = "";
urls.add(url.href);
} catch {
// Skip links that cannot be parsed as URLs.
}
}
console.log([...urls].join("\n"));
})();
The URL constructor resolves relative links against the document's base URL. Checking pathname allows a link such as /docs/spec.pdf?version=2 to match, even though the full address does not end in .pdf. See the MDN references for URL resolution and pathname.
Removing the fragment combines links such as spec.pdf#page=2 and spec.pdf#page=8. Query parameters remain intact because they may select a version or carry information the server needs. The Set removes identical resulting URLs; it does not compare file contents.
Copy the printed list into a text file, with one URL per line. Review it before downloading.
Know what this snippet misses
This is a deliberately narrow extraction rule. It finds anchors whose URL paths end in .pdf.
It will miss download endpoints such as /download?id=123, files behind intermediate document pages, links inside separate iframe documents, and resources loaded only after further interaction. A matching URL also does not prove that the server will return a PDF.
If a document is missing, inspect its link and open it normally. That tells you whether to adjust the extraction rule, visit another page, or use the site's own download flow.
2. Review the list before transferring files
Treat the URL list as a small input manifest. Check that it represents the collection you actually want.
- Remove unrelated documents and obsolete versions.
- Open a representative link to confirm it returns the expected file.
- Keep source page addresses in a separate project note for context and attribution.
- Keep temporary or signed download URLs out of public repositories.
Preserving query parameters matters here. Two URLs with the same path may return different versions. Conversely, two different addresses may return the same document. URL deduplication alone cannot resolve that distinction.
For a collection you need to reproduce regularly in CI, a maintained script with explicit inputs and logging may fit better. The browser workflow is useful when choosing the files is an interactive task.
3. Bring the selection into a download workflow
In Simple Mass Downloader, paste the URLs or import the TXT list. Review the resource list and select the files you want to save.
You can also start directly from a page: choose a resource type, click Scan this page, and review the results in the Dashboard. Search matches resource URLs and link text. Domain and extension filters help narrow the selection.
For resources spread across several pages, a multi-tab scan can collect from all tabs in the current window or from the tabs to the right of the current tab. Use that scope when you want a combined list; separate scans have separate results.
When links lead to document landing pages, Find files in pages can look for resources inside selected linked pages. The PDF download guide covers that workflow in more detail.
4. Choose filenames and a destination
Before starting, inspect the filename previews. Several source sites may all offer a file called download.pdf. Numbering and website names can help distinguish those files in the output.
Choose the save mode based on what you will do with the collection:
| Save mode | Useful for |
|---|---|
| Default Download Folder | Separate files managed through Chrome |
| ZIP archive | A modest collection you want to package together |
| Choose a folder | Saving directly into an authorized project directory |
ZIP archives are assembled in the browser, so memory and archive limits matter. Direct-folder saving requires the Dashboard to remain open. Keep Chrome running while the batch is active.
Try a few files first. Confirm the names and destination before expanding the selection.
5. Verify the saved result
A download workflow needs an output check. Open a saved PDF and confirm that it is the document you expected. Compare the successful and failed items against your selection.
Expired URLs, authentication requirements, and server request limits can still interrupt a batch. The extension provides automatic retries and completion checks for supported failures, but the source must still allow access. Review failed items, resolve the cause, and retry those items.
If the collection will become a test fixture or another reproducible project input, record the source version and consider storing checksums alongside your own manifest. A filename alone does not establish that two downloads contain the same bytes.
To try the browser workflow, the installation page includes small sample files. Start there, then apply the same steps to a few documents from your project.
What usually takes more work in your own download tasks: finding the right URLs, naming the files, or verifying what arrived?
Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes — full credit and traffic to the original publisher.