Link Extractor

Paste any block of text, HTML, or markdown. The extractor pulls every URL out of it and lists them in the output box. Filter by domain, drop duplicates, or copy the result as a clean list or a comma-separated string.

Copied to clipboard
0URLs found 0Unique 0Domains 0HTTPS

Sort by

Options

Private by default

Everything runs in your browser. Pasted HTML is parsed locally, not posted to a server.

How to extract links from text

1. Paste your input

Drop the text or HTML you want to mine for links into the input box. Plain text with URLs (a chat log, an article, an email), HTML markup with <a href> tags, or markdown with [label](url) all work. The regex behind the extractor matches anything starting with http:// or https:// regardless of whether it sits inside a tag, attribute, or plain text.

2. Pick filter options

Toggle HTTPS only to ignore http:// URLs (useful when auditing for mixed-content warnings). Drop duplicates removes repeated URLs. Strip tracking params removes utm_*, fbclid, gclid, and similar markers from each URL so the output is clean canonical links. Use the Include domains and Exclude domains filters to narrow down to a single domain or kick a few off the list.

3. Pick sort order

First appearance preserves the order the URLs showed up in your input (default). Alphabetical sorts the full URL string. By domain groups all URLs from the same domain together.

4. Copy or download

The output appears in the output box live as you change settings. Use Copy for a line-per-URL list, Copy as CSV for a comma-separated single-line string, or Download .txt to save the result as a file.

When you need a link extractor

Pulling URLs out of a block of text is a recurring chore in several workflows.

Auditing an outbound link list

Paste an article or a blog post and you have every external link in one place, ready to check for 404s, mixed content, or off-topic references.

Cleaning up a chat log or email

People share links in chat all day. Paste the conversation in to get a clean URL list without scrolling through the noise.

Migrating bookmarks or notes

Notes apps often store links inside paragraphs of context. Paste a note in, extract the URLs, and use them to populate a fresh bookmarks file or a CSV.

Filtering by domain for analysis

Auditing only outbound links from your own site? Drop the page HTML in, filter for everything that is not your domain (use the exclude field with your domain). The output is the outbound link set.

Stripping tracking parameters before sharing

Marketing URLs are full of utm_*, fbclid, gclid. Extract a list with Strip tracking params on and you have canonical clean links ready to publish.

What the extractor recognizes (and what it misses)

The extractor uses a single regex that matches http:// or https:// followed by any non-whitespace, non-bracket characters. That covers the vast majority of URLs in real-world text, including:

  • Plain URLs in body text: https://example.com/page?x=1
  • URLs inside HTML <a href> attributes
  • URLs inside markdown links: [label](https://example.com)
  • URLs with query strings, fragments, ports, and most special characters

What the extractor does not match:

  • URLs without a scheme: example.com alone. Add https:// before pasting.
  • FTP, mailto, tel, or other non-HTTP protocols. Adjust the regex via a download-and-modify workflow if you need them.
  • URLs broken across multiple lines. Paste the source on one line first.

Tracking parameter cleanup explained

Most URLs shared online carry tracking parameters: utm_source, utm_medium, utm_campaign, fbclid, gclid, mc_cid, and similar. They are useful for analytics on the sending side but they are noise when you are saving canonical links.

Turn on the Strip tracking params toggle in the sidebar and the extractor removes the following keys from every URL: any key starting with utm_, plus ref, fbclid, gclid, mc_cid, mc_eid, _ga, and yclid. If the URL ends up with an empty query string after the strip, the trailing ? is removed too.

The strip happens before deduplication, so two URLs that differ only by tracking params collapse into one in the output. That makes link auditing a lot cleaner.

Domain filtering for site audits

The include and exclude domain filters take a comma-separated list. The matcher uses substring containment on the domain (no protocol, no path, just the hostname after stripping www.), so a filter of example.com matches blog.example.com and shop.example.com too.

Include filter (whitelist)

Lists only URLs whose domain matches one of the entries. Use this when you want a quick "all links to X" list from a larger document.

Exclude filter (blacklist)

Lists every URL except those whose domain matches an entry. Use this for "all outbound links except mine" audits.

The two can be combined: include first, then exclude trims further inside the included set. For most jobs, one or the other is enough.

Output formats: list, CSV, and file

Three ways to grab the extracted URLs.

Line-per-URL list (Copy)

The default. Each URL on its own line. Best for pasting into a document, a spreadsheet column, or a script.

Comma-separated string (Copy as CSV)

All URLs on one line, separated by commas. Useful for inline lists, command-line arguments, or anywhere a single-line payload is needed.

Downloadable .txt file

Saves the line-per-URL list as links.txt for later use, archival, or sharing.

Related tools

Frequently asked questions

Does the extractor handle HTML and markdown links?

Yes. The regex matches any URL starting with http:// or https:// regardless of whether it sits inside an <a href> tag, a markdown [label](url) construct, or plain body text. Paste the source as-is.

Why are some URLs missing from the output?

Two common reasons. (1) The URL has no http:// or https:// scheme: bare example.com is not matched. Add the scheme before pasting. (2) The URL is broken across lines or has whitespace in the middle. Normalize the source first.

Can I filter by exact domain only, not substring?

The filters use substring matching, so example.com matches blog.example.com too. For an exact-domain filter, paste the result into a spreadsheet and use a more precise text-match filter there, or use the by-domain sort in the sidebar to group by domain and then visually trim.

What counts as a tracking parameter?

Strip Tracking Params removes any key starting with utm_ (utm_source, utm_medium, utm_campaign, utm_term, utm_content), plus fbclid, gclid, mc_cid, mc_eid, _ga, yclid, and ref. If your source uses a custom tracking key, paste the cleaned URLs into find-and-replace for one more pass.

Does the text leave my browser?

No. The extractor runs entirely in your browser using JavaScript. Your input, the URLs found, and the output are never sent to our servers and are not logged.

How does it handle URLs ending with punctuation?

After the regex match, the extractor strips trailing punctuation (period, comma, semicolon, colon, exclamation, question mark, closing brackets, quotes) from each URL. This catches the common case of a URL at the end of a sentence: See https://example.com. produces https://example.com, not https://example.com. with a trailing period.

Can I extract URLs from a PDF or a Word document?

Not directly. Open the document, select all (Ctrl-A / Cmd-A), copy, and paste into the input box. The text content with embedded URLs comes across cleanly enough that the extractor can find every link.

Is there a limit on input size?

No hard cap. The regex pass is fast and the dedup/filter step is O(n). The extractor comfortably handles inputs of several megabytes.

More wordcounter.ai tools

Other tools you might find useful.

Browse the full catalog →