Regex Tester & Debugger→Specialized Version
🔍

Hashtag Regex Tester

Test and explain the hashtag tester pattern

//g
Flags:
Examples:
Shipping today #buildinpublic and #SEO not#ahashtag #123numbers
#MatchIndexGroups
1 #buildinpublic14buildinpublic
2 #SEO33SEO

The Pattern

``regex (?:^|\s)#([A-Za-z][A-Za-z0-9_]*) `

Flags: g

What It Accepts and Rejects

InputResult
Shipping today #buildinpublic and #SEO✅ matches
not#ahashtag❌ rejected
#123numbers❌ rejected

The Boundary Before the Hash

(?:^|\s) requires the hash to be at the start of the string or preceded by whitespace, which is what stops not#ahashtag from matching. A bare #\w+ would match inside any word, and inside URLs with fragments.

The first character class is [A-Za-z] rather than \w, so #123numbers is rejected — most platforms treat a purely numeric tag as invalid, and a leading digit is how accidental matches on issue references (#4212) creep in.

The Limits of This Pattern

The pattern is ASCII-only, which excludes hashtags in most of the world's scripts. Twitter, Instagram and TikTok all accept Unicode letters, so #日本 and #café are real tags this rejects. Switching to \p{L} with the u flag fixes that.

Platforms also differ on length limits, underscores and whether tags are case-sensitive for display but case-insensitive for search. If you are matching tags against a database, fold case on both sides.

Testing Before Shipping

A regex that has only been tried against inputs you expect to match is untested. Every pattern needs three kinds of case:

1. Valid inputs that should match, including the awkward-but-legal ones. 2. Invalid inputs that should not, especially near-misses that differ by one character. 3. Adversarial inputs — very long strings, unusual Unicode, and nesting that could trigger catastrophic backtracking.

Paste your own examples into the tester above and watch which lines highlight. A pattern that matches everything you throw at it is usually too permissive rather than correct.

Catastrophic Backtracking

Nested quantifiers over overlapping character classes — (a+)+, (\w+\s?)* — can take exponential time on a non-matching input. On a server that is a denial-of-service bug, not a performance issue. If a pattern is applied to user input, bound the input length first and prefer explicit alternation over nested repetition.

Anchors, Greediness and Backtracking

Three behaviours account for most regex surprises, and this pattern shows all three.

Anchors. ^ and $ pin the match to the start and end of the input. Without them, \d{3} matches the 123 inside abc123def. With m in the flags — as here — they pin to each *line* instead, which is what lets one pattern be tested against a list.

Greediness. .* takes as much as it can and gives back only when forced; .*? takes as little as possible. On , the pattern <.*> matches the whole string and <.*?> matches just .

Backtracking. When a match fails, the engine reverses and tries other splits. Nested quantifiers like (a+)+ make that exponential, and a 30-character input can hang a server — a class of denial of service known as ReDoS. Avoid nesting quantifiers, and prefer explicit character classes over . wherever you can.

Testing It Properly

`javascript const pattern = /(?:^|\s)#([A-Za-z][A-Za-z0-9_]*)/g;

// A global regex keeps lastIndex between calls, so reusing one across // test() calls returns alternating true/false on the same input. pattern.lastIndex = 0;

// Named groups make the result readable const named = /(?\d{4})-(?\d{2})/; const { groups } = '2026-08'.match(named); `

Write the failing cases first. A pattern that accepts everything valid is easy; one that also rejects everything invalid is the hard half, and it is where the bugs are.

When Not to Use a Regex

Structured formats have parsers, and the parser is always more correct: new URL() for URLs, DOMParser for HTML, JSON.parse` for JSON, a date library for dates. Reach for a regex to *find* things in unstructured text, not to validate something a parser understands.

Frequently Asked Questions

Does this pattern handle every valid case?

The pattern is ASCII-only, which excludes hashtags in most of the world's scripts. Twitter, Instagram and TikTok all accept Unicode letters, so `#日本` and `#café` are real tags this rejects. Switching to `\p{L}` with the `u` flag fixes that. Platforms also differ on length limits, underscores and whether tags are case-sensitive for display but case-insensitive for search. If you are matching tags

Why does the pattern use the flags it uses?

Flags `g`. The `g` flag finds every match rather than stopping at the first. Changing them changes the results, so test with the flags you will ship.

Is validating this with a regex the right approach?

For a format check before doing real work, usually yes. For anything security-critical, a regex confirms shape and nothing else — parse the value with a real parser, or verify it against the system that owns it, before trusting it.

Related Tools

Explore other tools you might find useful:

More Regex Tester & Debugger tools

You might also need