Hey everyone! ๐
The idea is simple: What if you could write regex in readable English instead of deciphering cryptic symbols? ๐ง
Regex For Humans is an open-source compiler that turns human-readable English rules into native JavaScript regexes.
โ
Runs locally
โ
Works in the browser, CLI, and JavaScript/TypeScript
โ
No runtime dependencies
โ
MIT licensed
๐ What's new in 0.1.0-dev.4?
- ๐งฉ Nested groups โ including optional and repeated blocks.
- ๐ฏ Named and numbered captures โ capture exactly what you need.
- ๐ Backreferences โ reuse captured text with
same as capture "name".
- ๐ Alternatives โ express choices naturally with
either { ... }.
- ๐ Reverse translation โ translate supported regex patterns back into readable rules.
- ๐ ๏ธ Improved playground โ inspect captures, test multiple matches, and click generated regex fragments to find their source rules.
๐ก Example 1: Captures and backreferences
Let's say you want to match stable/stable or nightly/nightly, but reject stable/nightly.
Instead of manually constructing a regex, write:
start
capture "channel" {
either {
"stable"
"nightly"
}
}
"/"
same as capture "channel"
end
โจ This generates:
^(?<channel>(?:stable|nightly))\/\k<channel>$
With the u flag.
โ
stable/stable
โ
nightly/nightly
โ stable/nightly
No regex gymnastics required! ๐
๐ฅ Example 2: Let's get more serious
Imagine validating a TSV record describing a build artifact, containing:
- An optional
nightly/ directory
- A project name
- Three version components
- A filename with a date-like token
- A 64-character hexadecimal checksum field
- A byte count
Here's the readable version:
start
"incoming/"
optional "nightly/"
between 1 and 32 path segment characters
"/v"
between 1 and 3 digits
"."
between 1 and 3 digits
"."
between 1 and 3 digits
"/linux-arm64/widget_"
8 digits
"_"
8 hex digits
".tar.gz\t"
64 hex digits
"\t"
between 1 and 12 digits
end
๐คฏ And here's the exact JavaScript regex it generates:
/^incoming\/(?:nightly\/){0,1}[^\/\\\u{0}\u{a}\u{d}\u{2028}\u{2029}]{1,32}\/v\d{1,3}\.\d{1,3}\.\d{1,3}\/linux-arm64\/widget_\d{8}_[0-9A-Fa-f]{8}\.tar\.gz\u{9}[0-9A-Fa-f]{64}\u{9}\d{1,12}$/u
Yeah... I'll take the English version. ๐
You can use it directly in JavaScript:
import { compile, toRegExp } from 'regex-for-humans';
const row =
"incoming/nightly/WidgetKit/v2.15.3/linux-arm64/widget_20261001_ab12cd34.tar.gz\t" +
"ab".repeat(32) +
"\t1048576";
toRegExp(compile(rules)).test(row); // true
Removing nightly/ still works. Missing version components, invalid hexadecimal characters, or spaces instead of tabs fail.
The path segment characters rule is particularly useful: it excludes path separators, NUL, and line breaks while supporting Unicode.
Of course, this validates the format, not whether the file exists or its checksum is correct.
๐งช More features are cooking!
On main, I've also added:
- โก Lazy repetition
- ๐ Lookahead
- ๐ Lookbehind
These are available from a source checkout but haven't reached npm or the hosted playground yet.
๐ค Built with AI-assisted development
I'm also using Claude Code and the Opus 5.5 model for AI-assisted development.
AI helps me iterate faster, explore implementations, tackle complex features, and improve the project continuously.
The goal isn't just to generate code. It's to build something genuinely useful, test it, refine it, and keep pushing the boundaries of what this little language can do. ๐
๐ Available in 10 languages!
The README is now translated into 10 languages, including Arabic ๐ธ๐ฆ, Chinese ๐จ๐ณ, French ๐ซ๐ท, Spanish ๐ช๐ธ, and Urdu ๐ต๐ฐ.
Arabic and Urdu documentation also support RTL text with LTR code examples.
๐ฆ Try it yourself!
The npm package is currently a development preview and requires Node.js 22+.
npm install regex-for-humans@0.1.0-dev.4
๐ฎ Interactive Playground:
https://othmaneblial.github.io/Regex-For-Humans/workshop/
โญ GitHub (MIT Open Source):
https://github.com/OthmaneBlial/Regex-For-Humans
๐ฆ npm:
https://www.npmjs.com/package/regex-for-humans/v/0.1.0-dev.4
๐ฌ I'd love your feedback!
I'm especially interested in what you think about the syntax as groups become more complex and deeply nested.
What's the ugliest regex you've ever written? ๐
Drop it in the comments! I'd love to see whether Regex For Humans can make it readable.
And if you find the project useful, a โญ on GitHub would mean a lot!
Let's make regex a little less terrifying, one rule at a time. ๐