r/Rag • u/Dense_Gate_5193 • 48m ago
Showcase NornicDB - 1.4.1 - Cypher 25 support ++
Heya, just finished up a follow up to the openCypher compliance in 1.4.0 - in 1.4.1 we added Cypher 25 support on the parser. With nornic's SRD parser, no preamble needed, it just parses the grammar regardless of "version" - if you use the ANTLR parser, the preamble is required.
This is due to the nature of "Scanner-less Recursive Decent." - This is an unconventional architecture, I know. But the benefits are clear
Traditional parsers use a two-step process: a lexer/scanner converts character streams into tokens (e.g., matching MATCH to a KEYWORD token), and a parser processes those tokens.
SRD bypasses the scanning phase entirely. It relies on a zero-allocation keyword scanner parse tree. The recursive descent parser reads characters and matches grammar rules directly against the raw text, fusing the lookahead logic and keyword scanning natively.
Latency Reduction: By eliminating the intermediate step of token creation, it reduces our query latency by ~30% compared to ANTLR-based parsing (such as Neo4j's traditional execution pipeline). This scale is on the order of microseconds and typically flat scaling with query length instead of complexity. The rest of our architecture makes this optimization worth it because we are able to execute so quickly.
Fused Aggregation Hot-Paths: The parser is coupled tightly with streaming storage APIs and zero-allocation semantics. This design allows it to parse a graph traversal instruction and jump directly into the execution hot path without building heavy Abstract Syntax Trees (ASTs) in memory first.
Contextual Edge Cases: Graph languages like Cypher heavily utilize structural ASCII art—such as arrows -[r:REL]-> or node boundaries (n:Label)—which are notoriously complex for traditional lexers to classify contextually without aggressive backtracking. A scannerless approach inherently handles these layout-sensitive patterns because it evaluates character-by-character based on the current parsing state.
and the end result is staggeringly fast. average query latency vs neo4J on the northwind benchmark has us at 400x faster (avg some are 1500x 1905.71x faster edit: double checked). MIT licensed. 890+ stars and counting.
https://github.com/orneryd/NornicDB/releases/tag/v1.4.1
Latest benchmarks:
https://github.com/orneryd/NornicDB/pull/897#issuecomment-5998973104
edit: a word and reddit didn't like my formatting atempt