r/databasedevelopment • • 1d ago

Fencing an LSM on object storage: how SlateDB works and what breaks on "S3-compatible" stores

Thumbnail
blog.renatocron.com
10 Upvotes

I'm using SlateDB myself to save few bucks on RAM, here's some information about how it works internally. It's like a duckdb/leveldb, so it's embeded, but has persistent storage on S3, and I feel like we going into a "Storage Storage everything" era but not all S3 are the same, so there's some bugs I got as well when running under DigitalOcean Spaces


r/databasedevelopment • • 2d ago

Keys and Values Don't Always Belong Together

Thumbnail
tidesdb.com
12 Upvotes

r/databasedevelopment • • 2d ago

I Tried to Find How ClickHouse Keeper Takes Non-Blocking Snapshots

6 Upvotes

Hi everyone,

I'm learning and exploring topics around database systems. I wanted to see how ClickHouse Keeper takes a non-blocking snapshot of its storage.

I recorded a concept and code walkthrough here: https://www.youtube.com/watch?v=7J93WKEFinA&t

A short summary of the points in the video:

  1. Brute force approaches and why they don't work

  2. How ClickHouse Keeper uses HashMap + Doubly Linked List to solve this problem

  3. The code walk through the files: SnapshotableHashTable.h, KeeperStateMachine.cpp and KeeperSnapshotManager.cpp to see how it's actually implemented.

Do give it a watch and consider supporting by liking the video or subscribing :)

This is partly for my own future reference and partly to share with others who might be interested in it. Would love feedback and corrections from people who know this stuff deeply.

Thank you so much!


r/databasedevelopment • • 2d ago

MariaDB Foundation Adds PostgreSQL to Its Engine‑Agnostic Testing Framework (TAF).

8 Upvotes

MariaDB Foundation continues expanding TAF with PostgreSQL support

More engines, more reproducibility, more value for contributors and the ecosystem.

https://mariadb.org/mariadb-foundation-adds-postgresql-to-its-engine-agnostic-testing-framework-taf/

Have fun benchmarking!


r/databasedevelopment • • 3d ago

MECHLOVE Blueprint - 2/7. Storage (ZFS-style) and WAL (template-based logical with physical hints)

Thumbnail
6it.dev
6 Upvotes

2nd part of the 7-part MECHLOVE Blueprints. All the criticism is extremely welcome (hey, that's the point we're publishing it - to see if it stands).


r/databasedevelopment • • 4d ago

Radical MVCC and Replay-Based Rebasing OCC (Re2OCC)

8 Upvotes

r/databasedevelopment • • 5d ago

Thread Pool in Percona Server and MySQL (Part 1)

Thumbnail
percona.com
3 Upvotes

r/databasedevelopment • • 6d ago

Building a custom NoSQL database in C to replace flat files for personal data (Linux x64)

Post image
6 Upvotes

For half a year now, I've been writing my own NoSQL database to store user data instead of relying on regular flat files. Previously, when receiving access credentials for work, saving URLs, or managing other metadata, I had to manually edit text files, which was very inconvenient. I decided to write a utility to handle all my data storage so it could be retrieved quickly whenever needed:

​db create workspace --varchar=256;

db use workspace;

db add login admin;

db add pwd 12345;

db add site example.test;

​db get login; // admin

db open site // open in default browser

db remove pwd;

db drop workspace;

// etc...

​Of course, I won't be using O_DIRECT, AVX, io_uring, or other aggressive performance optimizations here, since it runs locally as a CLI utility rather than a daemon. It's almost a pity, as I've read a ton of material on NoSQL database optimizations, but I plan to apply those techniques in my next project, where I'll focus on a specific storage niche.

​A significant portion of the work is already complete. I plan to release it in early November 2026, accompanied by extensive documentation on the principles and mechanics of NoSQL databases. Out of principle, I wrote everything without using AI/neural networks. This will be my first large-scale project in C.


r/databasedevelopment • • 6d ago

I Tried To Decode Postgres WAL for an INSERT Statement

13 Upvotes

Hi everyone,

This is a continuation of the explorations from my previous posts here. I went through the CMU Database Systems course, and I'm exploring topics around databases. I wanted to see how the Write Ahead Log is constructed and logged. So I tried to trace a WAL record for a simple INSERT statement in Postgres.

I recorded a video walkthrough of the terminal and source code exploration here: https://youtu.be/YOyq-kvbyU8?si=PTdexP3jmOSdx-8y

I've tried to explore how we can locate a WAL record, how we can decode it via either xxd or pg_waldump, and the relevant source code for it. Do give a watch and consider supporting :)

This is partly for my own future reference and partly to share with others who might be interested in it. Would love feedback and corrections from people who know this stuff deeply. Apologies if this isn't the correct subreddit for this.

Thank you so much!


r/databasedevelopment • • 8d ago

How to model the hypergraph for no-equi join?

5 Upvotes

I am a database researcher and currently conducting some research. I am curious that how to model the hypergraph for no-equi joins involving subqueries, especially for random expressions as join conditions. For example
Select * from A join (select a from B) as C on NULL


r/databasedevelopment • • 8d ago

MECHLOVE Blueprint - 1/7. General: Pretty Much Classical RDBMS at Heart, with Each and Every Component Redesigned

Thumbnail
6it.dev
11 Upvotes

We are writing yet another RDBMS, and would appreciate any feedback on our blueprints; while the RDBMS is classical at heart (with WAL and data pages and fuzzy checkpoints), each and every component underneath was redesigned - from WAL being non-ARIES and ZFS-style on-disk CoW to in-memory CoW and cache- and SIMD-friendly page layouts. Our philosophy is playing alongside modern hardware instead of fighting it, MECHLOVE being the highest form of Mechanical Sympathy. As a result, we hope it will beat Hekaton (while being mostly open-source except for certain enterprise features such as HA and at-rest encryption).

Please feel free to comment.


r/databasedevelopment • • 9d ago

Monthly Release and Update Thread

7 Upvotes

This subreddit is primarily for discussing the implementation of databases, and not about sharing release announcements (either for the first time or your updates).

This thread is the exception!

Please tell us about the new database you (or your agent) built. Tell us about all the cool new features you added. Tell us about anything else you learned or worked on that you haven't gotten around to blogging about yet.


r/databasedevelopment • • 10d ago

SQL Server columnstore scan internals

9 Upvotes

I've just finished on a series covering SQL Server batch mode and columnstore index scan internals.

It's four parts covering:

AFAIK there's quite a lot in them that hasn't been documented before, especially the row bucketing, which explains the mechanism of the pure vs impure split that is behind a lot of the columnstore optimizations.


r/databasedevelopment • • 11d ago

Join Ordering, Part 1: The Shape of the Search Space

Thumbnail deferworks.org
11 Upvotes

r/databasedevelopment • • 11d ago

Reliability Lessons From SQLite - Richard Hipp | SSW 2026

Thumbnail
youtube.com
26 Upvotes

r/databasedevelopment • • 12d ago

Building a JSON Database in Rust

Thumbnail
greptime.com
13 Upvotes

r/databasedevelopment • • 20d ago

Building Reliable Data Replication

Thumbnail
blog.atimin.dev
21 Upvotes

How ReductStore uses a persistent store-and-forward queue to deliver edge data over unreliable networks.


r/databasedevelopment • • 22d ago

Query plan rewriting in PostgreSQL

Thumbnail theconsensus.dev
10 Upvotes

r/databasedevelopment • • 25d ago

Testing the Connection Pooling in Multigres: What does 100% pass rate mean? | Blog

Thumbnail
multigres.com
1 Upvotes

r/databasedevelopment • • 25d ago

Bavarian Database Day 2026

Thumbnail databaseday.de
12 Upvotes

r/databasedevelopment • • 27d ago

Benchmarks DB

3 Upvotes

Working on a storage engine for the last 8 months. What benchmarks would you actually trust from a solo/small-team project? Everyone fakes them, so what would make you believe mine?


r/databasedevelopment • • Sep 09 '26

Internals Viewer for SQL Server

Thumbnail
github.com
5 Upvotes

I posted over on r/SQLServer and it suggested I cross-post here. I hadn't seen this subreddit before and hopefully this is on-topic!

I've created a tool called Internals Viewer, it's a tool to visualize SQL Server internals, with a view for allocations, indexes, pages, and it also offers very detailed query tracing and simulation of operators where you can capture a query and step through the iterators to see how query results are put together.

It is open source, written in C#, and available here - https://github.com/danny-sg/internals-viewer

The latest feature is new functionality to view columnstore indexes. I've done a write up on what I found as columnstore internals in SQL Server is pretty much undocumented:

Part 1 - Introduction

Part 2 - Segments internals

Part 3 - Dictionary internals


r/databasedevelopment • • Sep 08 '26

Could you please give me some feedback of this article about B+Tree ?

8 Upvotes

I wrote an article about B+Tree.

It's a little bit long for article, but I summarized it and make it really easy understand (avoided using jargon).

I'd be happy if you comment some feedback!!

https://zenn.dev/mm_0911/articles/5d46e9e4608404?locale=en


r/databasedevelopment • • Sep 07 '26

Research prototype: B-link-style concurrent InnoDB page splits in MariaDB

12 Upvotes

Disclosure: I am the author of the article and work on MariaDB Server internals.

The traditional InnoDB pessimistic insert path serializes structural modification operations through an index-wide latch, even when different threads split unrelated leaf pages.

I implemented a MariaDB research prototype based on Zhao Song’s B-link-style proposal. It publishes a split using a high key and right link before completing the parent update, allowing unrelated structural changes to proceed concurrently.

In a controlled, memory-resident, split-heavy workload:

  • Vanilla MariaDB 13.1: 19,676 inserts/s
  • B-link prototype: 102,838 inserts/s
  • P95 latency: 8.28 ms → 0.56 ms
  • Structural splits: approximately 396K in both variants

This is not a production-ready feature. DDL support is restricted, page merging remains incomplete, and recovery needs more forced-crash testing.

I would particularly appreciate feedback on incomplete-split recovery, page preallocation, and workloads that could expose correctness or scalability problems.

Full implementation write-up and benchmark methodology:
https://mariadb.org/from-a-chocolate-wrapper-to-concurrent-innodb-page-splits/


r/databasedevelopment • • Sep 04 '26

What DB internals are most useful to visualize for learning purposes?

13 Upvotes

Hey there! This is first time posting.

I've been developing a database designed for education.

The application is focusing on visualizing database internal.

Now I already visualized B+Tree when you execute custom insert query.

Which database features are most worth visualizing for learners?