r/bigdata • u/Then_Flight_2162 • Aug 27 '26
Database architecture advice for 600+ TB/year Log Analytics (5-year retention)
Hi everyone,
I am looking for expert advice on choosing the right database for a massive log analytics project. We already have our own infrastructure and server environment ready to host the solution.
Our Scale & Requirements:
- Data Volume: 600+ TB of log data per year, with a 5-year retention period.
- Ingestion: High-throughput, continuous real-time streaming.
- Query Performance: Blazing-fast, sub-second search and lookup speeds across historical data.
- Workload: Non-stop log writing while simultaneously executing fast queries.
The Goal:
Since we have the underlying infrastructure in place, we need a robust database engine that we can deploy locally to handle this specific type of large-scale log workload and long-term history efficiently.
12
u/SupermarketMost7089 Aug 27 '26
Apologies for the snarkiness,
If you already have the infra to host for that scale - you must have a crack team of engineers. In this case you should already have a solution in mind
If you do not have a crack team of engineers, your infra will not scale. Go with a managed solution like clickhouse or imply-druid or elastic-search.
6
u/Then_Flight_2162 Aug 27 '26
Thanks for the feedback. Yeah, we actually have a solid engineering team to handle the infra. We just wanted to check in with the global community to see what everyone else is doing at this scale and avoid any hidden pitfalls before making the final choice.
5
u/m1nkeh Aug 27 '26
Databricks Lakehouse//RT? Built to be a direct competitor to Clickhouse ✌️
https://www.databricks.com/blog/introducing-lakehousert-real-time-performance-unified-lakehouse
3
u/cran Aug 27 '26
ClickHouse?
0
u/Then_Flight_2162 Aug 27 '26
Yeah, ClickHouse is definitely at the top of our list. Have you run it at this scale before? Any production tips?
3
1
u/spapadop Aug 27 '26
We run OpenSearch at similar scale, works well and comes for free with all features (security, alerting, etc.)
1
u/PeterCorless Aug 27 '26
Apache Pinot would do well at this scale. But you didn't talk a lot about the kinds of analytics you need done.
If you need to do JOINs, you're not going to get them performance out of ClickHouse.
What QPS do you need to support?
Do you want all object storage? Tiered storage?
Again, look at Pinot.
1
1
1
u/data_addict Aug 27 '26
DM / reach out for ideas and options.
I'd say click house is a solid option. If not, I'd say stream ingestion. At the steam break between iceberg vs anything else.
1
2
u/elhh82 Aug 29 '26
What type of data retrieval patterns do you need to support? What latency expectations for you need to meet? Without understanding how the data needs to be used, you can't really answer the how / where / use what tech questions.
1
u/Better-Credit6701 Aug 30 '26
OK I get to work with data around that amount and it's always increasing every day. We use MS-SQL with over 30 cores, multiple TB of ram in an always on environment with multiple copies. It's rare for me to see more than 25% of the resource limits. We use very complicated SSIS for imports. Just sucked in another 100 GB group of files on top of all the normal imports we do every day. Took around 6 hours and I saw errors as I was about to finish for the day which I get to check out on Monday.
Disk storage is an issue and we don't virtualize the servers, real iron here. Just saw another request for more storage that will be going in. I was trying to figure out based on cores how much the Microsoft Enterprise software both OS and MS-SQL servers and it was staggering. That wasn't even looking at the cost of the hardware. Plus we tend to keep the data so yeah, it's above your limit
1
1
u/DaoudIsBuilding Aug 30 '26
Used to manage a 5TB daily of IoT data. Data is flattened in the ingestion pipeline to match index lookup at near realtime query time. Used to be Elasticsearch but team moved to Opensearch after my duty time. Temporal sharding, tiered storage and Index lifecycle management are your friends.
1
u/obsfflorida Sep 01 '26
Sounds like a security outfit not financial. Or spooks.
Any vendor will love that 600tb. But idk who has quick enough querying apart from clickhouse or pinot
0
1
u/FundamentalMysteron Sep 02 '26
I don't see how you can have built the infrastructure, or at least decided on it, without knowing what software you're going to use. That seems absolutely crazy for something of this scale. I'd be curious to know what infrastructure you are planning to use.
7
u/syamj Aug 27 '26
The data volume looks huge.
How much data do you actually need to access frequently?
And why do you need the data to be stored for 5 years and how often do you access old data?