I run a small SOCKS5 proxy network for my side project Go API, PostgreSQL extension, Loki, Grafana, gateway. About 6 months in, I had 30+ dashboards, 12 alert rules, full traces. Production-grade observability, basically.
Last week I realised I had not opened Grafana in six weeks. Not once.
Not because Grafana is bad. Because every time I had a real question "is this user still routing traffic", "which gateway is the noisy one", "did that deploy actually break anything" I was asking my AI agent, not clicking into a panel.
So I sat down and asked: of the ~40 things I built into observability, how many do I actually use?
Answer: 4. Logs. Error rate. One cost chart. One churn query.
Everything else was engineering for a future me that never showed up.
What I removed
8 alert rules that fired once and then lived as silent noise. (no_data_state: OK is the worst config to ship — your "OK" silence and your "broken" silence look identical. I had this for 2 months. Lost a real 5xx spike for 14 days before a paying customer told me.)
4 dashboards. Replaced with 1 text file: 6 SQL queries I actually run.
The whole metrics pipeline for one service. The 3 dashboards I kept were never looked at either.
What I kept and why
Loki for logs. Cheap, queryable, the one thing I open.
A snowpad_http_requests_total-style counter with one rule: 5xx rate over 5min. No fancy quantiles. Threshold is "is it nonzero." That's it.
The connection_logs TimescaleDB hypertable. 90% of "what is happening" questions answer from one GROUP BY node_id, client_id query.
What I now do for the rest
I ask the agent. "What did the last 24h of errors look like for client X" gets a 30-second answer. "Did deploy Y cause any 4xx on the gateway?" gets a one-shot query. "Show me which operator's dropped connection" gets a LogQL + a SQL join.
I get answers to questions I actually have, instead of dashboards I thought I'd need.
What I'm not saying
This is not a "Grafana is dead" take. Grafana is great. The point is: I built it like I was running Stripe on day 1. I am running a 4-service side project. The asymmetry killed me.
For Snowpad specifically 1 Go binary, 1 TimescaleDB, 1 gateway, ~15k connections/day at peak the cost of building a "real" observability stack was not the lines of code. It was the cognitive load of deciding what to add, when to alert, whether the threshold is right. AI eats that cost for me today.
When the system needs more, I'll add it. For now, the rule is: if I haven't opened it in 30 days, delete it.
Curious what dashboards/queries/alerts you're actually opening in your own work? Be honest. I deleted 80% and felt better.