I want to store user event logs and run aggregation queries to see real-time trends. My current approach is just dumping everything into a single collection, but the aggregations are getting slow. What are the best patterns for analytics in MongoDB? Should I be using TTL indexes or pre-aggregated collections?
Achieving real-time analytics in MongoDB requires shifting from raw event storage to pre-aggregated documents or time-series collections to avoid performance degradation during complex query execution.
2 answers
MongoDB isn't a silver bullet for real-time analytics, and dumping raw events into a single collection will kill your latency every time. Stop trying to aggregate on-the-fly and move toward a pre-aggregated pattern where you process logs into time-bucketed documents. If you keep hitting the disk with massive collection scans, you are just throwing hardware at a design failure.
I noticed similar latency issues in my own projects, Herminia Garcia. It sounds like moving toward pre-aggregation is the safest path forward to avoid further hardware strain and performance bottlenecks.
This is so helpful, Herminia Garcia. I am currently drowning in massive collection scans and had no idea it was a design failure. I need to switch to time-bucketed documents immediately.
Using a single collection for raw events is the fastest way to turn a production database into a brick. You are fighting the storage engine instead of working with it.
I remember trying to scale a dashboard back in 2017 using a similar naive approach, and the lock contention basically nuked the entire application once we hit a few million records. We had to pivot to an out-of-band pipeline to push data into a dedicated rollup collection just to keep the UI from timing out for our users. Do yourself a favor and offload that aggregation work before your technical debt becomes a total system outage.
Herminia Garcia, your focus on time-bucketed documents makes sense, though I worry about the complexity of implementation. I might be overthinking the overhead, but this approach seems much more stable than my current setup.