Scaling to 10 Million Messages: Our Infrastructure Journey
How we built a system that handles 10 million messages per month with 99.9% uptime and sub-second response times.
Varun Sharma
Founder
The Growth Challenge
When we launched Agent Rush, we handled a few thousand messages per day. Now we process over 300,000 daily. Here's how we scaled without breaking.
Architecture Evolution
Phase 1: The Monolith (0-100K messages/month)
We started simple:
This worked great for our first customers. Simple to debug, easy to deploy.
Phase 2: Service Split (100K-1M messages/month)
As load increased, we split into services:
Each service could scale independently.
Phase 3: Event-Driven (1M-10M messages/month)
At scale, synchronous processing hit limits. We moved to events:
Message In → Queue → Process → Queue → DeliverBenefits:
Key Technical Decisions
1. Database Strategy
We use a hybrid approach:
2. AI Inference
Running LLMs at scale is expensive. Our approach:
3. Multi-Region Deployment
Indian users expect low latency. We run in:
Lessons Learned
1. Observability is Non-Negotiable
We track:
You can't fix what you can't see.
2. Graceful Degradation
When systems strain, we:
3. Chaos Engineering
We regularly break things on purpose:
Better to find weaknesses in testing than production.
Current Numbers
Scale doesn't require a massive team—it requires smart architecture.
Varun Sharma
Founder
Building the future of customer support at Agent Rush. Passionate about AI, product design, and creating delightful user experiences.