Neembadi – A Sprint to remember for all the right reasons

Jan 2026

By Shijith K

Neembadi was my third sprint since joining T4D last April. Without a doubt this is my most favourite sprint yet. It was a fun, productive, and above all, chill sprint.

Thanks to Sheetal and Diana, the sprint had the perfect balance of team bonding and productive work. And the team bonding didn’t feel forced, it was all natural. We had outings to Gandhi Ashram, Garba night, sessions from Gagan and Nupur, surprise birthday parties for Radhika and Noopur, and some of us played cricket. In between these we had time for BAU and individual platform team meetings.

But, none of this is what makes Neembadi the most memorable sprint for me. It’s a P0 issue and the learnings from it.

The day was going smoothly, some NGO support queries, team meetings and a fun session from Sushil on IT security, and BAU for Bhashini – failing 9 out of 10 times, nothing we haven’t seen before. But then other webhooks also started failing. Checked webhook logs, every webhook request seems to be waiting for a response from the provider. What was happening to all the providers?? Where can I check? The answer: Mighty Oban.

We checked the Oban webhook queues – it was congested, barely functioning. So why is it congested? Some triggers that were running for 10,000 users were using Bhashini to do NMT TTS, but Bhashini fails, and we retry again. Thus each job was taking 4 minutes to complete(fail).

So we found out the issue. Now how do we fix it? How do we unclog? We sat around a table and discussed, none of us had done this before. We had multiple ideas, in the end we decided to share the load between multiple queues. We chose a couple of queues that were, at the time, less used. We tried spreading the load across multiple queues. Now across three queues we had 9000+ jobs waiting to be processed. So all well and good, right? No..!! We missed one important detail – the problem was not that there were 10000 jobs in one queue, it was that each job was taking 4 minutes to complete and in parallel we only executed a maximum of 20 jobs. So now all three queues are jammed :|. Parallelizing failure just gives you… more failure. So to free up the queues there was no other way than deleting the jobs – because anyways they were all going to fail, as Bhashini was not performing well. But deleting is destructive so we should communicate with the NGOs. So we checked with Tejas and Radhika, and decided to delete the jobs and then communicate with the NGOs. Thus we ended up deleting all the jobs that were created as part of the trigger that was run, and saved the day.

But why is this a memorable situation? After all, it’s a P0 issue that affected our users. The reason is that I learned two lessons:

1. We are one hell of a team. We fixed it together. No, it’s not a technically difficult solution but what matters is the ability to think quickly and come up with a solution. We did that well, together – like a team should do.

2. Second is the most important lesson for me: I shouldn’t panic in such situations. I should stay calm so the team can stay calm. I wasn’t behaving like a leader should, and I learned that the hard way (Well, I learn most stuff the hard way, meh.!)

I returned home from Neembadi with these learnings. The learning plus the fun we had made Neembadi a great experience for me. Now I am looking forward to the next sprint, hoping I learn something new there too (maybe without having a P0 issue).

You may also like

Notes from Protsahan’s Girl Empowerment Centre

Driving Beneficiary Impact from Data: Lessons from 1000 Days Fund

Understanding the Fundraising Challenge for Grassroots CSOs