How do you answer "Tell me about a production bug you fixed"?
- Technology:
- Behavioral
Quick answer
Walk through how you noticed the issue, limited the impact, found the root cause, fixed it, and what you changed afterwards so the same class of bug could not happen again.
What the interviewer is evaluating
- A structured debugging approach
- Judgement about mitigation versus fixing
- Ownership and communication during an incident
- Learning and prevention afterwards
How to structure your answer (STAR)
- Situation: what broke and who was affected.
- Task: your role in the response.
- Action: mitigation, investigation steps and the fix.
- Result: impact, plus the prevention you added.
Example answer
Situation: after a release, some users reported being charged twice at checkout.
Task: I was on call and owned the checkout service.
Action: I first disabled the new retry logic with a feature flag, which stopped new duplicates. The logs showed that slow responses from the payment provider triggered our retry before the first request had finished. I made the payment request idempotent with an idempotency key, so a retry could never create a second charge, and wrote a test that simulated a slow provider.
Result: duplicates stopped, the affected customers were refunded the same day, and we added an alert on duplicate payment attempts. Since then, any external call we retry must be idempotent.
Discussion
This question tests your debugging process and how you behave under pressure. Interviewers want a calm, methodical story, not a hero story.
A strong answer has a clear order: detection (alert, logs, user report), mitigation first (rollback, feature flag, scaling), investigation (reproducing it, narrowing the cause with logs, metrics or git bisect), the actual fix, and prevention (tests, monitoring, a post-mortem).
Be specific about the technical cause, at a level the interviewer can follow. A memorable detail, such as a timezone edge case or a missing index that only hurt at peak traffic, makes the story credible.
Key points
- Mitigate impact before hunting for the root cause
- Explain how you narrowed the cause down
- Finish with prevention: tests, alerts, process changes
- Stay blameless
Common mistakes
- Jumping straight to the fix without explaining how you found the cause.
- Forgetting mitigation: users suffered while you investigated.
- Blaming the person who wrote the bug.
Follow-up questions
- How do you debug an issue you cannot reproduce locally?