OVERVIEW
QueuePilot visualizes background job throughput, failures and retries across services.
THE PROBLEM
Engineers had no visibility into queue health without digging through logs.
THE SOLUTION
A live dashboard aggregating queue metrics with alerting on failure spikes.
MY ROLE
Sole engineer — backend polling service, metrics schema and dashboard.
ARCHITECTURE
Spring Boot service polling queue backends, MySQL for historical metrics, React dashboard with polling updates.
CHALLENGES
Keeping the dashboard performant while polling multiple queue backends at once.
RESULT
Cut mean-time-to-detect queue failures from hours to minutes.
WHAT I LEARNED
Gained deeper understanding of observability tooling design.
ENGINEERING DECISIONS
▸Multi-backend polling without blocking the event loop
▸Alerting thresholds on failure-rate spikes
▸Historical metrics rollups in MySQL