Site Reliability Manager Resume
2.00/5 (Submit Your Rating)
Medford, MA
SUMMARY
- Experienced and passionate IT professional committed to delivering service and innovative technological solutions.
- Collaborative style working wif cross - functional teams, developing and implementing standard operating procedures, diligently troubleshooting issues to identify root cause and prevent recurrence.
- Passionate and driven toward continuous improvement and professional development by keeping current on new tools and technologies.
- Dependable and known for taking ownership, being results-driven, and using a practical and creative approach to problem-solving.
TECHNICAL SKILLS
Platforms: Windows, Linux
Applications: Splunk, VMware, AppDynamics, Sharepoint, AlertSite, SCOM, MSMQ
PROFESSIONAL EXPERIENCE
Confidential, Medford, MA
Site Reliability Manager
Responsibilities:
- Maintained workflow by leading infrastructure and application teams on managing and monitoring all core application systems in Confidential 's QA, Staging, and Production environments.
- Supported daily operations by coordinating communication between cross-functional teams to resolve incidents impacting service platforms.
- Prevented future network degradation and outages by leading meetings wif cross-functional teams to identify root cause analysis and re-instrument triggers.
- Sustained Service Level Agreements (SLAs) by managing on-call rotations and provided inputs to development teams and partners.
- Avoided repeat incidents and enhanced future responses by developing Blameless Postmortem process for root cause analysis and standard operating procedures.
Confidential
Principal Systems Engineer
Responsibilities:
- Lowered costs, improved web application services speed and performance by migrating multiple complex on-premise production and testing environments to AWS cloud.
- Improved stability and functionality of company application services by managing deployments wif project teams in engineering and packaging releases to company's core application servers, ensuring critical enhancements are continuously released.
- Met system stability and service level agreements wif clients by analyzing and improving performances, defining and monitoring core components and alert thresholds, setting key performance indicators.
- Reduced potential damage and quickly restored business operations by designing and implementing disaster recovery procedures for company's core web application.
- Increased overall production and support by gathering and researching data about internal hardware and software.
