Information Technology Site Reliability Engineer - Application Support
Sands
Job Description
Job Descriptions:
❖ Contribute to incident response, root cause analysis (RCA), and remediation, focusing on long term reliability improvements.
❖ Perform hands on troubleshooting across application, containers, database, and messaging and Data Lakehouse layers.
❖ Debug, diagnose, and fix application level issues encountered in production, including defects related to configuration, deployment, performance, resource utilization, and runtime behavior.
❖ Actively participate in on call rotations, incident response, and onsite troubleshooting root cause analysis, and remediation of production issues.
❖ Maintain monitoring, logging, and alerting solutions to provide real time visibility into application and platform health
Position Requirements:
❖ Bachelor Degree in Computer Science, Information Technology, Cybersecurity, or a related field
❖ 3+ years of experience in Site Reliability Engineering, Application Support Engineering role
❖ Hands-on experience in supporting production applications in live, customer facing environments.
❖ Proficiency in SRE fundamentals: SLIs/SLOs, error budgets, capacity planning, chaos testing, and toil reduction
❖ Experience with monitoring and logging platforms (e.g. Prometheus, Grafana, ELK stack)
❖ Proficiency in scripting or automation languages (e.g. Python, Bash, PowerShell)
❖ Demonstrated ability to leverage AI-assisted tools (e.g. GitHub Copilot, ChatGPT / enterprise AI assistants, or similar) to accelerate operation tasks while strictly adhering to enterprise security, data protection and compliance requirements
Job details and availability are governed by the employer’s official careers site.