Site Reliability Engineer
Tecsys Inc. - Bengaluru, India - posted 4d ago
Skills: Kubernetes, AWS, Azure, Terraform, Ansible, Gitlab, Jenkins, Datadog, Ci/cd, Infrastructure as code, Incident management, Monitoring, Observability
At Tecsys, the software we deliver keeps hospitals, pharmacies and distributors supplied with what they need, when they need it. Every role here connects to that outcome.
**Why This Role Matters **
As a Site Reliability Engineer you will join our Network and Security Operations Center (NSOC), a team at the heart of platform reliability for our mission-critical SaaS environments. You’ll respond to incidents, keep day-to-day operations running smoothly, and lead improvements and automation that cut repetitive work and technical debt. By finding the root cause of recurring issues and fixing them for good, you’ll make our platform more dependable for the healthcare and distribution customers who rely on it, and leave a lasting impact on the team.
**What You’ll Be Part Of **
Tecsys is trusted by mission-critical organizations in healthcare and distribution to power resilient, efficient and secure supply chains. We are a global provider of cloud-based, AI-driven software, publicly traded on the Toronto Stock Exchange (TSX: TCS).
Three values guide how we work. We earn trust daily by valuing the customer and acting with integrity. We build each other up by recognizing individual strengths and working as one team. And we explore boldly by staying curious and adapting for success.
You’ll join a culture driven by curiosity, continuous learning, and people who challenge what’s possible.
**RequirementsWhat You’ll Do **
Respond to and resolve P1–P4 incidents during your day hours, working closely with the NOC team across North America and India
Dig into repeat incidents to find the root cause and put fixes in place so they don’t come back
Handle recurring operational work, including onboarding and offboarding requests and monitoring changes
Run regular backup checks and infrastructure reliability checks
Build automation and improvements that reduce manual effort and technical debt for the team
Contribute to projects that leave a lasting impact, such as our backup architecture redesign, which strengthens how we protect backups from a security perspective
Develop tooling and automation on top of AWS and Azure to continuously reduce the need for manual intervention.
**What You Bring **
Bachelor’s degree in computer science or related field
Strong background with at least 3+ years of experience working in a technical SaaS environment
Hands-on experience with Kubernetes, AWS and Azure at scale
Hands-on expertise with Infrastructure as Code and automation tooling (Terraform, Ansible, or similar)
Experience with CI/CD pipelines (GitLab preferred, Jenkins acceptable) and monitoring and observability tooling (Datadog or equivalent), including metrics, alerting and dashboards
Experience with monitoring and alert handling
Solid incident management experience, including on-call rotations, escalations, and postmortem
Benefits
Benefits (Insurance and access to learning platform) and vacation from day one
Flexible work environment
Regular Employee Engagement initiatives
Room to learn, stretch and take on new challenges
Collaboration with global teams
We foster an inclusive workplace where everyone feels valued, respected, and empowered to thrive.
If this role fits what you are looking for, apply today. If you are not actively searching but want to hear more, we are glad to talk
Apply today and discover what you could build, solve, and achieve with Tecsys!