24 sep
|
Allied Global Technology Services
|
México
24 sep
Allied Global Technology Services
México
Ops Engineer – Infrastructure & Operations
Location: Canada or Mexico – Nearshore
Work Modality: 100% Remote
Openings: 2
Seniority: Mid to Senior Level – Individual Contributor
Reports to: Infrastructure & Operations Manager
About the Role
We are looking for an experienced Ops Engineer – Infrastructure & Operations to help maintain the health, reliability, and performance of a complex hybrid infrastructure environment spanning on-premises systems and Microsoft Azure.
This is a hands-on operations role, not a project-focused position. The primary responsibility is operational ownership: keeping systems available, responding effectively to incidents, maintaining infrastructure, and continuously improving reliability through monitoring and automation.
You will work directly with production infrastructure, handling monitoring, patching, upgrades, outage response, troubleshooting, and automation. Strong documentation practices are essential, including maintaining runbooks, diagrams, operational procedures, and change records.
Key Responsibilities
- Own the day-to-day health, availability, and reliability of hybrid infrastructure across Microsoft Azure and on-premises environments.
- Monitor infrastructure and proactively identify reliability, capacity, and performance issues.
- Perform infrastructure patching, upgrades, maintenance, and operational improvements.
- Participate in and lead incident and outage response, providing clear communication throughout the resolution process.
- Troubleshoot complex infrastructure issues and drive problems through root-cause analysis and resolution.
- Build and maintain Ansible playbooks to automate operational and infrastructure tasks.
- Develop automation and operational tooling using PowerShell, Python, or Bash.
- Build, maintain,
and improve monitoring dashboards and alerts using platforms such as Splunk, Datadog, or New Relic.
- Tune alerts to reduce noise and improve actionable operational visibility.
- Identify repetitive manual processes and replace them with reliable automation.
- Create and maintain high-quality runbooks, architecture/infrastructure diagrams, change records, and operational documentation.
- Participate in operational support and outage response as required.
- Continuously improve infrastructure reliability, observability, and operational efficiency.
Must-Have Qualifications
- 5–10 years of experience in Infrastructure Engineering, IT Operations, Systems Engineering, or a similar role.
- Proven experience operating within a complex hybrid environment combining cloud and on-premises infrastructure.
- Strong hands-on operational experience with Microsoft Azure, gained through daily production work rather than certification alone.
- Hands-on experience with Ansible, including building, maintaining, and troubleshooting production playbooks.
- Strong scripting skills in at least one of the following:
- PowerShell
- Python
- Bash
- Practical experience with a major observability and monitoring platform, such as:
- Splunk
- Datadog
- New Relic
- Experience creating dashboards, defining operational metrics, configuring alerts, and tuning monitoring rather than simply consuming existing dashboards.
- Demonstrated experience responding to production incidents and outages.
- Strong troubleshooting and root-cause analysis capabilities.
- Excellent documentation discipline, including runbooks, diagrams, operational procedures, and change records.
- Ability to remain calm during high-pressure incidents and communicate clear, concise updates to technical and business stakeholders.
Nice-to-Have Experience
- Experience with Cloudflare, including WAF, DNS, and CDN technologies.
- Exposure to on-premises infrastructure, including:
- Networking
- WiFi
- Servers
- End-user or infrastructure devices
- Experience with ServiceNow.
- Experience with Azure DevOps.
- Experience working within PCI or another regulated environment.
- Experience supporting infrastructure subject to formal change-management and compliance requirements.
What We’re Looking For
Strong candidates should be able to provide concrete examples of:
- An operational process they automated, why it was repetitive, and the measurable time or effort saved.
- A significant production outage they handled end-to-end, including troubleshooting, root cause, resolution, and preventive actions implemented afterward.
- A monitoring dashboard or alerting capability they personally built that became valuable to other engineers or operational teams.
- Situations where they took direct ownership of infrastructure availability and production reliability.
This position is best suited for engineers who have genuine operational ownership. Candidates whose experience is primarily focused on cloud project delivery, migrations, or implementations without responsibility for production uptime or incident response may not align with the role.
📌 Ops Engineer – Infrastructure & Operations (Based in Mexico or Canada) (México)
🏢 Allied Global Technology Services
📍 México