02 oct
|
Tekshapers
|
Guadalajara
02 oct
Tekshapers
Guadalajara
AWS Production Support Engineer The MSD Operations Specialist is responsible for managing daily service desk activities, ensuring seamless IT support operations, incident resolution, service request fulfillment, and maintaining high customer satisfaction. The role involves handling tickets, monitoring systems using observability tools, managing cloud and deployment platforms, coordinating with internal teams, and ensuring adherence to SLAs.
? Key Responsibilities
? ? Incident & Service Request Management
- Monitor and manage incoming tickets via ServiceNow / ticketing tools.
- Prioritize incidents and service requests based on severity and SLA.
- Troubleshoot and resolve L1/L2 issues or route to appropriate teams.
- Ensure timely updates and communication with users.
? Operational Support Activities
- Perform daily health checks on systems, applications, and services.
- Monitor queues and balance workload to avoid SLA breaches.
- Support infrastructure and cloud operations like:
o Active Directory o AWS Services (EC2, Lambda, ECS, EKS, CloudWatch)
- Work with deployment and infrastructure tools:
o Harness – Manage CI/CD pipelines, deployments, and release monitoring o Aptakube – Monitor and manage Kubernetes clusters and workloads
- Utilize observability and monitoring tools:
o Splunk – Log analysis, alerting, and troubleshooting o Datadog – Infrastructure monitoring, APM, dashboards, and alerting
- Support CDN and edge services:
o Akamai – Traffic routing, content delivery, caching, and issue troubleshooting
- Analyze logs, metrics, and deployment events to identify and resolve issues proactively.
? SLA & Performance Management
- Ensure all tickets are resolved within defined SLA timelines.
- Track KPIs such as:
o First Response Time (FRT)
o Resolution Time o Ticket backlog
- Monitor alerts from Datadog, Splunk, and infrastructure tools.
- Ensure deployment and system issues (via Harness/Aptakube) are addressed within SLA.
- Escalate critical or aging tickets proactively.
? Stakeholder & Customer Communication
- Provide regular status updates to users and stakeholders.
- Act as a single point of contact for support issues.
- Maintain professional communication with end users and leadership.
- Coordinate with cross-functional teams (Network, App, Security, DevOps).
- Communicate major incidents related to deployments, CDN (Akamai), or infrastructure outages.
? Documentation & Knowledge Management
- Maintain SOPs, runbooks, and troubleshooting guides.
- Update knowledge base for recurring issues and solutions.
- Document deployment issues, monitoring alerts, and resolutions.
- Maintain documentation for Harness pipelines, Akamai configurations, and Kubernetes troubleshooting (Aptakube).
- Ensure accurate ticket documentation and closure notes.
? Continuous Improvement
- Identify recurring issues using logs, monitoring, and deployment data.
- Suggest improvements in:
o CI/CD pipelines (Harness)
o Monitoring dashboards (Datadog/Splunk)
o Cluster management (Aptakube)
o CDN optimizations (Akamai)
- Improve alert tuning, automation,
and deployment reliability.
- Participate in process improvement and service optimization.
- Support audits, compliance, and reporting requirements.
? Required Skills
✅ Technical Skills
- Experience with ticketing tools (ServiceNow, Remedy, Jira)
- Knowledge of:
o Windows OS / MacOS basics o Active Directory & IAM o M365 (Exchange Online, Teams, SharePoint)
o Basic networking concepts (DNS, VPN, IP)
o AWS cloud services
- Experience with DevOps & Platform Tools:
o Harness (CI/CD and deployments)
o Aptakube (Kubernetes monitoring and management)
o Akamai (CDN and edge services)
- Experience with Observability Tools:
o Splunk o Datadog
✅ Soft Skills
- Strong communication and customer handling skills
- Problem-solving and analytical thinking
- Ability to work in shifts / 24x7 support environment
- Time management and prioritization
? Typical Day-to-Day Activities
- Logging in and reviewing ticket queues
- Monitoring alerts from Datadog / Splunk dashboards
- Checking Harness pipelines and deployment status
- Monitoring Kubernetes clusters via Aptakube
- Investigating CDN-related issues in Akamai
- Prioritizing incidents and responding to users
- Resolving L1/L2 issues or escalating
- Performing daily system/service checks
- Monitoring SLAs and updating dashboards
- Attending standups / shift handovers
- Documenting resolutions and closing tickets
? KPIs / Success Metrics
- SLA compliance %
- Ticket resolution rate
- Customer satisfaction (CSAT)
- Backlog reduction
- First-time resolution rate
- Alert response and resolution time
- Deployment success rate (Harness)
- System uptime and availability
- Reduction in recurring incidents
📌 AWS Production Support Engineer (Guadalajara)
🏢 Tekshapers
📍 Guadalajara