The Role
What You’ll Do
- Application performance tuning, outage troubleshooting / triage, short term / long-term fixes
- Work closely with development/architecture replace and reduce the footprint and cost of infrastructure on the cloud by migrating to our on-premise data centers
- Centralize logging in production environments using Graylog
- Ensure design, infrastructure, processes, and operational documentation are up to date
- Automate repetitive tasks, leaning out processes that drive down outages and deliver quick results
- Familiarity with Public & Private Cloud (AWS, Azure, GCP), and on-prem VMware virtualized systems environments.
- Networking basic knowledge, (F5, HAproxy)
- Ability to drive own bodies of work using Atlassian stack
What you will bring:
- Data-Driven & Communication: You’ll need strong communication skills to hold technical conversations in English, asking relevant questions to guide problem-solving processes.
- Problem Solving: You should possess strong critical thinking skills to navigate unfamiliar situations, providing clear, step-by-step troubleshooting guidance. You’ll lead on calls, mentor team members, and prioritize issues by evaluating their impact, urgency, and resource requirements.
- Application Services and Infrastructure: You’ll have to be capable of navigating through the IT infrastructure,
regardless of mínimal background on its architecture leveraging your previous experience
- Installation and Configuration of Software: Proficiency in both Windows and Linux environments, adept at multiple protocols/tools (HTTP, HTTPS, TLS, TCP, UDP, SMTP, ICMP, FTP, sFTP, SSH, RDP) expertise in technologies/practices like (VPN, Jumphost, DHCP, Active Directory) webservers, F5, HAProxy, and Apache, Nginx, IIS, Websphere, VMware, Git.
- Monitoring Experience: You’ll have experience setting up proactive monitoring systems to prevent recurring outages, identifying necessary configurations, and proposing short
- and long-term solutions to post-incident scenarios.
- Understanding of Application SDLC and On-Call Support: You’ll provide on-call support on a rotational basis, diagnosing, triaging, and fixing system issues during outages. Bonus: Knowledge of source control management and CI/CD automation processes (Ansible) is highly valued.
- Intellectual Curiosity and Can-Do Attitude: We’re looking for someone eager to contribute with contagious enthusiasm. In times of uncertainty, you’ll help clarify chaos, engage in outages, projects, and specialists when necessary, and bring thoughtful, clear-headed solutions.
Bonus Points: Hands-On experience with any Message broker technology such as RabbitMQ, Kafka, IBM-MQ, ActiveMQ, Tibco, MSMQ, RedHat AMQ, SQS, Redis, Graylog, LDAP, AWS, Azure, VMWare
LI-JG1
📌 Senior Application Services Engineer (México)
🏢 Solera
📍 México