Senior Technical Operations Engineer
Support and improve cloud infrastructure, operating systems, middleware, production environments, and enterprise technology services. The role focuses on system reliability, security maintenance, incident resolution, automation, and disaster recovery while partnering with engineering teams and third
Responsabilidades
- Install, configure, administer, and maintain servers, operating systems, middleware, cloud infrastructure, and internal applications
- Monitor system performance, capacity, production environments, error logs, dashboards, and batch processing
- Manage system configurations, backups, patches, version upgrades, and security improvements
- Administer identity and access management privileges and monitor access activity
- Configure and support workload automation, job scheduling, dependencies, and batch operations
- Support the full incident management lifecycle, including triage, escalation, investigation, resolution, and incident reviews
- Resolve critical customer and production system issues in collaboration with engineering teams and third-party vendors
- Implement backup, restore, disaster recovery, and recovery testing processes
- Support cloud infrastructure and independently implement continuous integration and deployment pipelines
- Maintain infrastructure, configuration, operational, disaster recovery, and technical process documentation
- Generate performance and incident reports and communicate technical information to varied stakeholders
- Identify and recommend improvements to operational processes, workflows, reliability, and service effectiveness
Requisitos
- Experience administering operating systems, middleware, servers, cloud infrastructure, and software environments
- Ability to configure systems, manage backups, patches, upgrades, and security maintenance
- Experience monitoring production environments, logs, dashboards, service capacity, and batch processes
- Knowledge of identity and access management, privileged account security, and secrets integrity
- Experience supporting incident management, escalations, root-cause analysis, and corrective action plans
- Working knowledge of disaster recovery, backup and restore processes, and recovery drills
- Ability to implement or support CI/CD pipelines and workload automation tools
- Ability to communicate technical information to technical and nontechnical stakeholders
- Ability to manage work independently, troubleshoot complex issues, and collaborate across teams and vendors
Beneficios
- Medical, life insurance, and retirement benefits
- Flexible benefits options
- Employee volunteer programs
- Disability accommodations and equal employment opportunity