About the role

We are looking for a Senior Cloud/DevOps Engineer to operate and improve the cloud infrastructure and reliability layers behind an enterprise data platform in a regulated healthcare environment.

The mandatory requirements are 5+ years of experience in Cloud Engineering, DevOps, or Site Reliability Engineering, advanced hands-on experience with AWS EKS and Kubernetes, experience administering and troubleshooting Argo Workflows, and strong English communication skills.

Must haves

  • 5+ years of professional experience in Cloud Engineering, DevOps or Site Reliability Engineering.
  • Strong hands-on experience operating AWS infrastructure in production environments.
  • Advanced experience with Kubernetes and Amazon EKS, including workload operations, troubleshooting, access, observability, capacity, and reliability.
  • Hands-on experience administering and troubleshooting Argo Workflows or comparable workflow orchestration platforms.
  • Strong Infrastructure as Code experience with Terraform and source-controlled infrastructure practices.
  • Experience building, hardening, and supporting CI/CD pipelines and production release processes.
  • Strong experience with monitoring, logging, alerting, and incident-routing tools such as Splunk, PagerDuty, Opsgenie, or comparable platforms.
  • Demonstrated ability to lead complex incident resolution, perform root-cause analysis, and translate findings into preventive improvements.
  • Proficiency in automation and scripting using Python, Shell, Bash, or similar languages.
  • Ability to make well-reasoned technical decisions, identify tradeoffs, estimate work, and drive improvements across a complex platform.
  • Experience mentoring engineers and collaborating effectively with Data Engineering, Security, Governance, Analytics, and business stakeholders.
  • Strong written and verbal English communication skills, with the ability to work directly with client stakeholders.
  • Availability to work within the LatAm service window of approximately 9:00 AM to 6:00 PM Eastern Time and participate in an agreed on-call rotation.

Nice to haves

  • Experience supporting data-platform infrastructure involving Snowflake, dbt, Fivetran, HVR, Tableau Cloud, or custom ingestion pipelines.
  • Familiarity with data-specific observability platforms such as SYNQ.
  • Experience modernizing or migrating legacy orchestration and ingestion solutions such as Boomi or AWS Data Pipeline.
  • Experience with service-management and change-control tools such as Freshservice and Jira.
  • Experience operating in healthcare, life sciences, financial services, or another regulated environment.
  • Familiarity with HIPAA, GDPR, FDA-related controls, least-privilege access, separation of duties, and audit-ready operational practices.

What you will do

  • Provide senior technical ownership for the Cloud / DevOps service tower during the LatAm coverage window, including day-to-day operations, complex troubleshooting, and L2/L3 escalation.
  • Operate, maintain, and improve AWS infrastructure supporting the Data Platform, including Amazon EKS, S3, EventBridge, SQS, API Gateway, Lambda, and related services.
  • Administer Kubernetes-hosted workloads and Argo Workflows, including deployment, scheduling, monitoring, troubleshooting, capacity management, resiliency, and recovery.
  • Define and improve standards for Infrastructure as Code, configuration management, CI/CD, release execution, rollback, and environment consistency, primarily using Terraform and Git-based delivery practices.
  • Lead the consolidation and improvement of observability across infrastructure and data workloads, linking alerts to operational evidence from Argo, dbt, Snowflake, and supporting runbooks.
  • Improve alert routing and escalation workflows across tools such as Splunk, Opsgenie, PagerDuty, Microsoft Teams, and data-specific observability platforms.
  • Lead or support major incident response, root-cause analysis, post-incident reviews, and corrective actions, with clear communication to technical and service stakeholders.
  • Design and implement reliability improvements such as selective auto-remediation, dependency-aware alert correlation, impact analysis, and automation of repetitive operational work.
  • Track and contribute to service metrics including availability, SLA compliance, alert volumes, workflow reliability, deployment outcomes, and mean time to restore service.
  • Apply disciplined change-management, access-control, secrets-management, auditability, and documentation practices appropriate for a HIPAA-, GDPR-, and FDA-regulated environment.
  • Create and maintain runbooks, operating procedures, architecture context, recovery procedures, and knowledge-transfer materials.
  • Mentor Middle-level engineers, review technical work, improve team practices, and promote consistent execution across the distributed team.
  • Participate in the Cloud / DevOps on-call rotation for critical incidents outside staffed service hours.

Perks

  • Growth without limits: build your skills through mentorship, internal TechTalks, challenging projects, and a dedicated annual learning budget
  • Competitive compensation: get recognition that reflects your skills and impact, with regular performance and compensation reviews
  • Flexibility: work 100% remotely with flexible hours that support focus, autonomy, and a healthy work rhythm
  • Meaningful, modern projects: build impactful products using modern technologies alongside global teams and leading brands
  • Collaborative culture: join a supportive environment with zero micromanagement where ideas are welcomed and contributions are recognized
  • Well-being & support: access local well-being programs and people-focused support tailored to your location

Didn’t find the perfect fit?

Subscribe to get notified about new roles that match your skills and interests.

Get role alerts

Hiring process timeline

We designed our hiring process to be fast, transparent, and convenient — so you always know what to expect and can move through the steps without unnecessary delays.

lightning-iconA quick note

After applying, keep an eye on your inbox — including your spam folder. Occasionally, emails from our hiring platform, LaunchPod, may land there.

1

Tell us about yourself

Submit a short application form.

2

Pass a quick test

Pass a 30’- 60’ test.

3

Record a short video

Introduce yourself on video to fast-track your application process.

4

Meet our team

Once you pass the video review, grab a spot on our calendar for a technical interview.

5

Get an offer

Welcome to AgileEngine!

FAQ

Have any questions?

What is the work format and schedule?

We are a remote-first company, so all our positions are remote. Schedules are flexible, but the main requirement is to have an overlap with your client’s schedule to ensure smooth collaboration. Specific details depend on the project and are discussed during the interview process.

What level of English is required?

We look for Upper-Intermediate (B2) proficiency or higher. Since you will be collaborating with international teams and global clients (including Fortune 500 companies), English is a part of your daily work.

Does AgileEngine provide work equipment?

It depends on your location. We provide equipment in Ukraine, Poland, Argentina, Colombia, Mexico, Brazil, Guatemala, Portugal, Spain, and India. In the USA, equipment is usually provided by the client; otherwise, you will be expected to use your own setup.

What opportunities for professional growth do you offer?

We support continuous growth through personalized development paths tailored to each expert. Also, you can expect:

  • An annual learning and development budget
  • Internal workshops, tech talks, and mentorship programs
  • Opportunities to switch projects or grow into new roles over time

What does a typical team look like, and what tools do you use?

Teams vary by project, but you’ll usually work with a mix of experts across different fields, along with a delivery or project manager who supports collaboration and onboarding.

For day-to-day work, teams commonly use tools like Jira and Google Workspace, along with communication platforms such as Google Chat. Depending on the client, you may also work with Slack, Confluence, Notion, or similar tools.

Need help?

Want more details? See full FAQ

If something doesn’t work or you have questions regarding your application process, drop us a line.

    Scroll to Top