DevOps & Site Reliability Engineer (GCP)

Qureos


Date: 3 weeks ago
City: Remote
Contract type: Full time
Remote

As a DevOps Engineer, you will design, implement, and maintain the infrastructure that supports our applications and services — and, just as importantly, you will keep that infrastructure reliable in production. You will work closely with development, QA, and IT teams to automate and streamline operations, build the observability that lets us catch problems before customers do, and lead the response when incidents happen. This is a hands-on role where you own the health of production, not only its build-out. We are hiring at a mid level (3–5 years) for someone with strong production instincts and the judgment to grow into a senior reliability owner as we scale globally.

Key Responsibilities
Infrastructure Management
  • Design, deploy, and manage cloud infrastructure on Google Cloud Platform (GCP).
  • Implement and maintain scalable container orchestration using Docker.

CI/CD Pipeline
  • Develop and maintain continuous integration/continuous deployment (CI/CD) pipelines to automate testing, building, and deployment processes through GitHub and Google Cloud Run.
  • Collaborate with development teams to integrate new features into the CI/CD process.

Monitoring & Observability
  • Set up and manage monitoring, logging, and alerting systems using tools like Elasticsearch, Kibana, Prometheus, and New Relic.
  • Build dashboards, metrics, and actionable alerts so engineering detects degradations before customers do — and so no critical signal (e.g. a database running hot for hours) ever goes unnoticed.

Reliability & Incident Response (SRE)
  • Define, measure, and own service-level objectives (SLOs/SLIs) and error budgets for critical services.
  • Lead incident response end to end: triage by severity, form a hypothesis and confirm it with metrics before taking action, mitigate, and drive to resolution. Act as incident commander on major incidents and coordinate communication across stakeholders.
  • Run blameless postmortems, identify true root causes, and track corrective actions to closure.
  • Participate in an on-call rotation; continuously reduce toil and mean-time-to-recovery (MTTR) through automation.
  • Plan for capacity, performance, and cost as we scale toward multi-region, global traffic.

Database Management
  • Administer and optimize MongoDB instances, ensuring data integrity, performance, and security.
  • Implement backup, recovery, and disaster recovery strategies for MongoDB and other databases.

Security & Compliance
  • Implement security best practices across infrastructure, applications, and data.
  • Ensure compliance with industry standards and internal policies.

Automation & Scripting
  • Automate infrastructure provisioning, configuration management, and system operations using GCP services.
  • Develop custom scripts as needed to enhance automation and operational efficiency.

Collaboration & Support
  • Work closely with development and QA teams to support the software development lifecycle.
  • Provide technical guidance and support to resolve infrastructure-related issues.


Qualifications
Education
  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent experience).

Experience
  • 3–5+ years of experience as a DevOps Engineer or in a similar role.
  • Strong experience with Docker, including container orchestration using Kubernetes.
  • Hands-on experience with Google Cloud Platform (GCP) services, including Compute Engine, Cloud Storage, Cloud Run, and GKE.
  • Experience with Elasticsearch for monitoring, logging, and search.
  • Proficiency in administering and optimizing MongoDB databases.
  • Demonstrated experience operating production systems and responding to incidents — not only building infrastructure. You can point to real production incidents you owned, quantify their impact concretely (users affected, duration, consequence), and describe what you changed afterward.

Skills
  • Strong scripting skills in Python, Bash, or similar languages.
  • Proficiency in infrastructure as code (IaC) tools such as Terraform or Ansible.
  • Experience with CI/CD tools such as GitHub Actions or CircleCI.
  • Sound production-debugging methodology: observe and form a hypothesis before acting, reason about how components fail together (e.g. how a queue backlog interacts with the database), and update your approach when new information appears.
  • Familiarity with defining alerts, dashboards, and SLOs; comfort reasoning about availability, latency, saturation, and error budgets.
  • Excellent problem-solving and troubleshooting skills.
  • Strong communication skills and ability to work collaboratively across teams.
  • Strong ownership and a blameless, collaborative posture under pressure — focused on diagnosing and fixing problems rather than assigning blame.

Nice to Have
  • Experience with additional cloud platforms (e.g. AWS or Azure).
  • Familiarity with Agile/Scrum methodologies.
  • Certifications in Docker, GCP, or related technologies.

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.

Post a resume

Similar jobs

Assistant Manager System Admin

K-Electric, Remote
12 hours ago
PURPOSE Perform maintenance activities for upkeep of Oracle database, Source Database (SDB), monitor performance tuning and database sizing. Create Management Information System (MIS) reports as per requirement and monitor availability of Source Databases & Historical Information System (HIS) database WITH the objective of smooth and proper functioning of Database System WITHIN the limits of established company policies and procedures, departmental...

TECHNICAL SALES SPECIALIST

Zelle Research & Analytical Services, Remote
12 hours ago
Location: Chennai DESIGNATION: Regional Sales Manager - Chennai Role Definition/PurposeThe Regional Sales Manager is responsible for managing the daily and long-term operations of a company across a geographic region. As a Regional Sales Manager you will often be responsible for setting and adjusting sales goals based on deep knowledge of individual’s selling patterns. They ensure that each team reaches its...

QA Engineer (Manual & Automaation) Minimum 4 years of experience

ASA Technologies, Remote
2 days ago
Our Client is a U.K based start up specifically in the domain of Legal Practice Management Software for LAW Firms. They are looking to hire a QA Engineer (Manual & Automation) with with great skillset having minimum 4 years of experience. Following are the other details;About the RoleWe are looking for an experienced QA Engineer with at least 4 years...