Site Reliability Engineer (Onsite, Lahore, PKR Salary)

HR POD Careers


Date: 2 hours ago
City: Lahore
Contract type: Full time
Requirements:

  • 5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale.
  • Bachelor's degree in Computer Science.
  • Deep Linux troubleshooting skills across the OS, networking, storage, and performance, with hands-on experience working as root on production systems.
  • Experience with Ubuntu is highly relevant, as it is used almost exclusively.
  • Hands-on experience operating GPU servers in production, including troubleshooting driver, device, and hardware-level issues, rather than only the workloads running on top of them.
  • Practical network troubleshooting experience, including diagnosing physical-layer faults.
  • Strong automation mindset with programming skills in Python or a comparable language.
  • Experience with configuration management, node provisioning, and infrastructure-as-code (IaC) using Ansible, Terraform, or similar tools.
  • Experience building observability and alerting solutions using Grafana and Prometheus.
  • Experience operating GPU clusters or AI infrastructure at production scale.
  • Production experience with Kubernetes or Slurm; experience with both is a bonus.
  • Background in HPC or research computing.
  • Familiarity with the NVIDIA GPU stack, InfiniBand/RDMA, and NCCL.
  • Experience with CLI-based AI coding agents such as Claude Code, rather than browser-based assistants alone.
  • Contributions to open-source projects within the cloud-native, HPC, or AI infrastructure ecosystem.


Responsibilities:

  • Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage.
  • Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues.
  • Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
  • Diagnose network faults end to end, including configuration, routing, and physical-layer issues such as cabling, transceivers, and link errors.
  • Configure and maintain workload managers that schedule customer jobs, including Kubernetes, Slurm, or both, along with the identity, storage, and networking services they depend on.
  • Build automation and tooling to eliminate operational toil, using a modern language such as Python or Go to design solutions, review implementations, and redirect approaches when needed.
  • Use AI-assisted engineering tools such as Claude to accelerate automation, runbook development, and incident analysis.
  • Automate provisioning, image deployment, configuration, and remediation using Ansible and infrastructure-as-code.
  • Design and operate observability using Grafana, Prometheus, and Loki, while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
  • Lead incident response, on-call activities, blameless postmortems, and reliability improvements that maintain customer SLAs.
  • Partner with Platform and Systems Engineering teams on capacity planning, rollouts, and continuous improvement.


Working Hours:

8 PM - 4 AM

How to apply

To apply for this job you need to authorize on our website. If you don't have an account yet, please register.

Post a resume

Similar jobs

Product Manager - PostEx

Taraki, Lahore
1 day ago
Our client PostEx is hiring a Product Manager in Lahore.About PostExPostEx is a technology-driven FinTech and logistics company building innovative products and solutions that simplify payments, commerce, and delivery experiences. We are looking for a Product Manager who can combine strong product ownership with data-driven research and product analysis to build solutions that deliver measurable business value.Position OverviewWe are seeking...

Fin-Crime Investigation & Reporting Officer

ACE Money Transfer, Lahore
6 days ago
Job DescriptionAs a Fin-Crime Investigation & Reporting Officer y will be responsible for data handling and analytics, with a strong focus on case management and investigations. The role involves monitoring transactions, analyzing data for potentially fraudulent activity, and maintaining accurate records to support appropriate compliance actions.Responsibilities:Monitor and analyze transactions for potentially fraudulent activity using strong analytical skillsConduct thorough investigations of...

Auditing Executive (Compliance)

Cedar Global Solutions, Lahore
6 days ago
Job detailsLocation: The Enterprise Building Near Thokar Niaz Baig LahoreMode: OnsiteTimings: 3pm to 12amCompensation: 130K - 150K (based on experience) Responsibilities  Develop and maintain audit programs aligned with regulatory requirements and organizational policies. Conduct compliance audits to assess adherence to laws, regulations, and internal policies. Document audit findings, including deficiencies and areas of noncompliance. Communicate audit results to stakeholders such...