Engineering

Site Reliability Engineer At Cowrywise

Cowrywise·Lagos, nigeria·Full Time·Internship
EngineeringFull TimeInternship

We're looking for a Site Reliability Engineer (SRE) to help build, maintain, and scale the infrastructure powering Cowrywise.You'll work closely with our engineering team to improve reliability, observability, security, and deployment processes across our systems.

Our infrastructure team specializes across four areas: Cloud, Databases, Platform, and Observability. We run primarily on AWS with some workloads on GCP. For this role, we're particularly interested in someone who can raise the bar on observability, helping us detect issues faster and resolve them with confidence.

What you'll do

Generally, members of the infrastructure team are able to do the following

  • Design, maintain, and improve cloud infrastructure and internal platforms
  • Improve system reliability, scalability, and performance across services
  • Build and maintain CI/CD pipelines and deployment workflows
  • Implement monitoring, logging, alerting, and observability systems
  • Respond to incidents, troubleshoot production issues, and lead root cause analysis
  • Automate operational tasks and infrastructure provisioning
  • Work with engineering teams to improve service architecture and operational readiness
  • Improve security posture, access controls, and infrastructure best practices
  • Manage containerized workloads and orchestration platforms
  • Maintain disaster recovery, backup, and high availability strategies

What we're looking for

Required

  • 4+ years of experience in an SRE, DevOps, or Platform Engineering role running production systems
  • Strong hands-on experience with AWS (compute, networking, IAM, storage, managed services)
  • Deep expertise in observability designing meaningful metrics, dashboards, alerts, and SLOs that actually catch problems before users do
  • Hands-on experience with New Relic, Grafana, and Prometheus (or equivalent tooling)
  • A track record of reducing MTTD and MTTR through better instrumentation, alerting, and incident response practices
  • Proficiency with Docker and containerized workflows
  • Solid scripting and automation skills (Python, Bash, Go, or similar)
  • Experience with infrastructure-as-code (Terraform, Pulumi, or CloudFormation)
  • Strong Linux fundamentals and networking knowledge
  • Experience building and maintaining CI/CD pipelines
  • Comfort leading incident response and writing clear post-mortems

Nice to have

  • Experience operating Kubernetes in production
  • Exposure to GCP or multi-cloud environments
  • Background in one of our specialization areas: Databases (Postgres, MySQL, Redis), Platform engineering, or Cloud architecture
  • Security-focused experience (IAM hardening, secrets management, compliance frameworks)
  • Experience in fintech or other regulated, high-availability environments

The people who succeed on this team

  • People who are proactive and take ownership
  • Engineers who automate before repeating manual work
  • People who stay calm and methodical during incidents
  • Engineers who care about clean systems and operational excellence
  • Strong collaborators who work well across teams
  • Curious builders who enjoy learning and improving systems continuously

Key skills

BA/BSc/HND

At a glance

Company

Cowrywise

Location

Lagos, nigeria

Employment

Full Time

Experience

Valid until

Not specified

Created

May 23, 2026

More opportunities

Similar roles you might like

Site Reliability Engineer At Heirs Insurance Ltd

Heirs Insurance Ltd

Lagos, nigeria

full-time

About the job The Site Reliability Engineer is the hands-on operator and automation builder who keeps the platform running. Reporting to the SRE Lead, you will be responsible for the day-to-day health of the platform --- monitoring, alerting, incident response, and the continuous automation work that makes the platform easier to operate and harder to break. You will work closely with the DevSecOps engineers, cloud engineers, and security team to ensure that reliability is built into every layer of the platform, not bolted on afterwards. This is a demanding role on a high-stakes platform --- and a rare opportunity to do SRE work that genuinely matters at scale. What You'll Do NSOC Operations & Monitoring --- Work as part of the 24/7 Network & Security Operations Centre (NSOC) --- actively monitoring the platform's health across all services, cloud environments, and network layers. Triage incoming alerts, distinguish signal from noise, escalate genuine incidents promptly, and maintain the situational awareness that keeps the team ahead of problems before they become outages. Incident Response & Triage --- Serve as a first responder for platform incidents --- detecting, triaging, and containing issues in real time. Execute runbooks under pressure, coordinate with engineering teams during active incidents, and own your part of the resolution. After every significant incident, contribute to the blameless post-mortem process --- documenting what happened, what worked, and what needs to change. Observability Implementation & Maintenance --- Implement, maintain, and continuously improve the platform's observability stack --- including centralised logging, metrics collection (Prometheus, Datadog, or equivalent), distributed tracing (Jaeger, OpenTelemetry, or equivalent), and alerting rules. Instrument new services from day one, tune alerting thresholds to reduce noise without missing real issues, and build dashboards that give the team the operational visibility they need. Automation & Toil Reduction --- Identify repetitive operational tasks and build the automation that eliminates them. Write scripts, tools, and workflows in Bash, Python, or equivalent to reduce toil --- replacing manual processes with reliable, repeatable automation. Toil that exists today should not exist next quarter. This is an ongoing discipline, not a project with an end date. Reliability & Chaos Testing --- Support the SRE Lead in running the platform's reliability engineering programme --- executing load tests, stress tests, and chaos engineering experiments to validate the platform's resilience under real conditions. Document findings, flag vulnerabilities, and work with engineering teams to address weaknesses before they become production incidents. Runbook Development & Maintenance --- Write, test, and maintain operational runbooks for all known failure modes and recovery procedures. A runbook is only valuable if it works --- regularly test each runbook against real or simulated conditions, update it when the platform changes, and make sure every procedure is accurate and actionable. Runbooks are not documentation --- they are operational tools that must be ready to use at 2am. SLO Monitoring & Reporting --- Monitor SLO attainment and error budget consumption across all platform services on an ongoing basis. Proactively flag services approaching error budget limits to the SRE Lead and relevant engineering teams. Track and report DORA metrics --- deployment frequency, lead time, change failure rate, and MTTR --- as part of the team's ongoing performance visibility. Platform Health & Capacity Monitoring --- Continuously monitor resource utilisation, service performance, and infrastructure health across both cloud environments (AWS and Azure). Identify trends that suggest emerging capacity constraints or degrading performance, and surface these findings to the SRE Lead and Platform Manager before they become incidents. CI/CD Pipeline Health --- Monitor the health and performance of the platform's CI/CD pipelines in collaboration with the DevSecOps team. Identify pipeline failures, flaky tests, or degraded build performance and escalate promptly. Reliable pipelines are part of platform reliability --- treat them accordingly. Documentation & Knowledge Sharing --- Maintain accurate, up-to-date operational documentation --- procedures, architecture notes, incident timelines, and lessons learned. Share knowledge actively with your SRE colleagues and across the broader engineering team. A well-documented platform is a more reliable platform. What We're Looking For Must Have 2+ years of SRE, DevOps, or platform operations experience in a production environment --- with hands-on responsibility for monitoring, alerting, and incident response. Hands-on experience with observability tooling --- logging (ELK, Loki, Datadog, or equivalent), metrics (Prometheus, Datadog, or equivalent), and alerting. You have built dashboards and tuned alert rules, not just viewed them. Experience responding to production incidents --- triaging, containing, and resolving issues under pressure. You understand incident severity classification and escalation paths. Scripting skills in Bash and Python (or equivalent) --- for writing automation, operational tooling, and toil-reduction scripts. Solid understanding of cloud infrastructure on AWS and/or Azure --- compute, managed services, networking, and the platform-level health indicators you should be monitoring. Genuine care about reliability --- you understand why SLOs and error budgets matter, and you approach operational work with the discipline that high-availability financial services demand. Nice to Have Experience with distributed tracing --- Jaeger, OpenTelemetry, AWS X-Ray, or equivalent --- and using trace data to diagnose performance and reliability issues. Exposure to chaos engineering tools --- Chaos Monkey, Gremlin, or equivalent --- or experience running load and stress tests in production or pre-production environments. Familiarity with Kubernetes operational patterns --- pod health, cluster observability, resource constraints, and Kubernetes-native alerting and scaling behaviours. Experience in a regulated environment --- fintech, banking, or similar --- with an understanding of the uptime, audit, and compliance expectations financial services carry. AWS and/or Microsoft Azure certification --- SysOps, Cloud Practitioner, or equivalent entry-to-mid-level cloud certifications. Familiarity with DORA metrics and what they tell you about engineering team health and deployment risk.

10 days ago

Senior Devops Engineer At Paystack

Paystack

Lagos, nigeria

full-time

About the Senior DevOps Engineer role As a Senior DevOps Engineer at Paystack, you will play a pivotal role in designing, implementing, and maintaining our cloud infrastructure, CI/CD pipelines, and automation frameworks. You will collaborate closely with cross-functional teams to optimise our development and deployment processes, improve system reliability, and drive continuous improvement initiatives. The ideal candidate will have a strong background in DevOps practices, cloud technologies, automation tools, database management, networking, and a passion for solving complex technical challenges. The selected engineer will join a team of DevOps engineers and report directly to the Lead DevOps Engineer. You will work closely with other members of the department including members of the larger infrastructure and security teams and quite heavily with other product engineering teams. We'll trust you to Own projects end to end, often across teams. Design technical solutions that are scalable and maintainable Design systems for high availability and fault tolerance. Build end to end automation pipelines. Tackle ambiguous, high risk problems proactively. Introduce new ideas and simplifies complex systems. Design secure systems, manages secrets and implements zero trust models. Participate in architectural evolution and software development activities. Help debug and diagnose issues that may arise in an intertwined distributed system environment. Own CI/CD architecture, minimises downtime and ensures security Lead major incident response and postmortems. You'll thrive in this role if you have 5+ years of experience in a DevOps or Site Reliability Engineering (SRE) role. Deep Expertise with a variety of AWS services Experience with IAC and tools like Terraform, designing reusable modules and enforces standards Real world experience with Kubernetes at scale Experience writing scripts and automation using Python, Node.js, Bash, Golang Experience with networking concepts and protocols, including TCP/IP, DNS, HTTP/S, VPNs, etc. Experience building tooling to reduce occurrences of errors and improve engineering productivity Experience with CI/CD tooling like ArgoCD, Github Actions, etc. Experience working with and monitoring and alerting tools like Prometheus and Grafana. Ability to clearly articulate, implement and document design decisions Bonus points if you have: Bachelor's degree in Computer Science, Engineering, or a related field might be beneficial but not required Experience in the fintech industry or with payment processing systems Knowledge of cybersecurity best practices and compliance requirements (e.g., PCI DSS) Experience with machine learning and data analytics. Benefits Competitive compensation package and benefits 13th month bonus TSG Equity compensation Full medical coverage Wellbeing stipend Generous leave and sabbatical policies Hybrid working environment Smart, kind colleagues who are invested in your growth.

10 days ago

Devops & Cloud Engineer At Corona Management Systems

Corona Management Systems

Abuja, nigeria

full-time

Job Description We're looking for a skilled DevOps and Cloud Engineer to help design, build, and maintain the infrastructure that powers our products. You'll work at the intersection of development and operations, automating processes, optimizing cloud environments, and ensuring our systems are reliable, secure, and scalable. This is a great opportunity for someone who enjoys solving complex infrastructure challenges and wants to have a direct impact on how we build and ship software. What You'll Do Design, implement, and maintain scalable, secure cloud infrastructure (AWS/Azure/GCP) Build and improve CI/CD pipelines to streamline deployment processes Automate infrastructure provisioning using Infrastructure as Code (Terraform, CloudFormation, or similar) Monitor system performance, availability, and reliability; respond to incidents and drive root-cause analysis Implement and maintain containerization and orchestration solutions (Docker, Kubernetes) Collaborate with software engineers to optimize application deployment and performance Strengthen security practices across infrastructure, including access controls, secrets management, and vulnerability remediation Manage configuration and environment consistency across development, staging, and production Contribute to on-call rotation and support incident response Document infrastructure, processes, and runbooks for team-wide visibility Requirements 3+ years of experience in DevOps, Site Reliability Engineering, or Cloud Engineering roles Hands-on experience with a major cloud provider (AWS, Azure, or GCP) Proficiency with Infrastructure as Code tools (Terraform, CloudFormation, Pulumi, etc.) Experience with containerization and orchestration (Docker, Kubernetes) Strong scripting skills (Python, Bash, or Go) Familiarity with CI/CD tools (Jenkins, GitLab CI, GitHub Actions, CircleCI, etc.) Solid understanding of networking, security, and system administration fundamentals Experience with monitoring and logging tools (Prometheus, Grafana, Datadog, ELK, etc.) Nice to Have: Relevant cloud certifications (AWS Certified Solutions Architect, Azure Administrator, etc.) Experience with GitOps workflows (ArgoCD, Flux) Familiarity with configuration management tools (Ansible, Chef, Puppet) Experience working in regulated or high-compliance environments Background in cost optimization and cloud governance

20 days ago

Engineering Manager (site Reliability) At Moniepoint Inc.

Moniepoint Inc.

Lagos, nigeria

full-time

About the role As an Engineering Manager, you will drive the successful delivery and execution of projects within your teams. You will manage end-to-end technical planning, ensuring that product requirements are translated into actionable tasks, while orchestrating collaboration between various stakeholders, including engineers, product managers, QA, and UX. This role requires a deep understanding of software design and development and the ability to plan, execute, and deliver product features in a timely and predictable manner. You will also be responsible for maintaining high technical standards, managing team bandwidth, and ensuring project milestones are met with efficiency and accuracy. What you'll get to do Own delivery and execution within the Site Reliability Engineering tooling team. Evaluate product requirements for feasibility and ensure they align with the existing product architecture, translating them into EPICs and technical stories. Work closely with Product Managers and Engineers to refine and groom tasks. Plan and organize sprints with clearly defined goals, using project planning tools to establish timelines, and delivery milestones, and identify task dependencies early. Foster engineering processes that promote seamless collaboration and teamwork. Track team velocity to ensure resources are effectively allocated, balancing bandwidth with task demands. Coordinate alignment and manage dependencies across multiple stakeholders to prevent bottlenecks and ensure smooth execution. Contribute to critical projects by ensuring appropriate design patterns and coding techniques are applied. Remain hands-on, participating in code reviews to uphold high-quality standards. Ensure monitoring and observability are in place for all owned services, meeting defined SLIs/SLOs. Partner with Product Managers to track and publish post-deployment product metrics, ensuring transparency with key stakeholders. To succeed in this role, you should have BSc in Computer Science, Engineering, or a related field Minimum of 8 years of experience as a Software Developer, Software Engineer, or similar role. 5+ years of Node JS / python experience. A minimum of 3 years of leadership experience is a must. Strong understanding of agile methodologies, sprint planning, and backlog management. Expertise in breaking down complex product requirements into structured EPICs, Stories, and Tasks. Solid experience with backend technologies. Experience with frontend is a plus. Knowledge of project planning tools for visualizing and tracking delivery timelines. Familiarity with engineering metrics and monitoring tools to assess team performance and product health. Capability to debug complex technical issues during incidents to identify solutions and run blameless RCA sessions. Understanding of deployment pipelines, continuous integration (CI), continuous deployment (CD), and their corresponding metrics. Ability to drive alignment across diverse technical and non-technical stakeholders. Exceptional ability to manage dependencies, mitigate risks, and communicate clearly with stakeholders. Proven track record of improving team velocity and fostering efficient delivery. Generic Skills: Problem-solving: Ability to assess complex problems, find solutions, and make sound decisions. Communication: Strong written and verbal communication skills, including technical documentation and stakeholder reporting. Adaptability: Able to thrive in a fast-paced, changing environment, adjusting strategies as needed. Attention to Detail: Meticulous in documenting technical requirements and ensuring all aspects of a project are accounted for. Supervisory skills: Team Management: Experience in managing and mentoring engineers, ensuring team growth and performance. Resource Allocation: Ability to assess bandwidth and manage resource distribution to optimize team performance. Feedback: Conduct regular performance reviews, providing constructive feedback and fostering a growth-oriented environment. Stakeholder Management: Lead project status reviews, manage expectations, and ensure smooth communication between teams and leadership. What we can offer you Culture - We put our people first and prioritize the well-being of every team member. We've built a company where all opinions carry weight and where all voices are heard. We value and respect each other and always look out for one another. Above all, we are human. Learning - We have a learning and development-focused environment with an emphasis on knowledge sharing, training, and regular internal technical talks. Compensation - You'll receive an attractive salary, pension, health insurance, paid leave, plus other benefits.

a month ago

Site Reliability Engineer At Cowrywise
Lagos, nigeria
Apply