Observability/Platform SRE Specialist job at Yas Tanzania
New
Today
Linkedid Twitter Share on facebook
Observability/Platform SRE Specialist
2026-09-19T11:20:42+00:00
Yas Tanzania
https://cdn.greattanzaniajobs.com/jsjobsdata/data/employer/comp_6055/logo/download%20(7).png
FULL_TIME
Dar es Salaam
Dar es Salaam
00000
Tanzania
Professional Services
Computer & IT, Science & Engineering
TZS
MONTH
2026-09-25T17:00:00+00:00
8

Company: Mixx by Yas

Position: Observability/Platform SRE Specialist

QUALIFICATIONS & EXPERIENCE

  • Bachelor's degree in Computer Science, Information Technology, Software Engineering, Telecommunications, or a related field.
  • Relevant SRE, DevOps, Cloud, Kubernetes, or ITIL certifications are an added advantage.
  • Minimum of 5-8 years' experience in SRE, DevOps, Observability, Infrastructure, or Platform Operations within a high-availability environment.

CORE RESPONSIBILITIES

  • Own and govern the observability platform, including Grafana dashboards, alerts, thresholds, and monitoring standards.
  • Manage and optimize Prometheus alerting, ensuring effective monitoring, accurate thresholds, and reduced alert noise.
  • Maintain end-to-end monitoring coverage, ensuring all production services are monitored and gaps are tracked and resolved.
  • Manage the ServiceNow incident taxonomy, keeping classifications aligned with services and operational needs.
  • Lead monitoring onboarding for new services and countries, including metrics, dashboards, alerts, and readiness validation.
  • Ensure monitoring continuity during platform changes and migrations, providing observability sign-off where required.
  • Analyze recurring incidents and trends, identifying whether issues are operational or platform-related.
  • Partner with Engineering and Product teams, providing technical insights, evidence, and recommendations for issue resolution.
  • Validate service restoration after deployments and incidents before operational closure.
  • Track and report platform gaps and improvement actions, ensuring ownership and timely resolution.
  • Conduct regular observability health audits, reviewing alert quality, coverage, taxonomy usage, and improvement opportunities.

CORE COMPETENCIES

  • Expertise in Prometheus, Grafana, alerting, and service monitoring.
  • Strong capability in log analysis, troubleshooting, and incident management.
  • Knowledge of Linux, Docker, Kubernetes, and platform reliability.
  • Experience with Git/GitOps, Python, Bash scripting, and automation.
  • Understanding of SLOs, SLAs, reliability improvement, and alert optimization.
  • Ability to identify risks, monitoring gaps, and drive continuous improvement.
  • Ability to influence, collaborate, and communicate effectively across technical and business teams.
  • Own and govern the observability platform, including Grafana dashboards, alerts, thresholds, and monitoring standards.
  • Manage and optimize Prometheus alerting, ensuring effective monitoring, accurate thresholds, and reduced alert noise.
  • Maintain end-to-end monitoring coverage, ensuring all production services are monitored and gaps are tracked and resolved.
  • Manage the ServiceNow incident taxonomy, keeping classifications aligned with services and operational needs.
  • Lead monitoring onboarding for new services and countries, including metrics, dashboards, alerts, and readiness validation.
  • Ensure monitoring continuity during platform changes and migrations, providing observability sign-off where required.
  • Analyze recurring incidents and trends, identifying whether issues are operational or platform-related.
  • Partner with Engineering and Product teams, providing technical insights, evidence, and recommendations for issue resolution.
  • Validate service restoration after deployments and incidents before operational closure.
  • Track and report platform gaps and improvement actions, ensuring ownership and timely resolution.
  • Conduct regular observability health audits, reviewing alert quality, coverage, taxonomy usage, and improvement opportunities.
  • Prometheus
  • Grafana
  • Alerting
  • Service Monitoring
  • Log Analysis
  • Troubleshooting
  • Incident Management
  • Linux
  • Docker
  • Kubernetes
  • Platform Reliability
  • Git/GitOps
  • Python
  • Bash Scripting
  • Automation
  • SLOs
  • SLAs
  • Reliability Improvement
  • Alert Optimization
  • Bachelor's degree in Computer Science, Information Technology, Software Engineering, Telecommunications, or a related field.
  • Relevant SRE, DevOps, Cloud, Kubernetes, or ITIL certifications are an added advantage.
bachelor degree
60
JOB-6aae700a0231a

Vacancy title:
Observability/Platform SRE Specialist

[Type: FULL_TIME, Industry: Professional Services, Category: Computer & IT, Science & Engineering]

Jobs at:
Yas Tanzania

Deadline of this Job:
Friday, September 25 2026

Duty Station:
Dar es Salaam | Dar es Salaam

Summary
Date Posted: Saturday, September 19 2026, Base Salary: Not Disclosed

Similar Jobs in Tanzania
Learn more about Yas Tanzania
Yas Tanzania jobs in Tanzania

JOB DETAILS:

Company: Mixx by Yas

Position: Observability/Platform SRE Specialist

QUALIFICATIONS & EXPERIENCE

  • Bachelor's degree in Computer Science, Information Technology, Software Engineering, Telecommunications, or a related field.
  • Relevant SRE, DevOps, Cloud, Kubernetes, or ITIL certifications are an added advantage.
  • Minimum of 5-8 years' experience in SRE, DevOps, Observability, Infrastructure, or Platform Operations within a high-availability environment.

CORE RESPONSIBILITIES

  • Own and govern the observability platform, including Grafana dashboards, alerts, thresholds, and monitoring standards.
  • Manage and optimize Prometheus alerting, ensuring effective monitoring, accurate thresholds, and reduced alert noise.
  • Maintain end-to-end monitoring coverage, ensuring all production services are monitored and gaps are tracked and resolved.
  • Manage the ServiceNow incident taxonomy, keeping classifications aligned with services and operational needs.
  • Lead monitoring onboarding for new services and countries, including metrics, dashboards, alerts, and readiness validation.
  • Ensure monitoring continuity during platform changes and migrations, providing observability sign-off where required.
  • Analyze recurring incidents and trends, identifying whether issues are operational or platform-related.
  • Partner with Engineering and Product teams, providing technical insights, evidence, and recommendations for issue resolution.
  • Validate service restoration after deployments and incidents before operational closure.
  • Track and report platform gaps and improvement actions, ensuring ownership and timely resolution.
  • Conduct regular observability health audits, reviewing alert quality, coverage, taxonomy usage, and improvement opportunities.

CORE COMPETENCIES

  • Expertise in Prometheus, Grafana, alerting, and service monitoring.
  • Strong capability in log analysis, troubleshooting, and incident management.
  • Knowledge of Linux, Docker, Kubernetes, and platform reliability.
  • Experience with Git/GitOps, Python, Bash scripting, and automation.
  • Understanding of SLOs, SLAs, reliability improvement, and alert optimization.
  • Ability to identify risks, monitoring gaps, and drive continuous improvement.
  • Ability to influence, collaborate, and communicate effectively across technical and business teams.

Work Hours: 8

Experience in Months: 60

Level of Education: bachelor degree

Job application procedure

Application Link:Click Here to Apply Now

All Jobs | QUICK ALERT SUBSCRIPTION

Job Info
Job Category: Computer/ IT jobs in Tanzania
Job Type: Full-time
Deadline of this Job: Friday, September 25 2026
Duty Station: Dar es Salaam | Dar es Salaam
Posted: 19-09-2026
No of Jobs: 1
Start Publishing: 19-09-2026
Stop Publishing (Put date of 2030): 10-10-2076
Apply Now
Notification Board

Join a Focused Community on job search to uncover both advertised and non-advertised jobs that you may not be aware of. A jobs WhatsApp Group Community can ensure that you know the opportunities happening around you and a jobs Facebook Group Community provides an opportunity to discuss with employers who need to fill urgent position. Click the links to join. You can view previously sent Email Alerts here incase you missed them and Subscribe so that you never miss out.

Caution: Never Pay Money in a Recruitment Process.

Some smart scams can trick you into paying for Psychometric Tests.