This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Observability Specialist based in Canada.
This is an opportunity to shape observability for large-scale financial technology systems serving millions of users across Africa.
You will own and evolve an observability platform spanning backend services, APIs, databases, Kubernetes, cloud infrastructure, and on-premises environments.
Your work will help engineering teams detect regressions earlier, investigate incidents faster, and make smarter reliability and performance decisions.
You’ll join a newly established Performance & Observability team with significant room to define standards, tooling, and ways of working.
The role combines hands-on engineering, platform ownership, automation, performance analysis, and close collaboration with product and infrastructure teams.
You’ll have strong autonomy, opportunities to mentor engineers, and the chance to solve complex problems at meaningful scale.
The environment values pragmatic engineering, fast iteration, simplicity, and technology that delivers measurable impact.
Accountabilities:
- Own, operate, and continuously evolve the observability platform across applications, APIs, databases, Kubernetes workloads, cloud infrastructure, and on-premises environments.
- Improve visibility into production behavior through effective metrics, logs, traces, profiling, dashboards, alerts, and service-level indicators.
- Partner with engineering and platform teams to define meaningful SLIs and improve alert quality, reducing noise while keeping alerts actionable and connected to user impact.
- Build internal tools, libraries, automation, and self-service workflows that enable engineers to instrument services, investigate incidents, analyze performance, and understand system dependencies.
- Identify reliability, latency, capacity, performance, and infrastructure cost issues before they become user-facing incidents.
- Establish observability standards, naming conventions, documentation, and training materials that can be adopted consistently across engineering teams.
- Manage and optimize observability technologies such as Datadog, Honeycomb, Sentry, Pyroscope, Prometheus, Grafana, and OpenTelemetry, balancing scalability with platform costs.
- Support the design and implementation of SLOs and help product teams adopt effective reliability practices.
- Investigate production issues and performance regressions, including profiling and analysis of application and infrastructure behavior.
- Collaborate with product, database, infrastructure, security, and engineering teams to improve system reliability and operational maturity.
Requirements:
- 5+ years of experience in observability, SRE, platform engineering, infrastructure engineering, backend engineering, or production systems engineering.
- Strong understanding of metrics, logging, distributed tracing, profiling, alerting, dashboards, SLIs/SLOs, and incident response workflows at scale.
- Experience building internal tools, automation, libraries, or platforms used by other engineers.
- Proficiency in at least one backend programming language, preferably Python.
- Hands-on experience with observability platforms such as Prometheus, Grafana, Datadog, OpenTelemetry, Jaeger, Tempo, Loki, Honeycomb, Sentry, Pyroscope, or similar technologies.
- Experience with OpenTelemetry instrumentation and collector configuration at scale.
- Familiarity with technologies such as PostgreSQL, CockroachDB, Redis, GraphQL, and Kubernetes.
- Experience analyzing and working with existing codebases and collaborating across technical teams.
- Strong communication and collaboration skills, with the ability to help other engineers build, operate, and troubleshoot reliable systems.
- Pragmatic problem-solving mindset, with good judgment around when to improve tooling, simplify solutions, or avoid unnecessary complexity.
- Ability to work autonomously, take ownership of projects, mentor less-experienced engineers, and operate effectively in a fast-moving environment.
- A proactive, growth-oriented approach and genuine interest in building reliable infrastructure with meaningful real-world impact.
Benefits:
- Remote work opportunity with flexibility and autonomy.
- Opportunity to work on large-scale, mission-driven financial technology infrastructure.
- High degree of ownership across projects, from problem definition through production monitoring.
- Professional growth through exposure to complex observability, reliability, performance, and platform engineering challenges.
- Opportunities to mentor and collaborate with experienced engineers across distributed teams.
- Inclusive and diverse international engineering environment.
- Opportunity to help define the scope, standards, tooling, and practices of a growing Performance & Observability function.
- Exposure to modern technologies including OpenTelemetry, Kubernetes, cloud infrastructure, Datadog, Honeycomb, Sentry, Pyroscope, and related observability platforms.
- Competitive compensation and benefits package, according to local employment terms.
- Flexible, autonomy-focused work culture designed to support productivity and continuous learning.