Sr. Specialist, SRE – Compute Platforms

Toronto, ON, CAOn-siteSeniorFound today
Apply to this job

Free credits included. Sign up to start applying with Jobfinder.

mainframez/osaixibm isite reliability engineeringobservabilityautomationvendor management

What You’ll Do:

The Sr. Specialist, SRE – Compute Platforms serves as the enterprise technical owner for CTC's compute platforms, including IBM Mainframe (z/OS), AIX, IBM iSeries, IBM NOI, New Relic, and associated business-critical services and applications.

This role is accountable for platform governance, service ownership, lifecycle strategy, observability governance, vendor oversight, and reliability outcomes across the compute environment. Acting as the internal authority for compute platforms, the position ensures services remain reliable, resilient, secure, observable, supportable, and aligned with enterprise technology strategy, risk management, and modernization objectives.

Operational execution, administration, monitoring response, maintenance activities, and day-to-day infrastructure support are performed by HCL/Harmony and other designated support providers. The Sr. Specialist provides governance, strategic direction, technical leadership, and vendor accountability to ensure services are delivered in accordance with enterprise standards, contractual commitments, and business requirements.

The role leads platform roadmap development, technology currency initiatives, service reliability improvements, observability strategy, operational governance, vendor management, and continuous improvement programs while promoting Site Reliability Engineering (SRE) principles, operational excellence, automation, and platform sustainability. The position drives platform maturity, reduces consultant and key-person dependency, and establishes clear ownership and accountability across the enterprise compute environment.

Mainframe Platform Ownership and Lifecycle Governance

  • Provide governance and oversight of HCL/Harmony-delivered Mainframe services, including z/OS lifecycle planning, platform currency, capacity strategy, LPAR governance, and infrastructure roadmap planning.
  • Assess current platform, operating system, middleware, and Mainframe-supported application versions; identify lifecycle, supportability, operational, security, and compliance risks; and develop renewal and modernization strategies.
  • Lead platform lifecycle governance activities, including hardware refresh planning, operating system upgrade strategy, maintenance governance, change oversight, and long-term roadmap alignment.
  • Govern service ownership, support models, operational accountability, and vendor-delivered support services for Mainframe-hosted and Mainframe-dependent applications.
  • Define, govern, and periodically review monitoring and observability requirements for Mainframe infrastructure, middleware, and business-critical services.
  • Provide governance and oversight of disaster recovery readiness, resiliency planning, recovery testing, and recovery capability validation delivered by managed service providers.
  • Identify and drive resolution of operational ownership, support model, documentation, and Statement of Work (SOW) gaps through the appropriate vendors, support teams, and stakeholders.
  • Provide technical leadership and governance during major incidents, problem investigations, and corrective action planning, ensuring responsible teams execute required remediation activities.
  • Serve as the primary technical escalation and governance authority for Mainframe-related risks, service concerns, lifecycle issues, and vendor performance matters.

Monitoring, Event Management and Observability Governance

  • Provide governance and strategic oversight of enterprise monitoring, event management, and observability capabilities.
  • Establish governance standards for monitoring coverage, alerting requirements, event correlation, escalation models, service health dashboards, and operational reporting.
  • Drive continuous improvement of observability maturity, service visibility, monitoring effectiveness, synthetic monitoring capabilities, and operational intelligence.
  • Govern event management practices, including alert quality, escalation effectiveness, incident correlation, operational readiness, and service monitoring standards.
  • Define strategic direction and adoption roadmaps for observability platforms, monitoring technologies, event management tooling, and automation capabilities.
  • Act as the Compute SRE representative for enterprise monitoring strategy, operational intelligence, and observability initiatives.

Managed Service Governance and Platform Accountability

  • Provide technical governance, vendor oversight, and escalation leadership for HCL/Harmony-managed services across Mainframe, AIX, IBM iSeries, storage, backup, monitoring, and supporting infrastructure platforms.
  • Validate vendor-delivered services against contractual obligations, SOW commitments, service level expectations, operational controls, monitoring standards, and governance requirements.
  • Lead vendor performance reviews, operational scorecards, service reporting reviews, incident follow-ups, and continuous improvement initiatives.
  • Identify ownership gaps, operational risks, monitoring deficiencies, contractual concerns, and escalation requirements for leadership review.
  • Ensure vendor-delivered services are properly documented, transitioned, supportable, and aligned with the enterprise operating model.
  • Drive vendor accountability for service reliability, incident response effectiveness, operational reporting quality, remediation commitments, and lifecycle objectives.

Platform Governance, Controls and Continuous Improvement

  • Ensure operational documentation, runbooks, support procedures, monitoring standards, and technical documentation are established, maintained, and periodically reviewed by responsible support teams and service providers.
  • Support audit, compliance, and control activities related to backup, recovery, disaster recovery, monitoring controls, operational governance, and infrastructure lifecycle management.
  • Identify opportunities to improve reliability, observability, automation, resiliency, operational efficiency, and platform sustainability.
  • Participate in change, incident, problem, risk, lifecycle, and service governance forums as the Compute platform owner.
  • Provide reporting and recommendations on platform risks, technology currency gaps, vendor performance, monitoring effectiveness, remediation priorities, lifecycle initiatives, and service readiness.
  • Lead continuous improvement initiatives focused on service maturity, reliability, observability, sustainability, risk reduction, and operational excellence.

What You Bring:

  • 10+ years of experience in platform ownership, infrastructure governance, Site Reliability Engineering (SRE), service management, enterprise technology leadership, or related disciplines.
  • Strong background in Mainframe platform governance, including z/OS environments, lifecycle planning, capacity management, service ownership, and vendor-managed support models.
  • Experience governing enterprise monitoring, observability, event management, or operational intelligence practices across large-scale technology environments.
  • Strong understanding of service reliability principles, alert management, event correlation, service health monitoring, operational reporting, and risk management.
  • Knowledge of AIX, IBM iSeries, enterprise middleware, and application hosting environments supporting business-critical workloads.
  • Experience governing enterprise storage and backup services, including capacity planning, recovery validation, resiliency requirements, and restoration testing oversight.
  • Experience supporting disaster recovery governance, DR exercises, recovery reporting, remediation tracking, and resiliency planning.
  • Experience working within MSP-governed or vendor-managed service environments.
  • Strong understanding of incident, problem, change, lifecycle, operational risk, vendor governance, and service management processes.
  • Demonstrated ability to lead technical escalations, coordinate across multiple stakeholder groups, and hold vendors accountable to service expectations.
  • Strong communication skills with the ability to translate technical risk into clear operational and business impact.
  • Proven ability to influence technical direction, challenge operational assumptions, and drive accountability across diverse stakeholder groups.

Preferred Qualifications

  • Experience supporting IBM Mainframe environments, including z/OS, middleware, enterprise schedulers, file transfer platforms, or related infrastructure services.
  • Experience with observability and monitoring platforms such as IBM OMEGAMON, ITM/Tivoli, Netcool, New Relic, Dynatrace, Splunk, Elastic, ThousandEyes, or equivalent technologies.
  • Experience defining monitoring standards, observability strategies, alert rationalization programs, service health reporting, event correlation approaches, and operational intelligence improvements.
  • Experience supporting infrastructure audit controls, disaster recovery governance, compliance remediation activities, and evidence management processes.
  • Experience supporting highly regulated, high-availability, or business-critical technology environments.
  • Familiarity with ServiceNow-based intake, incident, event, problem, change, risk, vendor, and lifecycle management workflows.
  • Experience governing infrastructure refreshes, hardware lifecycle initiatives, maintenance planning, and data centre change activities.
  • Experience working with outsourced infrastructure providers, including SOW governance, SLA management, contract oversight, and vendor performance management.
  • Experience supporting SRE, service reliability, observability, platform governance, or operational excellence programs across hybrid infrastructure environments spanning Mainframe and distributed platforms.

Location : Toronto, Ontario OR Calgary, Alberta (Hybrid: In-office 4 days a week)

We’re always looking for great talent! In addition to competitive pay, we offer:

  • Comprehensive benefits and retirement programs
  • Performance incentives, Continuing Education Programs
  • Other perks to support your well-being
  • Career growth opportunities and product discounts

Broadband Salary Range: $64,000 – $106,000.

Our typical hiring range is between $64,000 and $85,000. Salary decisions are also dependent on other factors such as your experience, industry benchmarks, internal equity and other role-specific requirements. For critical roles, the compensation offering will be reviewed to ensure alignment with market rate and conditions and the unique value you bring to the role.

#LI-AK1

This posting represents an existing vacancy within our organization.

We may use artificial intelligence tools as part of our recruitment process to assist in the initial screening of resumes. All hiring decisions, including candidate evaluation, selection, and disposition, are made by human recruiters.

About Us

Canadian Tire Corporation, Limited (“CTC”) is one of Canada’s most admired and trusted companies. With more than 90 Owned Brands, over 1,600 retail locations, financial services, exemplary e-commerce capabilities, and exciting market-leading merchandising strategies. We dream big and work as one to innovate with purpose for our customers at every level of our business, investing in new technologies and products, and doubling down on top talent to drive the company forward. We offer competitive salaries and wages to CTC employees, as well as store discounts, supported learning through our Triangle Learning Academy, Canadian Tire Profit Sharing, and retirement and savings programs for eligible employees. As part of our enhanced flex benefits program, we offer mental health benefits in the amount of $5,000 per year for benefits-eligible employees and their families, including total well-being, and mental health tools and resources for all employees. Join us in helping to make life in Canada better through living and working our Core Values: we are innovators and entrepreneurs at our core, outcomes drive us, inclusion is a must, we are stronger together and we take personal responsibility. It is an especially exciting time to join CTC and its family of companies where career opportunities are wide-ranging! Join us, where there's a place for you here.

Our Commitment to Diversity, Inclusion and Belonging

We are committed to fostering an environment where belonging thrives, and diversity, inclusion and equity are infused into everything we do. We believe in building an organizational culture where people are consistently treated with dignity while respecting individual religion, nationality, gender, race, age, perceived ability, spoken language, sexual orientation, and identification. We are united in our purpose of being here to help make life in Canada better.

Accommodations

We stand firm in our Core Value that inclusion is a must. We welcome and encourage candidates from equity-seeking groups such as people who identify as racialized, Indigenous, 2SLGBTQIA+, women, people with disabilities, and beyond. Should you require any accommodation in applying for this role, or throughout the interview process, please make them known when contacted and we will work with you to help meet your needs.

Canadian Tire Corporation, Ltd.Sr. Specialist, SRE – Compute Platforms
Apply to this job