Sr. Director, Platform & AI Infrastructure
Don't apply into the void.
Most applications for this Fortive role vanish into an ATS. With jobfinder-ai, your agent finds the actual hiring manager or founder behind this opening and sends a tailored email from your own inbox — so a real person reads your pitch and replies. We then follow up until you land on the calendar.
Reach the decision-maker — $5About the role
The Sr. Director, Platform & AI Infrastructure leads the platforms that power Gordian's products and our AI transformation: cloud infrastructure, data platforms, ML/AI infrastructure, incident response, and observability. This is a builder role. You'll stand up the AI platform that our product and engineering teams build on, modernize how we run production, and establish the incident response and observability programs that scale with us. Key responsibilities AI platform and production ML. Own the AI/ML platform: GPU capacity strategy, model serving and inference latency, training and fine-tuning infrastructure, MLOps and evaluation pipelines, vector and feature stores, and the RAG and agentic patterns our product teams build on. Partner with product engineering and architecture on build-vs-buy decisions across foundation model providers and open-source. Incident and observability management. Build out the incident response program: on-call structure, severity definitions, incident command, communication standards, postmortems, and follow-through on systemic fixes. Develop the observability stack across metrics, logs, traces, and synthetics. Set SLOs and report against them. Cloud infrastructure. Operate the Azure and OCI footprint. Infrastructure-as-code, capacity planning, and reliability across CPU and GPU workloads. Data platforms. Operations, performance, HA/DR, and roadmap for Oracle, SQL Server, MongoDB, and similar. Communication. Brief executives on reliability and risk. Lead internal communication during incidents. Speak to customers when major incidents affect them. People. Lead a globally distributed team of managers and senior ICs. Maintain a strong culture and leadership bench. Required 10+ years in platform engineering, SRE, infrastructure, or AI/ML infrastructure, with 5+ leading teams. Production experience running ML/AI workloads at scale, including GPU infrastructure, model serving, MLOps, or LLM/inference platforms. Familiarity with the modern AI stack: vector databases, RAG, agent frameworks, evaluation, and the build-vs-buy tradeoffs across foundation model providers and open-source. Built or rebuilt an incident response or observability program at scale. Measurable reliability improvements (MTTR, availability, change failure rate) in a cloud environment. Effective communicator with executives, the Board, customers, and the company during incidents. Zero-trust and modern identity platforms. Originally posted on Himalayas
Ready to reach the decision-maker?
Set this role as a target and your agent does the sourcing, finds the verified email, writes the pitch, and follows up — on autopilot.
Start your hunt