Back to jobs

Engineering, Infrastructure and Operations

Senior Staff Software Engineer - Data Platform - Kubernetes - Distributed Systems - Federal

  • ServiceNow
  • Santa Clara, CALIFORNIA, United States
  • Full-time
  • Salary not listed

About the role

Company Description It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people.   Job Description Position Location: This is a  Flexible (Hybrid)  position.   Flexible  positions require 2 days per week in a ServiceNow office location.  We have offices in several locations, including San Francisco, CA; Pleasanton, CA; Santa Clara, CA; San Diego, CA; and Kirkland, WA Please Note:    This position will include supporting our US Regulated Markets. “This position requires passing a ServiceNow background screening, USFedPASS (US Federal Personnel Authorization Screening Standards). This includes a credit check, criminal/misdemeanor check and taking a drug test. Any employment is contingent upon passing the screening.   Due to Federal requirements, only US citizens, US naturalized citizens or US Permanent Residents, holding a green card, will be considered. About the team: We are seeking a  Senior Staff Software Engineer  (IC5 Level) to join our Data Platform Engineering organization.   You will be a hands-on technical lead who will lead  t he design and delivery of shared, multi-tenant platform services — Postgres, queueing/streaming, and key-value stores — on Kubernetes, and serves as the technical anchor for their availability, resilience, and operating model. What you get to do in this role: You will lead the design and delivery of shared platform services — managed Postgres, queueing/streaming, and key-value/cache — that run on Kubernetes across a large global fleet and that product teams across the company depend on, owning them from design through production operation. You will act as the technical lead on major initiatives within your team, breaking down ambiguous problems — “offer HA Postgres as a service to every cluster,” “make our queueing tier survive a zone loss” — into clear, executable designs with explicit availability, durability, and cost targets. You will define the HA, failover, disaster-recovery, and multi-region architecture for the services you own, and the Kubernetes-native automation (operators, controllers, CRDs, self-service APIs) that provisions, scales, upgrades, and fails them over without human intervention. You will partner with principal and distinguished engineers to align your work with the broader platform architecture and standards, and with product teams to set consumption contracts, tenancy models, and SLOs for shared services. You will identify technical risks early and drive them to resolution, with a strong focus on reliability, scalability, and operability — leading failure-mode analysis, game days, and post-incident reviews for stateful systems. You will spend significant time hands-on — designing, coding, and reviewing the core systems your team builds, such as operators, controllers, infrastructure automation, and platform services. You will mentor mid-level and junior engineers and raise the engineering bar through code reviews, design feedback, and pairing — particularly around distributed-systems and data-service design.   Qualifications To be successful in this role you have: Experience leveraging or critically thinking about how to integrate AI into engineering and platform work — whether using AI-powered tooling, automating operational workflows, building agentic systems for fleet visibility and operations, or reasoning about AI’s impact on how infrastructure is built and run. 12+ years of software development experience with a Bachelor's degree; OR 8+ years with a Master's degree; OR 5+ years with a PhD; OR equivalent work experience. 8+ years building production software, with solid experience operating distributed systems and running stateful workloads on Kubernetes at scale. A track record of leading the design and delivery of at least one shared, multi-tenant infrastructure service used broadly by other teams — a relational database service (Postgres or similar), a message queue or streaming platform (Kafka, NATS, RabbitMQ, or similar), or a key-value/cache service (Redis/Valkey, etcd, or similar) — including its HA, failover, and operating model. Deep understanding of high availability and failure handling: replication topologies, leader election, quorum and consensus, split-brain avoidance, backup/restore and point-in-time recovery, and designing to explicit RPO/RTO and durability targets. Strong system design skills — you can reason rigorously about consistency models, partitioning and rebalancing, capacity planning, tenant isolation, and noisy-neighbor mitigation, and communicate the trade-offs clearly to engineers and stakeholders. Hands-on experience with at least one major hyperscaler (AWS, Azure, GCP), including its core compute, networking, storage, and IAM primitives. Strong working knowledge of containers and Kubernetes (including stateful primitives: StatefulSets, CSI/persistent storage, PDBs, topology spread), CI/CD and GitOps-based delivery, and infrastructure-as-code. Strong programming skills in Go (or strong systems-language skills with a willingness to work primarily in Go). It also helps if you have: Experience building or extending Kubernetes operators that manage stateful systems (e.g., CloudNativePG, Zalando/Crunchy Postgres operators, Strimzi, Redis/Valkey operators), and opinions on when to adopt versus build. Deep expertise in one of the target systems: Postgres internals (WAL, streaming/logical replication, vacuum, connection pooling with PgBouncer/PgCat, major-version upgrades); Kafka/NATS (partitioning, ISR/replication, exactly-once semantics, consumer scaling); or Redis/Valkey (cluster mode, persistence, eviction, hot-key handling). Experience running data services across multiple regions or clusters — cross-region replication, failover orchestration, and DR testing. Experience with zero-downtime upgrades, schema/data migrations, and fleet-wide rollouts of stateful services. Experience with observability and SLOs for stateful systems — replication lag, saturation, tail latency, error budgets — and with capacity and cost management for shared infrastructure. Experience designing multi-tenancy: quotas, isolation, chargeback/showback, and self-service provisioning APIs or CRDs. Experience with container networking (CNI) and/or service mesh, and with workload identity, mTLS, and secrets management as applied to data services. Experience with managed Kubernetes (EKS/AKS/GKE), managed data services (RDS/Aurora, Cloud SQL, MSK, ElastiCache), and infrastructure-as-code tools such as Terraform or Crossplane. For positions in this location, we offer a base pay of $149,800 - $262,200 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location. Additional Information Work Personas We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service. Equal Opportunity Employer ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity,  veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.   Accommodations We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact globaltalentss@servicenow.com for assistance.  Export Control Regulations For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.  From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.