Voltage Park Logo

Voltage Park

Manager of Infrastructure Operations

Posted 24 Days Ago
Remote
2 Locations
180K-240K Annually
Senior level
Remote
2 Locations
180K-240K Annually
Senior level
Lead and mentor the 24/7 Infrastructure Operations team, ensuring system stability and performance while implementing best practices and automation.
The summary above was generated by AI

Voltage Park is seeking a highly skilled and proactive Manager of Infrastructure Operations to lead our 24/7 Infrastructure Operations team responsible for the stability, scalability, and performance of compute, storage, and platform infrastructure. This role plays a key part in delivering always-on, high-performance environments that support AI/ML training, inference, and HPC workloads at scale. The ideal candidate combines technical depth with strong leadership skills and a passion for operational excellence. 

This position offers full remote flexibility, although candidates must be based in the continental US and available to work during PST hours. Unfortunately, we are unable to provide sponsorship for this role.

Responsibilities:

  • Establish and uphold the standard practices for our expanding InfraOps team.

  • Lead and mentor a 24/7 infrastructure Operations team responsible for monitoring, maintaining, and supporting our infrastructure.

  • Develop and maintain operational runbooks, escalation procedures, and documentation for critical systems.

  • Collaborate with Infrastructure Engineering, Network operations, and Datacenter Operations and Customer Success teams to support infrastructure rollouts, upgrades, and scaling efforts.

  • Oversee observability systems (monitoring, logging, alerting) and drive continuous improvements in automation and root-cause analysis.

  • Drive adoption of “Infrastructure as Code” and automated workflows to reduce manual intervention.

  • Implement and enforce best practices for system availability, performance tuning, capacity planning, and lifecycle management.

  • Be available for on-call support during urgent system incidents.

  • Ensure compliance with security, regulatory, and organizational standards across all environments.

Qualifications:

  • Proficiency in Puppet, Terraform, and Ansible.

  • Strong scripting skills in Bash, Python, or Go.

  • Extensive experience in setting up, deploying, and managing Kubernetes clusters.

  • Proven track record of architecting, building, and delivering complex systems from inception.

  • Ability to strike a balance between pragmatic development and ideal architectures.

  • Skilled at navigating trade-offs between design, risk, cost, and outcomes.

  • Deep understanding of network protocols, network programming, Unix variants, monitoring, and security systems.

  • Excellent written and verbal communication skills.

Leadership Requirements:

  • Demonstrated ability to inspire and lead a team towards common goals, fostering a positive and collaborative work environment.

  • Proven track record of effectively delegating tasks, providing constructive feedback, and developing team members' skills.

  • Strong decision-making skills, capable of guiding the team through complex technical challenges and strategic initiatives.

  • Ability to communicate a clear vision and align team efforts with broader company objectives.

  • Experience in conflict resolution and team building, promoting diversity, equity, and inclusion within the team and the organization.

Culture:

  • Enjoy collaborating with a growing motivated team focused on execution.

  • Comfortable operating with a high degree of autonomy and able to independently prioritize tasks aligning with company objectives.

  • Possess a breadth of knowledge in your domain while also embracing the opportunity to take on diverse responsibilities.

  • Value the importance of clear communication and documentation in driving success.

Team Charter:

The 24/7 Infrastructure Operations Team ensures the stability, scalability, and performance of Voltage Park’s compute, storage, and platform systems across data centers, cloud, and edge. Supporting AI and HPC GPU environments, the team delivers proactive monitoring, automation toolsets, and continuous optimization to maintain high availability and operational excellence at all times to ensure the best possible customer experience.

Voltage Park is an equal opportunity employer and makes employment decisions on the basis of merit. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or any other characteristic under federal, state, or local law. If you require an accommodation during the job application process, please notify your recruiter. 

Compensation Range: $180K - $240K


#BI-Remote

Top Skills

Ansible
Bash
Go
Kubernetes
Puppet
Python
Terraform

Similar Jobs at Voltage Park

4 Days Ago
Remote
2 Locations
115K-145K Annually
Senior level
115K-145K Annually
Senior level
Artificial Intelligence • Cloud • Hardware • Machine Learning • Other • Software • Infrastructure as a Service (IaaS)
As a Security Engineer, you'll design and implement security solutions, automate processes, conduct vulnerability assessments, and support incident response efforts.
Top Skills: AnsibleF5Palo Alto NetworksPowershellPuppetPython
5 Days Ago
Remote
2 Locations
120K-180K Annually
Senior level
120K-180K Annually
Senior level
Artificial Intelligence • Cloud • Hardware • Machine Learning • Other • Software • Infrastructure as a Service (IaaS)
As a Platform Engineer, you'll maintain platforms, develop automation software, and ensure system reliability, leveraging strong Linux administration and scripting skills.
Top Skills: AnsibleBashCephDebianDockerElk StackGrafanaKubernetesLibvirtLinuxMaasNfsPostgresPrometheusPythonReactRedisTailwindTerraformUbuntu
5 Days Ago
Remote
USA
130K-180K Annually
Senior level
130K-180K Annually
Senior level
Artificial Intelligence • Cloud • Hardware • Machine Learning • Other • Software • Infrastructure as a Service (IaaS)
The Product Marketing Manager will develop messaging, create engaging content, manage digital strategies, and collaborate with teams to drive product awareness and adoption.
Top Skills: Adobe Premiere ProAfter EffectsCanvaFinal Cut ProHootsuiteHubspotIllustratorPhotoshop

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account