← all jobs

[Remote] Senior Software Engineer, Platform Infrastructure

Work from home Full-time role Hiring

Note: The job is a remote job and is open to candidates in USA. Moonlite AI delivers high-performance AI infrastructure for organizations involved in intensive computational research and data processing. They are seeking a Senior Software Engineer to build the infrastructure platform that connects physical infrastructure with customer services, focusing on automation, orchestration, and API development for large-scale computation.

Responsibilities

  • Design and build systems that bridge physical infrastructure (bare-metal servers, storage clusters, network fabric) with customer-facing services, enabling programmatic management of compute, networking, and storage at scale
  • Design and implement systems for provisioning and managing research computing environments including Kubernetes and SLURM clusters, enabling automated deployment, resource scheduling, and workload orchestration for distributed AI training and HPC workloads
  • Implement comprehensive orchestration systems that coordinate across compute, storage and networking to deliver unified experience for complex research workloads
  • Design and build network provisioning automation including intelligent VM placement decisions for optimal network topology, automated VLAN and subnet configuration, and software-designed networking orchestration for high-performance interconnects
  • Develop robust APIs and SDKs that enable researchers and engineering teams to programmatically provision and manage infrastructure resources across all platform domains
  • Implement comprehensive observability, telemetry, and logging systems that provide visibility into infrastructure health, performance, and utilization across the infrastructure footprint
  • Build and optimize platform services that deliver consistent high-throughput low-latency networking for demand research applications and data-intensive workloads
  • Work closely with engineering, infrastructure, and product to define requirements, drive infrastructure-product-rollouts, and improve resource lifecycle management
  • Implement platform-wide compliance and security features supporting SOC 2, ISO 27001, and enterprise regulatory requirements including comprehensive audit logging, access controls, and data residency management

Skills

  • 5+ years in software engineering with a proven track record of infrastructure platforms, distributed systems, or cloud platforms for production environments
  • Strong familiarity with Kubernetes architecture, container orchestration concepts, and experience deploying workloads in Kubernetes environments. Understanding of pods, deployments, services, and basic Kubernetes operations
  • Strong understanding of infrastructure fundamentals including compute orchestration, storage systems, networking technologies, and how they integrate together to deliver complete platform experiences
  • Experience with systems programming languages (Go, C/C++, Rust, Python) for performance-critical components is a strong plus
  • Strong experience with linux in production environments, including systems administration, performance tuning, and troubleshooting
  • Deep knowledge of bare-metal infrastructure, provisioning systems, out-of-band management, and virtualization technologies (KVM, Kubernetes, etc)
  • Proven experience designing and building APIs, SDKs, and automation frameworks that enable programmatic infrastructure management
  • Strong familiarity with cloud environments (AWS, GCP, Azure) and understanding of how to translate cloud-native patterns to bare-metal infrastructure
  • Experience with Infrastructure-as-code tools (Terraform, Ansible) and building automated deployment pipelines
  • Self-starter who can navigate ambiguity, balance pragmatic shipping with good long-term architecture, and independently drive complex technical initiatives
  • Strong written and verbal communication skills, including ability to write clear technical communication and collaborate across teams
  • Growth mindset with continuous focus on learning and professional development
  • Background provisioning or managing research computing environments (Kubernetes, SLURM, or HPC clusters)
  • Experience building internal platforms, infrastructure-as-a-service, or developer tooling
  • Background with GPU computing platforms and AI/ML infrastructure requirements
  • Knowledge of high-performance networking technologies (InfiniBand, RDMA, SR-IOV)
  • Experience with observability and monitoring platforms (Prometheus, Grafana, ELK stack)
  • Familiarity with both cloud-native and bare-metal infrastructure deployment models
  • Understanding of enterprise compliance requirements and security best practices
  • Extra points for experience with financial services technology infrastructure and understanding of trading system requirements

Benefits

  • Startup equity
  • A 6% 401(k) match
  • Fully covered health insurance premiums
  • Other comprehensive offerings to support your well-being and success as we grow together

Company Overview

  • Moonlite AI is a technology company. It was founded in 2024, and is headquartered in Chicago, Illinois, USA, with a workforce of 2-10 employees. Its website is https://www.moonlite.ai.
  • More open positions

    [Remote] Enterprise Account Manager

    Work from home Full-time role

    [Remote] Technical Services Consultant (USA)

    Work from home Full-time role

    [Remote] Data Scientist

    Work from home Full-time role

    [Remote] Program Manager-100% remote

    Work from home Full-time role

    [Remote] Account Manager, SMB

    Work from home Full-time role

    Experienced Remote Data Entry Specialist – Flexible Work Arrangements and Competitive Compensation

    Work from home Full-time role

    Coder Certified (Remote) - Surgery

    Work from home Full-time role

    Financial Advisor & Planner - Personal Strategy

    Work from home Full-time role

    Office Clerk

    Work from home Full-time role

    Job Senior Logistics Strategy and Integration Consultant

    Work from home Full-time role

    Healthcare Technology Sales (Hospital Channel Sales Exp Req, $75K-$90K + Uncapped Comm)

    Work from home Full-time role

    [Hiring] Talent Acquisition Director @Everise

    Work from home Full-time role

    At-Home with Oaknoll Care Coordinator, Full-Time

    Work from home Full-time role

    [Remote] Senior DevOps Engineer AWS, Kubernetes, Cloud Infrastructure - Remote

    Work from home Full-time role

    Remote Sales Closer | Financial Industry | $15k+/mo w/o Degree

    Work from home Full-time role

    (5) Entry Level GIS Technicians needed

    Work from home Full-time role

    [Remote-Position] Manager - Corporate Solutions Chief of Staff

    Work from home Full-time role

    Boutique Director-5th Ave Flagship

    Work from home Full-time role

    Staff Data Engineer, Platform Engineering

    Work from home Full-time role

    Remote Premier Service Consultant – Inbound Sales, Customer Care & Technology Solutions (Work From Home)

    Work from home Full-time role

    Legal Counsel - Regulatory Compliance, Product and Privacy

    Work from home Full-time role