← all jobs

[Remote] Data Infrastructure Engineer

Work from home Full-time role Hiring

Note: The job is a remote job and is open to candidates in USA. Guidehouse is a company seeking a Data Infrastructure Engineer to build and operate their data platform for AI/ML analytics. The role involves designing and implementing data ingestion pipelines, managing a data lake on AWS, and ensuring data governance and quality for analytics purposes.

Responsibilities

  • Build & Operate Data Pipelines (Batch + Streaming)
  • Design and implement batch and streaming ingestion from APIs, relational databases, file drops, event streams, and external partners
  • Build and optimize ETL/ELT pipelines to produce curated, analytics-ready datasets for reporting and ML consumption
  • Implement incremental processing patterns, change data capture (CDC) approaches where appropriate, and data contract standards
  • Deliver a Modern Lakehouse (Data Lake / Delta Lake)
  • Build and manage a scalable lakehouse on AWS object storage (e.g., S3) using open table/file formats and delta/lakehouse concepts (e.g., ACID tables, schema evolution, time travel patterns)
  • Optimize performance and cost through partitioning, compaction, lifecycle policies, and efficient compute/storage usage
  • Establish environment standards for dev/test/prod and consistent promotion across stages
  • Metadata, Governance, Lineage & Quality (Trust Layer)
  • Implement a managed metadata repository for dataset cataloging, ownership, glossary/definitions, tagging, and discoverability
  • Enable end-to-end lineage (source → transformations → consumption) to support auditability and impact analysis
  • Implement governance controls including policy-based access, data classification, retention, and secure data handling
  • Build operational data quality checks (freshness, completeness, validity, anomaly detection) and publish SLAs/SLOs
  • AWS Automation + CI/CD for Data Pipelines
  • Implement automated cloud provisioning in AWS using Infrastructure as Code (IaC) for consistent environments and secure-by-default baselines
  • Build and enhance CI/CD for data pipelines, including automated tests, validation gates, promotion workflows, and rollback strategies
  • Improve observability with metrics/logs/alerts, dashboards, runbooks, and incident response readiness
  • Cross-Team Collaboration & Documentation
  • Work closely with engineering, security, networking, and application teams to support mission needs and delivery timelines
  • Maintain high-quality engineering documentation including SOPs, system diagrams, and secure configuration baselines
  • Summarize and present findings and recommendations—both written and verbal—to technical and non-technical stakeholders

Skills

  • Must be able to OBTAIN and MAINTAIN a Federal or DoD 'PUBLIC TRUST'; candidates must obtain approved adjudication of their PUBLIC TRUST prior to onboarding with Guidehouse. Candidates with an ACTIVE PUBLIC TRUST or SUITABILITY are preferred
  • Bachelor's degree in Engineering, IT, Computer Science, or related field (or equivalent experience)
  • Minimum of FOUR (4) years experience building production data pipelines and/or data platforms
  • Strong experience implementing data ingestion and ETL/ELT workflows, including data modeling and transformation best practices
  • Hands-on experience building a data lake / delta lake (lakehouse) on AWS (or equivalent cloud) using object storage and modern table formats/patterns
  • Proficiency in SQL and one programming language commonly used for data engineering (Python preferred; Scala/Java acceptable)
  • Experience with metadata management and governance: cataloging, lineage, ownership, access controls, classification and policy enforcement
  • Experience implementing automated AWS provisioning using IaC and operating across multiple environments
  • Experience building or operating CI/CD pipelines for data workflows (testing, packaging, deployment automation, environment promotion)
  • Solid security fundamentals: IAM/least privilege, encryption, secrets management, secure SDLC practices
  • Hands-on experience with Databricks
  • Hands-on experience utilizing modern DevOps practices, including tools like Git, Terraform, Jenkins, AWS CodePipeline, and Docker
  • Experience utilizing AI-assisted coding tools (e.g., GitHub Copilot, ChatGPT, Cursor, Kiro) to safely accelerate implementation while maintaining strict code quality through testing, code reviews, and security practices
  • Knowledge graph and Graph RAG experience, including: Graph modeling and ontology/taxonomy alignment, Entity resolution and relationship extraction, Hybrid retrieval approaches combining graph traversal with semantic/vector search to improve grounding and explainability

Benefits

  • Medical, Rx, Dental & Vision Insurance
  • Personal and Family Sick Time & Company Paid Holidays
  • Parental Leave
  • 401(k) Retirement Plan
  • Group Term Life and Travel Assistance
  • Voluntary Life and AD&D Insurance
  • Health Savings Account, Health Care & Dependent Care Flexible Spending Accounts
  • Transit and Parking Commuter Benefits
  • Short-Term & Long-Term Disability
  • Tuition Reimbursement, Personal Development, Certifications & Learning Opportunities
  • Employee Referral Program
  • Corporate Sponsored Events & Community Outreach
  • Care.com annual membership
  • Employee Assistance Program
  • Supplemental Benefits via Corestream (Critical Care, Hospital Indemnity, Accident Insurance, Legal Assistance and ID theft protection, etc.)
  • Position may be eligible for a discretionary variable incentive bonus

Company Overview

  • Guidehouse offers consulting services for public and commercial markets with expertise in management, technology, and risk consulting. It was founded in 2018, and is headquartered in Washington, District of Columbia, USA, with a workforce of 10001+ employees. Its website is https://guidehouse.com.
  • More open positions

    [Remote] Senior Project Manager

    Work from home Full-time role

    [Remote] Global Account-Based Marketing (ABM) Manager

    Work from home Full-time role

    [Remote] Test Engineer

    Work from home Full-time role

    [Remote] Client Project Manager

    Work from home Full-time role

    [Remote] Lead Software Engineer

    Work from home Full-time role

    IT technician specializing in support and networks - full Remote / Home office

    Work from home Full-time role

    Experienced Customer Service Representative - 100% Remote Opportunity with Uncapped Bonuses and Weekly Commission-Based Pay

    Work from home Full-time role

    Software Engineer - PlanetScale Postgres

    Work from home Full-time role

    Werkstudent Marketing/Kommunikation - IT & Digitalisierung (d/m/w)

    Work from home Full-time role

    [Remote] Staff Software Engineer

    Work from home Full-time role

    [Remote] Remote Life Insurance Consultant

    Work from home Full-time role

    Professional Service Veterinarian

    Work from home Full-time role

    Administrative Coordinator - Client Accounting & Advisory Team

    Work from home Full-time role

    Local recruitment: Environmental Scientist II

    Work from home Full-time role

    BMC Remedy/Helix Developer

    Work from home Full-time role

    Senior Staff Software Development Engineer

    Work from home Full-time role

    Enterprise Architect

    Work from home Full-time role

    [Remote] Enterprise Account Executive, Wisconsin/Illinois

    Work from home Full-time role

    [Remote] P-3 Flight Engineer Instructor - Homeland Security

    Work from home Full-time role

    Remote Data Entry Specialist – Legal Document Management & e‑Fulfilment Support (Work‑From‑Home)

    Work from home Full-time role

    Cloud Architect for T-Cloud Public (m/f/d)

    Work from home Full-time role